Tokens, in a blink.⚡️ Ultra-low latency inference. contact@tilert.ai

Filter
Exclude
Time range
-
Minimum likes
Great to see this result highlighted by @SemiAnalysis_ 🚀 We’re excited about how far we can push decode interactivity on GPUs — and there’s still much more to explore with TileRT. More to come!
ALERT🚨 TILERT from @TileRT_AI TO BOOST DECODE INTERACTIVITY BY 1.9X AT THE SAME PER-TOKEN COST ON THE SAME NVIDIA BLACKWELL GPUs. What does this mean for Groq/Cerebras/SambaNova? 👇️ 1/5🧵
4
1,423
Huge thanks to @SemiAnalysis_ for the deep dive into TileRT and for putting our work through InferenceX. We're excited to see what ultra-high-interactivity inference on GPUs can unlock — and this is just the beginning.
Ultra-High Interactivity on NVIDIA GPUs? TileRT InferenceX Can TileRT software on NVIDIA GPU compete with Cerebras, Groq LPU, SambaNova? Batch Size 1, Disaggregated engine, High throughput prefill engine, High interactivity decode engine newsletter.semianalysis.com/…
1
3
1,082
🚀 Excited to bring TileRT Decode to the @vllm_project ecosystem. Pair vLLM prefill with TileRT's latency-optimized decode through the V1 connector interface—no forks, no patches. Thanks to the @vllm_project team and @inferact for the collaboration!
Once prefill and decode are disaggregated, the decode side becomes a choice: pluggable, swappable per workload. The @TileRT_AI team just shipped a concrete example: vLLM prefill paired with TileRT's latency-optimized decode engine through vLLM V1's connector interface, with zero changes to vLLM. Native vLLM decode stays the default for throughput. TileRT decode is there for latency-bound work (agents, real-time assistants). Same OpenAI-compatible surface, same prefix caching, so switching is a routing change. No fork, no patches, both decode pools behind one stock vLLM prefill pool. TileRT reports around 618 tok/s single-user decode on GLM-5.1-FP8 (8× B200) with MTP, roughly 2× its no-MTP baseline. At peak acceptance, up to nearly 800 tok/s. Thanks to the @TileRT_AI team and @inferact for the collaboration. 🔗 vllm.ai/blog/2026-07-14-vllm…
1
2
6
2,003
Replying to @vllm_project
Thanks to the @vllm_project team for making this possible. The V1 connector interface opens the door to a more modular inference stack, and we're excited to bring a latency-optimized decode backend to the ecosystem. More to come! 🚀
1
3
11
2,098
GLM高速版也是我们支持的,感谢关注♥️
2
62
⚡️ System & Model Co-design: Deep technical synergy with the MiMo team on FP4/FP8 mixed quantization and production-grade DFlash.
5
1,325
⚡️ Heterogeneous Workers & Warp Specialization: Breaking the serial pace to orchestrate specialized worker groups not just within a single SM, but scaling across the entire GPU execution domain.
1
5
1,271
⚡️ Tile-grained Pipelining: Deeply overlapping memory movement, tensor computation, and communication at the physical tile level.
1
5
1,149
⚡️ Persistent Kernels: The entire compute pipeline runs continuously inside the GPU, enabling full-stack continuous prefetching and erasing operator boundaries.
5
865
How did we push a 1 Trillion parameter MoE model past the 1,000 TPS barrier on a standard 8-GPGPU node with @XiaomiMiMo? 🚀 It’s not just a faster kernel. It’s a total execution model revolution. Key technical breakthroughs inside TileRT:
🚀 1,000+ TOKENS/S ON A 1T MODEL! 🚀 We are thrilled to release Xiaomi MiMo-V2.5-Pro-UltraSpeed in collaboration with @TileRT_AI , breaking the 1,000 tokens/s output speed on a 1 Trillion parameter model for the FIRST TIME! Not wafer-scale integration like Cerebras. Not pure on-chip SRAM chips like Groq. We achieve 1,000 tps on a 1T MoE model using just a SINGLE, STANDARD 8-GPGPU NODE. Read the full technical deep dive:mimo.xiaomi.com/blog/mimo-ti… Want to experience the future of real-time AI? 👉 Apply for UltraSpeed now: platform.xiaomimimo.com/ultr… ⏳ Limited-Time Access: Application-based · Jun 8 – Jun 23 (PDT) 💬 Chat Experience: Completely FREE for a limited time — try the blazing-fast web chat now. ⚡ UltraSpeed API: Just 3x the price for a ~10x boost in output experience. 🤝 Enterprise & Large-Scale Needs: business-mimo@xiaomi.com
8
7
57
10,455
Replying to @XiaomiMiMo
Proud to core-build this with the MiMo team! Breaking 1,000 TPS on a 1T model with standard 8-GPU nodes is just the beginning of the Speed Scaling era. Technical deep dive coming on our channel! 🚀⚡️
1
78
9,564
Proud to core-build this with the MiMo team! Breaking 1,000 TPS on a 1T model with standard 8-GPU nodes is just the beginning of the Speed Scaling era. Technical deep dive coming on our channel! 🚀⚡️
🚀 1,000+ TOKENS/S ON A 1T MODEL! 🚀 We are thrilled to release Xiaomi MiMo-V2.5-Pro-UltraSpeed in collaboration with @TileRT_AI , breaking the 1,000 tokens/s output speed on a 1 Trillion parameter model for the FIRST TIME! Not wafer-scale integration like Cerebras. Not pure on-chip SRAM chips like Groq. We achieve 1,000 tps on a 1T MoE model using just a SINGLE, STANDARD 8-GPGPU NODE. Read the full technical deep dive:mimo.xiaomi.com/blog/mimo-ti… Want to experience the future of real-time AI? 👉 Apply for UltraSpeed now: platform.xiaomimimo.com/ultr… ⏳ Limited-Time Access: Application-based · Jun 8 – Jun 23 (PDT) 💬 Chat Experience: Completely FREE for a limited time — try the blazing-fast web chat now. ⚡ UltraSpeed API: Just 3x the price for a ~10x boost in output experience. 🤝 Enterprise & Large-Scale Needs: business-mimo@xiaomi.com
7
8
40
7,540
Replying to @zRdianjiao
Huge milestone. Grateful to @Zai_org for the partnership. Flagship quality at 400 tok/s is just the start of what we can do together. 🔥
2
36