Tokens, in a blink.⚡️ Ultra-low latency inference. contact@tilert.ai

Pinned Tweet
Proud to core-build this with the MiMo team! Breaking 1,000 TPS on a 1T model with standard 8-GPU nodes is just the beginning of the Speed Scaling era. Technical deep dive coming on our channel! 🚀⚡️
🚀 1,000+ TOKENS/S ON A 1T MODEL! 🚀 We are thrilled to release Xiaomi MiMo-V2.5-Pro-UltraSpeed in collaboration with @TileRT_AI , breaking the 1,000 tokens/s output speed on a 1 Trillion parameter model for the FIRST TIME! Not wafer-scale integration like Cerebras. Not pure on-chip SRAM chips like Groq. We achieve 1,000 tps on a 1T MoE model using just a SINGLE, STANDARD 8-GPGPU NODE. Read the full technical deep dive:mimo.xiaomi.com/blog/mimo-ti… Want to experience the future of real-time AI? 👉 Apply for UltraSpeed now: platform.xiaomimimo.com/ultr… ⏳ Limited-Time Access: Application-based · Jun 8 – Jun 23 (PDT) 💬 Chat Experience: Completely FREE for a limited time — try the blazing-fast web chat now. ⚡ UltraSpeed API: Just 3x the price for a ~10x boost in output experience. 🤝 Enterprise & Large-Scale Needs: business-mimo@xiaomi.com
7
8
40
7,536
Great to see this result highlighted by @SemiAnalysis_ 🚀 We’re excited about how far we can push decode interactivity on GPUs — and there’s still much more to explore with TileRT. More to come!
ALERT🚨 TILERT from @TileRT_AI TO BOOST DECODE INTERACTIVITY BY 1.9X AT THE SAME PER-TOKEN COST ON THE SAME NVIDIA BLACKWELL GPUs. What does this mean for Groq/Cerebras/SambaNova? 👇️ 1/5🧵
4
1,423
Huge thanks to @SemiAnalysis_ for the deep dive into TileRT and for putting our work through InferenceX. We're excited to see what ultra-high-interactivity inference on GPUs can unlock — and this is just the beginning.
Ultra-High Interactivity on NVIDIA GPUs? TileRT InferenceX Can TileRT software on NVIDIA GPU compete with Cerebras, Groq LPU, SambaNova? Batch Size 1, Disaggregated engine, High throughput prefill engine, High interactivity decode engine newsletter.semianalysis.com/…
1
3
1,082
Ultra-High Interactivity on NVIDIA GPUs? TileRT InferenceX Can TileRT software on NVIDIA GPU compete with Cerebras, Groq LPU, SambaNova? Batch Size 1, Disaggregated engine, High throughput prefill engine, High interactivity decode engine newsletter.semianalysis.com/…
4
13
134
44,748
🚀 Excited to bring TileRT Decode to the @vllm_project ecosystem. Pair vLLM prefill with TileRT's latency-optimized decode through the V1 connector interface—no forks, no patches. Thanks to the @vllm_project team and @inferact for the collaboration!
Once prefill and decode are disaggregated, the decode side becomes a choice: pluggable, swappable per workload. The @TileRT_AI team just shipped a concrete example: vLLM prefill paired with TileRT's latency-optimized decode engine through vLLM V1's connector interface, with zero changes to vLLM. Native vLLM decode stays the default for throughput. TileRT decode is there for latency-bound work (agents, real-time assistants). Same OpenAI-compatible surface, same prefix caching, so switching is a routing change. No fork, no patches, both decode pools behind one stock vLLM prefill pool. TileRT reports around 618 tok/s single-user decode on GLM-5.1-FP8 (8× B200) with MTP, roughly 2× its no-MTP baseline. At peak acceptance, up to nearly 800 tok/s. Thanks to the @TileRT_AI team and @inferact for the collaboration. 🔗 vllm.ai/blog/2026-07-14-vllm…
1
2
6
2,003
How did we push a 1 Trillion parameter MoE model past the 1,000 TPS barrier on a standard 8-GPGPU node with @XiaomiMiMo? 🚀 It’s not just a faster kernel. It’s a total execution model revolution. Key technical breakthroughs inside TileRT:
🚀 1,000+ TOKENS/S ON A 1T MODEL! 🚀 We are thrilled to release Xiaomi MiMo-V2.5-Pro-UltraSpeed in collaboration with @TileRT_AI , breaking the 1,000 tokens/s output speed on a 1 Trillion parameter model for the FIRST TIME! Not wafer-scale integration like Cerebras. Not pure on-chip SRAM chips like Groq. We achieve 1,000 tps on a 1T MoE model using just a SINGLE, STANDARD 8-GPGPU NODE. Read the full technical deep dive:mimo.xiaomi.com/blog/mimo-ti… Want to experience the future of real-time AI? 👉 Apply for UltraSpeed now: platform.xiaomimimo.com/ultr… ⏳ Limited-Time Access: Application-based · Jun 8 – Jun 23 (PDT) 💬 Chat Experience: Completely FREE for a limited time — try the blazing-fast web chat now. ⚡ UltraSpeed API: Just 3x the price for a ~10x boost in output experience. 🤝 Enterprise & Large-Scale Needs: business-mimo@xiaomi.com
8
7
57
10,455
⚡️ System & Model Co-design: Deep technical synergy with the MiMo team on FP4/FP8 mixed quantization and production-grade DFlash.
5
1,325
⚡️ Heterogeneous Workers & Warp Specialization: Breaking the serial pace to orchestrate specialized worker groups not just within a single SM, but scaling across the entire GPU execution domain.
1
5
1,271
⚡️ Tile-grained Pipelining: Deeply overlapping memory movement, tensor computation, and communication at the physical tile level.
1
5
1,149
⚡️ Persistent Kernels: The entire compute pipeline runs continuously inside the GPU, enabling full-stack continuous prefetching and erasing operator boundaries.
5
865
感谢鸭哥的博客推荐,我们还会有新的惊喜,欢迎继续关注❤️
智谱 GLM-5.1 高速版 API 达到 400 tokens/s。这不是优化得更快,而是从执行模型层面重构了 GPU 推理。深度分析了 TileRT 的技术原理,以及推理速度为什么正在成为 AI API 的第二条竞争轴。 yage.ai/share/glm51-highspee…
1
3
850
TileRT❤️Z.ai 🚀🚀🚀
GLM-5.1-highspeed is coming, 400 tokens per second. Very expensive, but bring a new possibility.
3
769
TileRT retweeted
GLM-5.1-highspeed is coming, 400 tokens per second. Very expensive, but bring a new possibility.
41
25
598
52,057