Inference serving you can trust.

Israel
MoonMath.ai retweeted
On this holiest day in the Jewish calendar, I want to share a list of PRs to @sgl_project and @vllm_project focused on low-level performance optimizations for @AIatAMD CDNA3 × Kimi K3 that we at @zroai_ opened over the past year. We touched everything from bugs to new attention kernels. We are now working on CDNA4 (thanks, AMD!). I’d also like to take this opportunity to ask forgiveness from AMD: you ship great hardware, and while your software is not as good, you do your best to improve it, and you are much more welcoming to startups than I gave you credit for. github.com/vllm-project/vllm… github.com/vllm-project/vllm… github.com/vllm-project/vllm… github.com/sgl-project/sglan… github.com/sgl-project/sglan… github.com/sgl-project/sglan… github.com/sgl-project/sglan… github.com/sgl-project/sglan…
3
21
1,318
MoonMath.ai retweeted
Abliteration is a technique for selectively removing a model’s learned refusal behavior while preserving its underlying capabilities. We’ve been running dolly1-security on Zro: an abliterated GLM-5.3 variant, tuned specifically for cybersecurity workloads. We kept it quiet while external cybersecurity experts tested it and gave us the green light. Dolly1-security runs on our own hardware and, as usual, it runs fast. The demo here was run by an AI agent powered by Dolly1 itself. The agent orchestrated the experiment, calling both Dolly1 and DeepSeek along the way. Astra refused.
2
13
1,376
MoonMath.ai retweeted
Dear mathematicians, if you don’t want to be spied on, and get fast inference, use Zro
Dear mathematicians, if you don't want to be spied on use venice.ai
1
4
260
MoonMath.ai retweeted
Auto routing is now live on Zro. Run your agent with --model zro/auto and Zro selects from available models for you. For the full routing pool, enable all regions on your API key. zro launch claude --model zro/auto
1
2
8
327
MoonMath.ai retweeted
1/ Here’s a demo showing how I generated tokens for free on @OpenRouter using the @FireworksAI_HQ endpoint for GLM-5.3
2
6
20
6,296
MoonMath.ai retweeted
We’re excited to share that AMD has provided Zro with evaluation compute to accelerate Kimi K3 inference serving on AMD MI325X (CDNA3). Our goal is to push the performance of frontier-scale LLM serving on AMD hardware, and upstream the work so the broader ecosystem can benefit. We started by building a Kimi K3 serving baseline on @sgl_project and are now working across several layers of the stack, from MXFP4 MoE execution and attention to KV-cache efficiency and speculative decoding. The first pieces of this work are already upstream for review in SGLang, split across 4 PRs: → MXFP4 MoE on gfx942 github.com/sgl-project/sglan… → Split-KV verify optimization github.com/sgl-project/sglan… → Chunked prefix KV github.com/sgl-project/sglan… → Zro MLA attention backend github.com/sgl-project/sglan… More performance results to come. Thanks @AIatAMD and @sgl_project for supporting the work.
4
13
2,079
MoonMath.ai retweeted
GLM‑5.3 Flash is now live on Zro. We’re giving it special treatment from day one, with a focus on fast, reliable serving. More performance improvements are coming over the next few days.
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: huggingface.co/zai-org/GLM-5… API: docs.z.ai/guides/llm/glm-5.3… Coding Plan: z.ai/subscribe ZCode: zcode.z.ai/en Chat: chat.z.ai AutoClaw: autoclaw.z.ai
2
5
15
1,187
MoonMath.ai retweeted
Joining the party with our first PR merged: a fused sparse-MLA prefill path for DSA. Running GLM-5.2 on 8xTPUv7x, this reduced long-context prefill time by up to 4.05× at 48K context while maintaining accuracy. github.com/sgl-project/sglan…
Big news: @Google and @RadixArk are partnering to bring @sgl_project to Google Cloud TPUs! ✅ Run SGLang on TPU today via SGL-JAX ✅ Coming soon: SGL-torchtpu for a PyTorch-native experience ✅ No more migration tax—just ultimate flexibility for devs Learn more 🛠️: goo.gle/3U2JVOC #GoogleCloud #TPU #SGLang #AI #DevRel
4
6
1,750
MoonMath.ai retweeted
We implemented this banger, and it's now merged into @sgl_project main 🎉 PR: github.com/sgl-project/sglan…
🚀 Excited to share our CVPR 2026 paper: 🌈Spectrum: Adaptive Spectral Feature Forecasting for Diffusion Sampling Acceleration Diffusion models generate stunning images/videos — but sampling is still slow because every output requires many expensive DiT forward passes. What if we could skip most of them? 🔗 Project: hanjq17.github.io/Spectrum/ 📄 Paper: arxiv.org/abs/2603.01623 Community-contributed ComfyUI available for 10+ image/video diffusion models: github.com/hanjq17/Spectrum#… Amazing collaboration with Juntong Shi, Puheng Li @lphLeo623 , Haotian Ye @haotian_yeee , Qiushan Guo @QiushanGuo_HKU , and Stefano Ermon @StefanoErmon ! Check out our poster session at ExHall A 664 on June 7th!! ✈️ Denver, CVPR 2026
4
15
2,625
MoonMath.ai retweeted
🚀
If opencode go's deepseek inference has also slowed down for you and you're looking for an alternative check out @zroai_ zro.moonmath.ai - it's fast af shoutout @OmerShlomovits for letting me try it out!
1
1
6
341
MoonMath.ai retweeted
Tldr- 1. For the model we tested, Fireworks appears to use off-the-shelf @vllm_project with a relatively simple routing setup 2. As a customer, you can exploit that routing behavior to improve your own inference economics
2
4
38
11,765
MoonMath.ai retweeted
one command to launch Prime Intellect and use open-weight frontier models optimized for agentic usage. most powerful model: > zro launch prime --model kimi-k3 fastest model: > zro launch prime --model deepseek-v4-flash-0731
Introducing Prime Agent: A self-improving RLM harness for coding and long-running autonomous tasks. Designed to be both token-efficient and expressive through programmatic tool calling, context as a variable, multi-agent messaging, and a self-modifiable harness state.
9
7
69
254,100
BIG:
Today we added @AMD MI325X as a new serving backend for @zroai_. Kimi K3 is the first model running on it, powered by @sgl_project. We believe this is the first production support for the official Kimi K3 weights on CDNA3. With 256 GB of HBM per GPU, the full model fits on a single 8×MI325X node. This is likely the lowest-cost hardware configuration capable of serving Kimi K3. The benchmark below shows solid serving performance, even before adding any of MoonMath’s custom kernels. Huge thanks to @digitalocean for their partnership. ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Here is everything we changed to make Kimi K3 work on CDNA3 with SGLang: 1. Rerouted MXFP4 experts to AITER’s Triton GEMMs: AITER’s default 4-bit MoE path selects a FlyDSL kernel on gfx942, then fails during compilation. Its A4W4 kernels are gfx950-only. AITER states this in the source, and AMD’s MI35x tests note that “MXFP4 does not register on gfx942.” Without this change, K3’s MoE has no working execution path on CDNA3. We enable the Triton route automatically on gfx942. 2. Fixed an AITER GEMM crash on sliced activations: K3 passes the tuned GEMM path a 1,536-wide slice of a 2,112-wide tensor. The launcher rejects its strides and throws; under graph capture, that becomes fatal instead of falling back. The tricky part is that PyTorch ignores the stride of size-1 dimensions, so the view reports as contiguous and .contiguous() does nothing. The fix explicitly checks for the canonical strides required by the launcher. 3. Fixed a recent regression in K3’s SiTU activation: An August 1 upstream commit added an unguarded CUDA-only include, breaking the kernel build on every AMD GPU. Because the failure appears during graph capture, it looks like a graph issue rather than a missing kernel. We route ROCm to SGLang’s equivalent Triton implementation. Upstream, a two-line ifndef USE_ROCM guard would fix MI355X as well. 4. Selected the required graph-capture mode up front on ROCm: Upstream replaced graph recapture with a validator that raises an error. Under speculative decoding, capturing a mode that is too weak now kills the run instead of triggering a recapture. We force the correct hidden mode before capture begins. 5. Added support for 12 MLA heads per rank: With TP8, K3’s 96 heads become 12 heads per rank. AITER has no MLA kernel for that shape: it rejects the configuration at startup and can later call abort() from C++ without a Python stack trace. We bypass the startup assertion and zero-pad the query heads from 12 to 16 so a supported kernel can run. 6. Enabled the configuration the stack actually requires: The working setup needs: SGLANG_USE_AITER=1 SGLANG_AITER_K3_OPT=1 <head-padding flag> --trust-remote-code SGLANG_USE_AITER=1 enables the AITER MoE route. Without SGLANG_AITER_K3_OPT=1, expert weights are padded to 256 and TP8 runs out of memory during loading. Missing --trust-remote-code fails only after loading roughly 1.42 TiB of weights.--kv-cache-dtype fp8_e4m3 is not inherently required by K3 or ROCm; our current attention path requires it because it does not yet include a BF16-KV kernel. All credit goes to the team - I am here simply to report their work 💐
2
322
MoonMath.ai retweeted
The updated dsv4-Flash-0731 feels unfairly good :) We added day-@zroai_ support serving it from our EU infra. And, just like with Kimi K3, we cut the price 🤠
🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta! 🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! 👇 🔷 The official V4-Flash now natively supports the Responses API format and is fully adapted for Codex! Check out the configuration details in our official API docs: api-docs.deepseek.com/quick_…
2
8
2,563
MoonMath.ai retweeted
We believe we’re one of the first non-Day 0 serving platforms to support Kimi K3. So we cut prices 🤠 We’ve built a full system around this amazing model. Proud of the team! and shout-out to @sgl_project and @AIatAMD !! Try it: zro.moonmath.ai Please help us get listed on @OpenRouter 🙏
Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale. Model weights: huggingface.co/moonshotai/Ki… Tech report: github.com/MoonshotAI/Kimi-K… Tech blog: kimi.com/blog/kimi-k3
36
19
393
911,540
MoonMath.ai retweeted
Great work by the AMD @sgl_project team on enabling nightly disaggregated serving CI to improve code quality! It has already caught and prevented 2 massive bugs from reaching customers, as we explained before 👇️ 1/7🧵
4
4
164
37,464
We launched on @producthunt this week, finished as #2 Product of the Day! Zro is private, multi-region-hosted inference for coding agents with zero data retention and no training on your prompts. Link in first comment. Thanks to everyone who's supported us ❤️
8
2
17
1,521
👀👀
Humbled! Casually launched Zro on Product Hunt for family and friends. Next thing, PH adopts the privacy-by-math narrative and features us. Zro is now #2 on the PH leaderboard, and we are by no means ready for all this traffic 😅
2
1
556