Fastest inference anywhere: open models on-prem, hosted or on-device. Backed by @ycombinator W26

Meet Wally. Our inference stack for open frontier models, built to be the fastest place to run them. Performance snapshot: GLM-5.3 Flash: 380 tok/s. GLM-5.3 Max: 790 tok/s Qwen3.8-27B: 485 tok/s DeepSeek-V4.1 Flash: 615 tok/s 1/7
63
51
705
61,561
RunAnywhere retweeted
It's even cooler with open models. You can run GLM-5.3 Flash, DeepSeek V4.1 Flash and MiMo-V2.6-Pro inside Claude Code, super fast. $5 free credits to try it, so why not?
The seaon has changed and Claude Code is now cool again! I love what's cool and what's not keeps flipping every 6 months in AI, it's so fast
3
6
86
We raced GLM-5.3 Flash on Wally against OpenAI's GPT-6 Luna. Same prompt, live, both at max reasoning. Wally built the page in 118 s. Luna took 363 s. An open model, 3x faster. $5 free to try it.
6
4
33
1,854
RunAnywhere retweeted
The funniest part is closed-source model APIs declining to defend against the attacks due to safeguards, so they used a version of GLM 5.2 to defend themselves. This is exactly why we are building Wally(@RunAnywhereAI), an inference stack built around the idea that teams should be able to run the weights they choose, without restrictions, at the fastest speeds possible while staying incredibly efficient.
Thank you @jnbarrot & @UN for inviting me to share our lessons to the Security Council Being the first company to disclose an agent cyberattack taught us that we need a lot more transparency in AI and more open-source AI to fight asymmetry and empower defenders!
3
3
15
616
RunAnywhere retweeted
Codex down? Don't stop working. MiMo-V2.6 Pro beats GPT-6 Sol on xhigh on Artificial Analysis and it's live on wally at blazing fast speeds! $5 free credits, and at our pricing that's plenty to get through your work. runwally.com
3
7
381
We raced GLM-5.3 Flash on Wally against @Zai_org 's own paid fast tier, FlashX. Same prompt, live. Wally: 363 tok/s. FlashX: 143 tok/s. 2.5x faster, 3.5x cheaper. Why are you still using slow AI? Try Wally, $5 free.
Faster GLM-5.3-Flash is now live: up to 200 tokens/s. Model code: glm-5.3-flashx. Priced at 2.5× GLM-5.3-Flash on both the Coding Plan and API. Open to all API users. Coding Plan users can apply here: docs.google.com/forms/d/e/1F…
10
10
90
15,173
Feel the speed yourself, Drop this into your favorite coding agent to get started: "set up runanywhere.ai/SKILL.md" runwally.com
1
3
764
RunAnywhere retweeted
In two months, Open models took about half the spend Anthropic lost. Open models are so good now that teams are moving money out of closed models and into them. So we built Wally to run the best of them as FAST as they go. $5 free credits: runwally.com
Spend of OpenAI vs Anthropic vs Open (Vercel AI Gateway, last 2 months) • Anthropic still #1 in spend, but went 69% → 40% • OpenAI: 10% → 24% in spend • GPT-6 Astra + GPT 5.6 Sol are ripping • OpenAI now leads in tokens # • Kimi K3 + DeepSeek took ~half of Anthropic's loss • Opus 5.5 is up to 10% of spend in 2 days • OpenAI is 62% of image generations Watch here: token-race.vercel.app
1
3
12
541
RunAnywhere retweeted
Same prompt: "build a Doom-style FPS..." GPT Sol 6 Max vs MiMo-V2.6-Pro on @RunAnywhereAI's Wally. MiMo: 2.2x faster, 5.7x cheaper, and the game just looks more "Doom". @XiaomiMiMo shipped their smartest open model yet, so we put it on Wally. Now it beats the frontier on speed and cost. Once you feel these speeds going back to gpt/claude models just does not make sense. Free to start👇
MiMo-V2.6-Pro is live on Wally. Day 1. @XiaomiMiMo's new open-weights flagship, top of the @ArtificialAnlys Intelligence Index for open-weight models. 320 tok/s, 2.4x the model maker's own API, in <24h. 1/4
4
10
397
Run your company brain on Wally using @garrytan GBrain, at 10x speed and 10x less cost. 5$ free on signup!
Replying to @garrytan
My current setup is Hermes running the new @XiaomiMiMo MiMo v2.6 pro, using Wally (@RunAnywhereAI ) at 350 tok/seconds, running at lightning speed, with connected to GBrain.
1
3
11
670
RunAnywhere retweeted
We just put the smartest open-weight model inside Wally, and now it's blazing fast too! A one-shot Crossy Road build using MiMo-V2.6 Pro, pitted against @XiaomiMiMo's own API. Wally finished 2.8× faster. If you don't believe our numbers, try Wally yourself. $5 free on signup.
MiMo-V2.6-Pro is live on Wally. Day 1. @XiaomiMiMo's new open-weights flagship, top of the @ArtificialAnlys Intelligence Index for open-weight models. 320 tok/s, 2.4x the model maker's own API, in <24h. 1/4
4
13
1,114
MiMo-V2.6-Pro is live on Wally. Day 1. @XiaomiMiMo's new open-weights flagship, top of the @ArtificialAnlys Intelligence Index for open-weight models. 320 tok/s, 2.4x the model maker's own API, in <24h. 1/4
Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. 🔹 Two omnimodal models, advancing through scaled reinforcement learning 🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks 🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models 🔹 Stronger coding, computer use, 3D reasoning and creative capabilities 🔹 Open model weights, technical report, RL environments and training code Blog:mimo.xiaomi.com/mimo-v2-6
7
13
58
12,480
@ArtificialAnlys lists exactly one endpoint for MiMo-V2.6-Pro: Xiaomi's own API, at 134 tok/s(09/21/2026). Wally, same weights, measured the same way, p50: 320 tok/s. 2.4x the model maker's own API on decode. On day one. 3/4
1
1
6
279
Wally lets you run the open frontier models on day ZERO, Faster than anyone else. To get Started: Drop this into your favorite coding agent: "set up runanywhere.ai/SKILL.md" wally claude-code -m mimo-v2.6-pro $5 in free credits runwally.com 4/4
1
6
241
RunAnywhere retweeted
MiMo-v2.6-Pro, launching in a couple hours on @RunAnywhereAI's Wally Fastest speeds ,as always(exact numbers with the launch) Not just fastest inference, also fastest to every new model launch from here on. Stay Tuned!
Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. 🔹 Two omnimodal models, advancing through scaled reinforcement learning 🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks 🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models 🔹 Stronger coding, computer use, 3D reasoning and creative capabilities 🔹 Open model weights, technical report, RL environments and training code Blog:mimo.xiaomi.com/mimo-v2-6
1
4
5
458
RunAnywhere retweeted
btw this costed ¢28 on @RunAnywhereAI Wally, and we're 3.1× faster than @FireworksAI_HQ and 4.5x than Model maker @Zai_org .
Instead of the benchmarks telling you how fast we are, check this out: I built this whole site with GLM-5.3 Flash on Wally. The build was only 2 minutes, QA and fixes took 8 more minutes. 10 minutes for this quality is insane. Might never go back to Claude.
2
2
6
470
RunAnywhere retweeted
We just launched Wally, our inference stack for open frontier models. We're not just fast. We're the fastest. For GLM-5.3 flash, wally has 43% more throughput than Nebius, 3.1x Fireworks, 4.5x Z ai. The craziest part... This entire launch video was made by GLM-5.3 Flash running on Wally + Blender.
Meet Wally. Our inference stack for open frontier models, built to be the fastest place to run them. Performance snapshot: GLM-5.3 Flash: 380 tok/s. GLM-5.3 Max: 790 tok/s Qwen3.8-27B: 485 tok/s DeepSeek-V4.1 Flash: 615 tok/s 1/7
9
3
24
2,860
Meet Wally. Our inference stack for open frontier models, built to be the fastest place to run them. Performance snapshot: GLM-5.3 Flash: 380 tok/s. GLM-5.3 Max: 790 tok/s Qwen3.8-27B: 485 tok/s DeepSeek-V4.1 Flash: 615 tok/s 1/7
63
51
705
61,561
Every model on wally runs at state-of-the-art speed. That matters the most for realtime apps:Voice pipelines with an LLM in the loop, coding agents, browser automation, live copilots. On our GPUs, or yours(BYOC/on-prem) 6/7
2
1
14
3,745
Get Started: Drop this into your favorite coding agent: "set up runanywhere.ai/SKILL.md" or Curl : curl -fsSL raw.githubusercontent.com/Ru… | sh Windows: irm raw.githubusercontent.com/Ru… | iex wally login wally claude-code -m glm-5.3-flash $5 free on signup, no card. Standard per-token after. runwally.com 7/7
2
2
23
3,592