One API to access the best AI models in the world. Faster agents, lower costs.

Which is your pick? Both live on commonstack.ai now.
4
11
17
593
New research we’re excited about: MERA powered by commonstack.ai Evolve small models with SkillBook + routing + adapters so they take on more of the agent workload over time, instead of always defaulting to the biggest (and most expensive) model. Faster agents. Lower costs. Better systems. Paper + code 👇
🚀 MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale How can we make small models stronger for on-device agent deployment? MERA uses stronger models to guide an iterative loop of RL/GRPO, skill learning, and router optimization. Student failures become verified demonstrations, reusable SkillBook procedures, and LoRA updates, helping the small model take on more work over time. 🔥 Results: Qwen2.5-Coder-1.5B: 28.7% → 49.7% coding pass Qwen3.5-2B on TAU-2: 14/35 → 18/35 Fine-tuned 2B matches an unadapted 4B model Don’t just route around small models. Evolve them. 📄 arxiv.org/abs/2608.10333 💻 github.com/yh-yao/MERA-Evolv…
10
5
15
1,035
Opus 5. Refreshed homepage. More to come.
8
7
25
2,907
Commonstack retweeted
used @commonstack_ai and built lumenforge a bio lumen experimental architecture studio for fun. created with @Kimi_Moonshot latest K3 which has been a absolute stunner when it comes to assisting in design. hybrid seed > activation > articulation > symbiosis > metropolis
Kimi K3, the latest frontier by Moonshot AI is now available on @commonstack_ai! It is packed with 2.8T parameters with KDA enabling up to 6.3x faster in decoding. Built for self evolving and long horizon agentic tasks. Give the moon a shot 🌕
10
7
44
1,068
Kimi K3 Subscriptions closed? We’ve got you covered. Run your next task with Kimi K3 through Commonstack now.
Kimi K3, the latest frontier by Moonshot AI is now available on @commonstack_ai! It is packed with 2.8T parameters with KDA enabling up to 6.3x faster in decoding. Built for self evolving and long horizon agentic tasks. Give the moon a shot 🌕
6
20
1,091
A glimpse into a month of production traffic through Commonstack: -12.9M+ requests -168B+ tokens -70 models -20%+ savings off list prices On to the next.
9
10
36
4,399
GPT-5.6 family is live on Commonstack. Route your agents to Sol, Terra and Luna. openai/gpt-5.6-sol openai/gpt-5.6-terra openai/gpt-5.6-luna Same key as your Claude, Gemini, DeepSeek and Kimi calls.
4
14
32
4,359
Glad to be powering AgenticTrading, an open-source playground for LLM trading agents. Backtest and run on Alpaca paper trading, built with Dr. Xiao-Yang Liu's @Ai4Finance Open Finance Group at Columbia University. Check out more below👇
Excited to introduce AgenticTrading! 🚀 An open-source experimental playground for LLM-powered trading agents. Build, test, and deploy agents that reason, trade, and perform in realistic market environments. We are thrilled to collaborate with Dr. Xiaoyang Liu's group at Columbia University to push the boundaries of AI in finance. Stay tuned for our findings. Explore our platform: agentic-trading-lab.vercel.a… Read our Medium post: medium.com/@kuailedefcl/agen… GitHub Repo: github.com/Open-Finance-Lab/… We welcome more collaborators to join us! Let's shape the future of open finance together. 🌟 #AgenticTrading #AIinFinance #LLM #AlgorithmicTrading #OpenSource
11
7
25
1,170
Commonstack retweeted
A lot of routing work evaluates isolated prompts, but real agent systems are fundamentally multi-step and budget-constrained. Cool to see benchmarks moving toward execution-grounded, end-to-end evaluation instead of just token-level proxies. TwinRouterBench is a strong step toward realistic agentic routing evaluation — especially the separation between static supervision and dynamic SWE-bench execution. Excited to see where this goes!
Excited to share that TwinRouterBench has been accepted to the #RLEval Workshop at #CAIS2026 🎉 As LLM apps become long-horizon agents, one request can trigger many model calls across planning, tool use, retrieval, coding, and verification. That makes per-step LLM routing a core infrastructure problem: sending each call to the cheapest sufficient model without breaking downstream success. TwinRouterBench introduces: ⚡ Static track: 970 router-visible prefixes from 520 instances across SWE-bench, BFCL, mtRAG, QMSum, and PinchBench 🚀 Dynamic track: live SWE-bench Verified evaluation with official task resolution + realized API spend Key result: a router trained on static labels achieves comparable SWE-bench resolve rate while cutting API cost by ~53% vs. an unrouted Opus 4.6 baseline. Paper: arxiv.org/html/2605.18859v1 Code: github.com/CommonstackAI/Twi… Dataset: huggingface.co/datasets/Amor… Website: commonstackai.github.io/Twin… #LLM #AgenticAI #LLMRouting #Benchmark #SWEBench
1
3
10
1,740
Great to see TwinRouterBench accepted to the #RLEval Workshop at #CAIS2026! Per-step routing is quickly becoming essential infrastructure for agentic systems: each planning, coding, retrieval, and verification call should use the cheapest sufficient model without hurting final task success. Proud to open-source TwinRouterBench and contribute a practical benchmark for this problem.
Excited to share that TwinRouterBench has been accepted to the #RLEval Workshop at #CAIS2026 🎉 As LLM apps become long-horizon agents, one request can trigger many model calls across planning, tool use, retrieval, coding, and verification. That makes per-step LLM routing a core infrastructure problem: sending each call to the cheapest sufficient model without breaking downstream success. TwinRouterBench introduces: ⚡ Static track: 970 router-visible prefixes from 520 instances across SWE-bench, BFCL, mtRAG, QMSum, and PinchBench 🚀 Dynamic track: live SWE-bench Verified evaluation with official task resolution + realized API spend Key result: a router trained on static labels achieves comparable SWE-bench resolve rate while cutting API cost by ~53% vs. an unrouted Opus 4.6 baseline. Paper: arxiv.org/html/2605.18859v1 Code: github.com/CommonstackAI/Twi… Dataset: huggingface.co/datasets/Amor… Website: commonstackai.github.io/Twin… #LLM #AgenticAI #LLMRouting #Benchmark #SWEBench
9
13
32
2,259
How do you evaluate an LLM router fairly? Most benchmarks look at prompts, but routers operate at an agentic-step level. A router that saves money but breaks the task could be worse than no router. We open-sourced TwinRouterBench to measure this honestly. 🧵
6
16
44
2,469
Conflict of interest? acknowledged! We know our router (UncommonRoute) currently leads the leaderboard. Open submissions, locked pricing, public scoring code. If a different router wins, the leaderboard will say so.
3
8
270
Step-level routing matters. The benchmark to measure it is open today. Bench: github.com/CommonstackAI/Twi… The current leader: github.com/CommonstackAI/Unc… Paper coming soon on ArXiv.
1
6
215