Accelerate inference, model shaping, and pre-training on a research-optimized platform.

San Francisco, CA
TogetherLink, coming soon.
5
6
55
16,877
Together is #1 on @OpenRouter token share across top open coding models, today: 🥇 @Zai_org GLM 5.3 Flash: 29.2% 🥇 @deepseek_ai V4.1 Flash: 25.6% 🥇 @Kimi_Moonshot Kimi K3: 18.9% Running coding agents on open models? Look no further. openrouter.ai/provider/toget…
10
10
70
31,835
Qwen3.8-Flash is 40% off through the rest of the month, making now a perfect time to run your evals. Designed for high-volume applications like coding and coworking assistants, @Alibaba_Qwen's model optimizes quality at a low cost. Try it today: api.together.ai/models/Qwen/…
1
1
9
2,134
We are proud to be launch partners for NVIDIA Open Agent Safety Platform. We believe safe and responsible AI development is critical as the use of agents scale and they get more prevalent. Together AI is committed to safe AI and we have built platform capabilities for secure development and deployment of agents. We will continue investing in this space, including our work with NVIDIA on OpenShell.
Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry. Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come. But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility. This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems. Together, we are building the foundation of the AI economy. Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world. nvda.ws/4hcoq7m
5
3
22
3,301
Together AI retweeted
Our serverless APIs for open weights models have become exquisite machinery (that we plan to write more about soon!) One effect is that @togethercompute now one of the largest originators of open tokens in the US, with a wild week over week growth curve that reflects the pace of OSS adoption. Will serve 1.8T tokens today on @OpenRouter alone. You can get started with $5 — api.together.ai openrouter.ai/provider/toget…
11
14
77
8,525
Reliable, fast inference is essential. As Ted shared at Apsara Conference 2026, Together makes it easy for teams to get access to frontier open models without needing to manage any of the infrastructure complexity. See for yourself api.together.ai
Making AI simpler for businesses. At #ApsaraConference2026, Ted Cui, VP of Engineering and Inference Platform at Together AI, spoke about AI becoming more accessible, with model selection becoming less of a concern for businesses. Find out more: click.alibabacloud.com/m/200… #AgenticEra #AgentNative #AIAgentsAtApsara #BringYourAgent #AlibabaCloud
4
3
12
4,497
Together AI retweeted
Just trained Tev1 0.8B, a tiny Jev-like classifier. Here it is running completely locally on my mac with @ollama & classifying some tasks. It's extremely fast: only ~50ms E2E latency. Video is not sped up! Releasing weights & benchmarks very soon so you can try it yourself :)
Announcing tev1-4B-experimental, a Jev-like classifier finetuned on top of Qwen3.5 4B for only $17. I'm releasing everything: the weights, data recipe, & a full tutorial on how to train your own. You can try it today on Together serverless at $0.042/1M input & $0/M output.
35
39
425
46,881
"We discovered early on that it's really hard to reliably deploy AI in the real world. Frontier models are really strong, but they're really slow for real-time use cases." See how @DecagonAI partners with Together to power its voice AI concierge.
4
13
3,858
We're releasing tev1-4B-experimental, a Jev-like classifier finetuned on top of Qwen3.5 4B. We're making it available on Together serverless at $0.042/M input & $0/M output. Also releasing the data recipe & a tutorial on how to finetune your own (Tev1 cost $17 to train!).
43
82
1,285
135,818
Note that this is an experimental model meant to show how you can train your own specialized decision models fairly easily! You can also check out the model weights here, along with the full tutorial on finetuning your own classifiers above: huggingface.co/togethercompu…
2
33
4,283
Qwen3.7-Max and Qwen3.8-Flash from @Alibaba_Qwen are now 40% off on serverless through September 30. Use Max for long-horizon agent work and Flash for high-volume, cost-sensitive workloads. Get started today! api.together.ai
5
12
2,894
New on Dedicated Model Inference: canary rollouts. Upgrade the model behind a live endpoint without downtime. Traffic moves from your current deployment to the new checkpoint in gated steps (default 5% → 25% → 50% → 100%). Health checks run before any traffic shifts. After every step, metric gates compare the new model's p95 latency and error rate against the old one. If a gate trips, the rollout pauses at the canary share and waits for you: resume, promote to 100%, or roll back. Three strategies: canary, blue-green, and rolling. Available now via the tg CLI, REST API, and Python SDK. Learn how to start a rollout: together.ai/blog/canary-roll…
6
6
16
7,402
Together AI retweeted
Introducing theopenfrontier.com! Figure out which open model is best for your use case. Compare models across coding, agents, long context, vision, finance, and more. Then see how they compare on cost + quality, including what you could save by moving to open models.
41
16
250
18,296
coding agents are moving fast from prototype to production. the infrastructure question is what's left. join @parthsareen from @ollama and @zainhas Hasan from Together AI at @AIconference for a breakout on what it actually takes to build coding agents on open models, and run them at scale.
4
5
25
7,054
This week in London, we got into the real economics of running AI in production. Cost, model selection, routing, reliability. Straight talk from teams doing this at scale. Thanks to @MiniMax_AI, @deel, and everyone who joined us. Until next time.
7
2
13
4,708
deepseek v4.1 flash on together ai is leading on ttft for pre-warmed queries for agent workflows, that startup time shows up again on every model turn, so shaving it down matters more as the loop gets longer we’ve been pushing hard on the serving path to keep that overhead as low as possible
omg @togethercompute I love you. this is time to first token for every provider for Deepseek v4.1 flash. you guys CRUSH on pre-warmed queries. there's no other game in town
8
4
29
6,497