Inference infrastructure for AI-native teams. Join us on Discord discord.gg/K2deYSXNu

San Francisco, CA
Catalyst is live built for teams shipping agents in production train & deploy frontier LLMs in minutes, using the data your application is already generating Get Started: docs.inference.net/introduct…
Introducing Catalyst: a developer platform to monitor, train & deploy self-improving AI models built for teams operating AI products at scale Catalyst can automatically: - collect traces from your agents - curate training data & evals - train & deploy models on par w/ Opus 4.6
3
1
28
35,102
October 1st 🚀 we’re taking over a sick theater for a private screening of Interstellar 🚌 Coach bus leaves the Inference office @ Embarcadero at 5pm 🍻 Come pregame with us beforehand 🍴 Catered food + beers + Inference merch at the theater 🚌 Back at the Inference office by 9:30pm Come hang , spots are limited make sure to RSVP :) luma.com/go1i4gpy
2
1
1
1,604
Inference retweeted
excited to share that @inference_net is officially live on the @vercel ai gateway! vercel.com/ai-gateway/models…
3
5
36
2,143
Inference retweeted
My dear friends, at our current growth rate @inference_net will cross $10M ARR at precisely 6pm on Oct 1, 2026 to celebrate, we've rented a 140 seat movie theatre for a private screening of Interstellar join us for a evening of vibes and fun luma.com/go1i4gpy
12
5
115
17,244
Inference retweeted
introducing fast inference (fast.inference.net) fast inference is an LLM API for devs who want to go faster access the fastest Kimi K3, GLM, DeepSeek, etc for $9/mo integrations with Claude Code, Codex, OpenCode & more speeds compounds. waiting for Claude does not
47
17
280
159,441
Inference retweeted
I’m running an experiment with a handful of folks who spend > $5k/mo on coding models to better understand the state of open source models for complex coding tasks. If you’re interested in driving Kimi K3 for a few days and giving feedback, shoot me a DM Paying $500 each
9
2
38
3,457
Inference retweeted
It’s crazy how much a 3x speed difference makes for coding agents. here's a side-by-side of both models re-creating the github home page. glm is halfway done before sonnet even starts. glm 5.2 fast (@inference_net) = ~300 tps sonnet 4.5 (via anthropic) = ~100tps even crazier: glm 5.2 doesn’t have image support and still works for this
1
4
14
1,427
Inference retweeted
We're releasing Inference AutoEvals⚡️ automatically find the best model for any agent by replaying traffic across a suite of models. works with every model/provider on average, AutoEvals cuts spend by 30% & improves accuracy by ~10% with a one-line code change Available now👇
9
13
109
19,341
Inference retweeted
Want to try Kimi K3 in production but worried how it might change your product? Don’t worry, we got you: 1. Install Inference Gateway (docs.inference.net) 2. Keep sending traffic to your current provider 3. Gateway automatically starts sorting through your live data using an RLM to generate evals for your app. This takes ~24 hours. 4. Gateway starts mirroring live traffic to Kimi K3 to run evals. Traffic is mirrored - you’re still using your old model in prod. 5. Once evals look healthy, you get a Slack notification letting you know it’s safe to switch. 6. Switch model id in your code to “kimi-k3” Congrats, you just saved 50% on your monthly token bill, and you own your LLM stack end to end.
16
20
297
96,194
Inference retweeted
We’re releasing Inference AutoRouter. Always the right model for your task AutoRouter evals dozens of models and routes requests based on cost/accuracy One endpoint. Hundreds of models Append `:auto` to the model name to enable AutoRouter Available in private beta today 👇
12
7
116
13,017
Inference retweeted
5 things @samhogan does differently > signs company documents from the terminal > trained 5 function-specific @inference_net agents > keeps his entire to-do list in one apple note (over a year old) > moved his entire team to @opencode for llm portability > blocks twitter on his phone from 9-5 (300+ muted words) ep 6 of show me your stack is live!
7
4
53
43,928
Inference retweeted
We're releasing Inference AutoTune Distill any frontier model into a 1-30B parameter task-specific SLM with only 25 lines of code automatically route requests to reduce cost and latency by >90% ~2 hours and <$250 to train. You own the weights Available in private beta today
142
208
2,646
354,667
Inference retweeted
we're helping a customer spending $60k/mo move from OpenAI & Anthropic to open source models they use almost every model offered by the labs, so we needed to find replacements for all of them after generating evals, this is what we landed on new cost: $12k/mo, 80% savings
152
97
1,560
225,065
Inference retweeted
I'll be on @MTSlive today at 4 pm discussing GLM 5.2 adoption, what this means for frontier labs, and The Shape of Inference in a post-GLM 5.2 world
META BRAIN MODEL | NEW AGENT BILL | FABLE BACK SOON? nitter.net/i/broadcasts/1nxeLLOWp…
4
21
6,512
Inference retweeted
Execs at Google are probably calling the US government right now and begging them to withhold the next Gemini release
74
73
2,086
121,970
Inference retweeted
We’re issuing up to $5k in @inference_net credits to teams running large agentic workloads on Opus / Codex 5.5 and want to try GLM 5.2. DM me @atbeme for info. - reduce spend by ~80-90% - faster and more predictable latency - set yourself up for fine tuning down the line
13
6
180
43,374
Inference retweeted
Want to try GLM 5.2 in production but worried how it might change your product? Don’t worry, we got you: 1. Install Inference Gateway (docs.inference.net) 2. Keep sending traffic to your current provider 3. Gateway automatically starts sorting through your live data using an RLM to generate evals for your app. This takes ~24 hours. 4. Gateway starts mirroring live traffic to GLM 5.2 to run evals. Traffic is only mirrored - you’re still using your old provider in prod. 5. Once evals look healthy, you get a Slack notification letting you know it’s safe to switch. 6. Switch model identifier in your code to “glm-5.2” Congrats, you just saved 90% on your monthly token bill, and you own your LLM stack end to end.
78
68
1,471
537,155
Inference retweeted
Looking like we'll 4x revenue in the next 90 days, thnx GLM 5.2
9
6
247
26,442
Inference retweeted
I love how the jupyter notebook crowd successfully rebranded basic software engineering and rote data annotation as “Bleeding-edge ML Research” lmao absolute legends
4
1
35
3,565
Inference retweeted
We just launched agent signals! Now you can get notified when a user crashes out on your agent. Write any prompt and it will run on your trace data as it comes in, classifies it, and then optionally notifies you. Track your classifications over time to see trends in agent/user behavior. Check out this guide to get it set up: docs.inference.net/guides/me…
1
2
834
Inference retweeted
We’re releasing HALO Desktop 😇 it's the best way to find bugs in your agents import traces from Langfuse or Arize & have HALO create a report with failure modes give the report to your favorite coding agent to build it runs locally on your machine. 100% free and open source
9
6
79
18,431