Inference is everything.

San Francisco and New York
Baseten retweeted
The open frontier only works if we can make autonomy safe by default. @JensenHuang has done important work here that we're proud to support. Excited to join the Open Secure AI Alliance and to partner with @nvidia on the launch of the Agent Safety Platform.
Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry. Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come. But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility. This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems. Together, we are building the foundation of the AI economy. Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world. nvda.ws/4hcoq7m
1
2
7
621
We're proud to be a launch partner for the NVIDIA Open Agent Safety Platform and a member of the Open Secure AI Alliance. Baseten is committed to contributing to the open frontier, including open safety standards and frameworks for inference and training. Two weeks ago, we launched our safety infrastructure research effort with Base Labs. Today, we've contributed to OpenShell and built a template to run it inside our sandboxes. Much more to come.
Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry. Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come. But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility. This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems. Together, we are building the foundation of the AI economy. Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world. nvda.ws/4hcoq7m
3
6
24
1,300
Every day, more companies are building with the open frontier. We're thrilled to support them. Congrats @genspark_ai on the launch of GenCode! Power your coding agents with closed or open-weight models, with open models at 90-95% lower cost.
We have been waiting to say this. 🚀 GenCode is live. A coding agent inside the Genspark Super App: Claude, GPT, DeepSeek, or open weights, all in one account. Zero API keys to fiddle with. 💡 Every coding agent picks your model for you. GenCode does not. That choice is yours. 🎯 Frontier and open weights, pick per task. Open weights at 1/10 to 1/20 the cost. Right model for the task. Not the one you were handed. genspark.ai/gencode
1
6
2,028
We're officially in our Sheila era. Excited to welcome our new CMO to Baseten! baseten.co/blog/sheila-vashe…
10
8
101
33,369
LangSmith Fine-Tuning from @LangChain is in public beta today, and its open-source CLI, smithtune, trains on Baseten Loops. It turns your LangSmith traces into a fine-tuning dataset and can deploy the finished checkpoint to Baseten with one command. Congrats on the launch!
Introducing LangSmith Fine-Tuning and the smithtune CLI. LangSmith now handles the entire fine-tuning process. Use your traces to train specialized models that cut cost and latency. Now in Public Beta. langchain.com/blog/langsmith…
1
20
4,705
NVIDIA Nemotron 3 Diarization is available in the Baseten Model Library, day 0! - 500+ concurrent hour-long diarization streams supported on a single RTX PRO 6000 - Configurable algorithmic latency from 0.32s to 30.4s - Labels for up to 8 speakers, no clustering required How we optimized @NVIDIAAI's model for real-time streaming, VAD, and transcription: baseten.co/blog/nvidia-nemot…
1
1
23
1,653
RL teaches models to work longer, but reasoning is dependent on domain-specific post-training. Baseten's Head of Model Training @oneill_c sat down with @dwarkesh_sp to explain horizon generalization and what's next at the frontier. Full episode here: piped.video/watch?v=PrSf7IOY…
3
25
2,051
"Your user data, your signal from these models, the improvements to these models themselves will become the core IP of every company in the world." @tuhinone sat down with @alexeheath to talk about the future of inference, Base Labs, bringing Blaxel on board, and why every company will want to own its intelligence.
Baseten’s CEO on why every company will want to own its AI / @tuhinone on the rise of AI agents, the demand for inference, and building a new kind of hyperscaler. Lately, I’ve been spending a lot of time thinking about the AI inference market and how big it could get. Agents like Muse, Instinct, Town, and Grok Bot aggressively use browsers, run their own computers, and burn through far more tokens than a traditional chatbot. As more people put them to work (Muse is number two in the App Store), the demand for inference, or the computing needed to run these models, should grow enormously. Tuhin and I discuss the rise of agents using browsers and virtual machines to get things done, and @baseten's recent acquisition to help power that shift. We also talk about the data center backlash, why companies are embracing Chinese open models, his plans for Baseten’s new research lab, and why he thinks inference becomes the only market left after AGI. Timestamps: 00:00 What Is AI Inference? 06:05 Competing With the Cloud Giants 09:24 Why Companies Want to Own Their AI 15:30 Building Baseten Before the AI Boom 23:45 DeepSeek and the Race for Open AI Models 29:32 Baseten’s Growth and Expansion 33:05 AI Agents and the Blaxel Acquisition 38:32 Data Centers and the AI Backlash 43:25 What Happens to Inference After AGI? 45:12 When AI Agents Become Customers Thanks to the show's premier sponsors: @Atlassian, @meetgranola, and @mercury.
2
25
6,451
We sat down with the team at @elise_ai to talk about why they're betting on specialized, post-trained models, how to evaluate them, and what it takes to run them in production. We're doing round two at SF Tech Week on October 7th. @oneill_c, our Co-Head of Training, is doing a fireside chat with Mario Martone, who leads applied research at EliseAI, on fine-tuning and continuous retraining for housing and healthcare. Join us for drinks and Mario Kart afterward! RSVP here: partiful.com/e/qZ9tQNSPOgpvx…
11
4
25
2,012
We're proud to partner with the team at @p0 to power fast, accurate, low-cost web search for open-weight models via Baseten Grounded Inference.
Parallel Search is now available in @baseten. Pick Parallel for the most accurate web search at the lowest price. Your agents will thank you for the upgrade.
4
4
43
8,484
We're excited to work with @youdotcom to offer web search with open-weight models at a fraction of the price of closed frontier labs.
We're partnering with @baseten to make the case for open-weight models. Paired with @youdotcom web search, they're a real alternative to closed models at a fraction of the price. Already calling Baseten's Chat Completions endpoint? Add one line to your tools array: tools=[{"type": "baseten__you__search"}] Same model, same request. No search vendor to wire up, no retry path, no resending a growing context every turn.
1
18
2,495
Excited to partner with @KeenableAI to power open-model web search via Baseten Grounded Inference!
Keenable Search is now native to @baseten Models search, read, and reason inside one inference request.
1
14
1,874
Excited to partner with the @ExaAILabs and @ExaDevelopers teams to help us power Baseten Grounded Inference, our new server-side web search tool for open-weight models.
Use Exa with Baseten! Add Exa's web search to any open-source model using Baseten Hosted Tools. How to install 👇
3
1
29
4,002
Many of our customers run workloads with speaker attribution, with the strictest demands on both quality and speed. We partnered with the team at @pyannoteAI to deliver both: some of the highest-quality diarization models on the market, with 3.2x higher throughput and 9.6x lower latency.
3
3
18
1,735
Excited to see Vercel push the open-weights ecosystem forward. 💚
v0 is now model-agnostic. Frontier models, cheap models, open models, and fast models, all from Vercel's AI Gateway. Choose from Claude, GPT, Kimi, GLM, Grok, DeepSeek and more to build your apps.
2
20
2,757
Baseten retweeted
Open models on Baseten can now search the web, within the inference path, thanks to our partnership with Exa, Keenable, Parallel, and You.com. A lot more to come here soon.
Until now, adding web search to open-source models meant hand-wiring orchestration, managing separate keys, and paying latency taxes on every round trip. Today, we're solving that. We're excited to introduce Baseten Hosted Tools and Baseten Grounded Inference to bring real-time web search server-side to open models running on Baseten through a single configuration: - 15% lower latency compared to client-side execution - No extra vendor key required - Zero orchestration We're launching this preview version with four leading web search partners, @ExaAILabs, @KeenableAI, @p0, and @youdotcom. Get frontier-level web search parity for your open-weight models. More here: baseten.co/blog/introducing-…
4
3
47
3,495
Web search is now a first-class citizen for open models, only on Baseten. Baseten Hosted Tools and Baseten Grounded Inference bring real-time, server-side web search to your favorite open models via multiple providers: @youdotcom, @p0, @KeenableAI, and @ExaAILabs.
1
2
22
2,669