The Pydantic Stack: Pydantic Validation, Pydantic AI, Pydantic Logfire, Pydantic Evals, and Pydantic AI Gateway

London
Pinned Tweet
One Jev request, three scores: a policy label, a rubric level, and the probability the reply asks for a password, all recorded in Logfire. Modeled cost of 1M evaluated replies: $62, vs $4,542 with Braintrust before data charges. Blogpost pydantic.io/n4PWy
5
6
53
3,178
We just added realtime speech support for Gemini 3.8 Live (v2.50.0) and OpenAI GPT-Live (v2.51.0) to Pydantic AI. Same agent, tools, and limits over one live audio connection. Mic in, speech out, tools on your backend mid-call. When it ends, ordinary message history you can hand to a text agent. Swap models with the model string. pydantic.io/XjyLb
2
25
1,524
Replying to @d_philla @rakyll
Just use Monty from @pydantic
1
1
2
1,465
Pydantic AI agents now hold live voice conversations on Gemini 3.8 Live and OpenAI's GPT-Live. Your mic streams in, speech streams back, and the agent's own tools run mid-call. Moving the same agent between the two is one model string. pydantic.dev/docs/ai/realtim…
8
5
60
2,969
Jev is now on Pydantic AI Gateway. One key with the same spend and guardrails as your other models. Still typed questions + probabilities, not chat. Read the post: pydantic.io/qMS2y
2
2
38
2,106
Imagine what you could do with a sandbox that starts in 1 millisecond...
Fuck it, still early but here goes ... We've just released Monty v1 - a Python sandbox that starts in 1 millisecond, not 1.5 seconds. I just ran 10k sandboxed scripts in 674ms, something that would take a cloud sandbox > 3 hours. This removes the biggest drawback of letting agents write code. The future is fast. Even better, it's open source, you can install it from PyPI, npm or Crates now. Serviced platform coming soon. Please get in touch if you want to be a design partner! Who should try it? ⚑ if you care about startup time, use Monty ⚑ if you care about long-lived sessions, use Monty - Monty can be dumped and resumed at any external function call ⚑ if you care about accessing functions in the agent/host, use Monty - Monty makes it trivial to expose local functions into the sandbox ⚑ if you care about scale, use Monty - Monty workers use as little as 2MB of memory, meaning you can run thousands of concurrent sandboxes on a single machine ⚑ if you care about security, use Monty - we've run 3 rounds of bounty program and thousands of researchers have tried to break into our sandbox, meaning it should be secure to run untrusted code Who should avoid it? 🚫 if you like to take a coffee break while waiting for sandboxes to start, DO NOT use Monty 🚫 if you enjoy the challenge of routing API requests from sandboxes through your corporate network to access state in your agent without exposing secrets to the sandbox, DO NOT use Monty 🚫 if your agent really needs to install packages from PyPI, Monty won't help you yet (spoiler: it probably doesn't) pydantic.dev/docs/monty/get-…
5
3
70
6,345
Whichever new model has you most excited, Pydantic AI has you covered. Claude Opus 5.5 from @AnthropicAI and GPT-6 Sol and Luna from @OpenAI, supported on launch day. And Jev from @typesafeai, not even a week old. Upgrade, then change one string. github.com/pydantic/pydantic…
5
4
54
2,368
Redesigned dashboards are live in Pydantic Logfire: A new chart palette, headline numbers you can compare at a glance, auto-saving drafts, and one-click cloning for built-in dashboards. Read up on what changed, what's next, and the design thinking behind it: pydantic.io/QdevM
1
26
1,697
Pydantic retweeted
Replying to @pydantic
@pydantic has been type-safe from the beginning! Now with Jev powering your Pydantic AI agent calls, they’ll be returning faster than ever ⚑️⚑️
Pydantic AI agents now run on Jev, the classifier from @typesafeai. Jev doesn't write text, it answers typed questions. So the output_type you already wrote is the question, and the answer comes back as your model, one confidence per field. pydantic.io/8iAZ9
11
23
272
59,206
Pydantic retweeted
As a Python enthusiast, I get excited about code that feels truly Pythonic. Great SDKs, natural APIs, and thoughtful design and architecture. That craft means even more in the age of AI-generated code 🐍 Thanks for having me @modal @pydantic @samuelcolvin
1
6
358
big fans of pydantic! we use it a lot at typesafe
Pydantic AI agents now run on Jev, the classifier from @typesafeai. Jev doesn't write text, it answers typed questions. So the output_type you already wrote is the question, and the answer comes back as your model, one confidence per field. pydantic.io/8iAZ9
1
4
20
3,708
Try Monty V1 beta now! > 𝚞𝚟 𝚊𝚍𝚍 πš™πš’πšπšŠπš—πšπš’πšŒ-πš–πš˜πš—πšπš’==𝟷.𝟢.πŸΆπš‹πŸΈ > πš—πš™πš– πš’πš—πšœπšπšŠπš•πš• @πš™πš’πšπšŠπš—πšπš’πšŒ/πš–πš˜πš—πšπš’@𝟷.𝟢.𝟢-πš‹πšŽπšπšŠ.𝟸 > πšŒπšŠπš›πšπš˜ 𝚊𝚍𝚍 πš–πš˜πš—πšπš’@𝟷.𝟢.𝟢-πš‹πšŽπšπšŠ.𝟸
I've published @pydantic Monty V1 beta. We had to significantly refactor the wire protocol based on bounty program feedback, hence the beta release. Masses changed since v0.0.23, please try now! github.com/pydantic/monty/re…
3
4
48
5,609
[1/4] We put @pydantic’s Pydantic AI through a production-style deployment to see whether typed agent primitives could reduce repetitive glue codeβ€”and how its behavior compared with LangGraph across a 160-scenario test matrix.
3
1
5
1,117
Pydantic retweeted
In the process of migrating the repo I built at Braintrust for @HamelHusain and @sh_reya's AI evals course (github.com/braintrustdata/ai…) to an all-@pydantic stack: PydanticAI + Logfire + pydantic-evals. Been living in these tools for a while now and the progress on observability and evals is impressive. I'll be blogging and posting findings here as I go, including where I think models like @typesafeai's Jev fit into the end-to-end evals pipeline. Stay tuned ...
3
11
88
4,915
Jev picks the route now, and fills it in. Give it two output types and a tool, and it works out which one the text calls for, then writes that route's fields or arguments itself. No language model in the chain.
16
10
102
4,839
Your evals judge doesn't have to write text either. LLMJudge and GEval now run on a model that has none, so Jev can score your cases: a rubric becomes a typed question. LLMJudge(rubric='The ticket is urgent', model='typesafe:jev-latest')
1
2
11
756
Pydantic retweeted
Jev by @typesafeai is incredible. But I wanted moar type-safety and something moar pythonic. To that end... Introducing jev.BaseModel, a drop-in replacement for Pydantic BaseModel. Coerce unstructured state to structured outputs with jev in milliseconds.
7
6
101
6,730
Pydantic retweeted
Jev vs. sonnet with @pydantic AI. cost: sonnet $0.0026 vs. jev $0.001 time: sonnet 2.36s vs. jev 643ms Shown in Pydantic Logfire. code (just 17 lines): gist.github.com/samuelcolvin…
12
6
96
6,523
Pydantic retweeted
I created an MCP server so my team can draft emails for me. Just one tool: draft_email(to, subject, body_markdown, cc, bcc) I check my drafts and click send. Running on @render free tier. instrumented with @pydantic Logfire, using GH/samuelcolvin/cloudkv for storage. Free to run. (not open source because it's mostly vibe coded)
7
1
30
2,825
Working on a Pydantic AI integration now. Watch this space.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev β€’ 20-200x faster β€’ 40-400x cheaper (w/ output tokens free) β€’ Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
6
3
96
5,561
Pydantic AI agents now run on Jev, the classifier from @typesafeai. Jev doesn't write text, it answers typed questions. So the output_type you already wrote is the question, and the answer comes back as your model, one confidence per field. pydantic.io/8iAZ9
2
4
572