Exploring agent swarms @ tardigrade.sh, backed by @spc; Prev. co-founded jenni.ai, smithery.ai (acquired)

Pinned Tweet
Smithery has been acquired by Arcade! It’s been an incredible year building Smithery alongside the MCP community. I'm proud that we've built something that helped process millions of tool calls and made 10K+ MCPs from developers around the world more discoverable. My co-founder @kamathematic is joining Arcade to keep building toward our vision of making tool calls the new clicks. As for me, I’ll be taking some time to explore what’s next. More soon!
@SmitheryDotAI, the leading public registry and hosting platform for MCP, is now part of Arcade.dev. Smithery made finding and running an MCP server take one click. Pairing that developer experience with the security, reliability, and governance that Arcade is known for, gives enterprises both to lead the next generation of agentic AI. Smithery co-founder @kamathematic is joining us to keep building it. More from @TheMostlyGreat: arcade.dev/blog/smithery-joi…
26
4
85
9,928
So Anthropic models are good when it's a minor version bump rather than a major version bump!
1
227
Henry Mao retweeted
You can now property-test your Tardigrade agents by generating and checking possible trajectories against safety rules! Since each agent is a state machine, we use fast-check to find states that break defined invariants. This is particularly useful in critical settings where you can't simply rely on the vibes of your prompts. This takes us one step closer to formally verifying agent harnesses. Read the complete docs here: tardigrade.sh/docs/verifying…
4
4
32
843
Henry Mao retweeted
I built Vybez A small TypeScript sugar that weaves Jev decisions ergonomically into your existing code. Most workflows mix semantic decisions and deterministic rules; Vybez makes that strongly typed and seamless in TypeScript. github.com/clavia-labs/vybez Try: bun add vybez
5
4
12
567
Henry Mao retweeted
We’ve been using tons of TLA+ for Tardigrade, very bullish on it. Engineers should get into math more than ever because the best specification for a system is math. Not some text in a markdown file. Specifying systems by Leslie Lamport (author of TLA) is a great book for those getting started. lamport.azurewebsites.net/tl…
I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs and race conditions. Video attached. TLA+ also works well. I sometimes combine Lean and TLA+ to look for issues around data flow, concurrency, and state mgmt. I don't know either language well, but Claude is excellent at both. This approach is super useful for formally modeling your code and finding bugs that a human probably wouldn't have spotted. Is formal verification the future of coding (or at least, bug finding)?
14
18
188
25,294
Multi-agent simulations are going to be prevalent for RL envs!
Tardietown! 👋 A fun experiment where you create a town of agents coordinating using a reddit-like forum with karma for work. Still very rough around the edges. Working on it mainly as a demo for tardigrade. Everything in the town is an actor hosted as a durable object; the residents, library, packages etc. Open sourced. Run it locally, deploy it, clone it and create your own towns and simulations! github.com/clavia-labs/tardi…
3
358
Henry Mao retweeted
So far, I'm impressed with Jev as a verifier and reward model. It makes RL for hard-to-verify domains so much cheaper to iterate! @typesafeai Allows you to easily implement CheckEval-style judges: arxiv.org/abs/2403.18771
5
25
213
14,900
I built Vybez A small TypeScript sugar that weaves Jev decisions ergonomically into your existing code. Most workflows mix semantic decisions and deterministic rules; Vybez makes that strongly typed and seamless in TypeScript. github.com/clavia-labs/vybez Try: bun add vybez
5
4
12
567
Inspired by:
I've seen people describe Jev as an "AI if statement". But what if it actually WAS an if statement? Introducing Probably: a programming language powered by Jev: probably-lang.southpolesteve… Jev baked into the language. “feels” asks a question. “match” routes between descriptions. “while” keeps going until something stops feeling true. This is obviously a toy, but it's fun to think about what something like Jev unlocks. Jev makes the decisions, an LLM does the writing, and a little program ties it together.
1
175
One trick to get "reasoning" out of Jev @typesafeai is to loop it: feed Jev's output score back into its input state, and iterate a few times to refine its confidence. It's the Universal Transformers (arxiv.org/abs/1807.03819) idea from 2018.
The dial I want is not more compute on every call, it is the option to spend more only on the unsure ones. I already branch on confidence in my game. Let me hand the 0.5s back for a second, harder look and leave the 0.9s alone.
4
1
53
8,169
Calling Jev @typesafeai system one limits our imagination of what classifiers can do. I can easily see these models have a reasoning dial to trade more compute for performance. In that sense it becomes more than just system one.
6
1
36
16,438
We're definitely still at the phase where intelligence is being metered
1
251
Maybe the future looks like Interaction Models in the front end with a bunch of Jevs in the backend. The majority of agent invocations are structured tool calls anyway, so it makes sense to specialize specifically for that.
2
302
Not sure what architecture this is, but reminds me of "back in the day" when I finetuned masked language models like BERT! When output shape is known in advance, MLM approaches can be pretty efficient!
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
1
7
895
Henry Mao retweeted
Docs for Tardigrade components is up! In Tardigrade, a component is the smallest unit of an agent's behavior. This page interactively walks you through the concepts behind them: state machines, effects, and event sourcing. tardigrade.sh/docs/component…
5
6
97
4,044
There's no magic in "pure synthetic data" for post-training, just as there are no perpetual motion machines. The alpha either comes from ground truth that you source or from handcrafted recipes (you are the alpha). Ideally, you want both to increase the leverage of the source.
Making RL environments for frontier models is kinda like On Policy Self Distillation. You're injecting some privileged information to create tasks that models wouldn't be able to conceive of otherwise. The difference is the info is encoded in the env, rather than the prompt.
2
2
586
Making RL environments for frontier models is kinda like On Policy Self Distillation. You're injecting some privileged information to create tasks that models wouldn't be able to conceive of otherwise. The difference is the info is encoded in the env, rather than the prompt.
5
902
So, how many ChatGPT subscriptions do you have now?
83% 1
0% 2
17% 3
0% 4+
6 votes • Final results
242
Compared to the Harvey LAB dataset @harvey, which is synthetically generated and larger in scale, this leaderboard is evaluated on real data crafted with lawyers in the loop.
Can you really trust an AI's legal advice? I partnered with Legal Benchmarks @aguozy to launch their new leaderboard measuring frontier legal capabilities. Here's how we built the task environment and calibrated an LLM judge against lawyer preferences. 👇
4
461
I was slightly surprised that Claude Opus 5/Fable 5.1 models outperformed Astra on this benchmark Astra won in legal writing style (i.e., form), but Opus/Fable won on legal correctness (i.e., substance). Muse Spark also did really well
Can you really trust an AI's legal advice? I partnered with Legal Benchmarks @aguozy to launch their new leaderboard measuring frontier legal capabilities. Here's how we built the task environment and calibrated an LLM judge against lawyer preferences. 👇
1
362