Head Of Research @AtlanHQ, Author, Forbes Asia 30 under 30. Truth seeking, tweets+typos are my own.

San Francisco
Multiplayer AI will be one of the biggest design problems of 2027. The phrase alone means something different to every person I talk to.  Is it shared memory? Context? Org design? Something else entirely? On Oct. 1, we’re bringing a group of sharp AI builders together in San Francisco to explore what real multiplayer AI could look like. Real prototypes and demos.
14
12
71
8,985
Spotted Mozzeria truck in Dogpatch today. The best marinara pizza by far! Or maybe it was the best combination of the chilli oil along with the chilly breeze that did it. #sf
3
54
Rishi Gaurav Bhatnagar retweeted
We're launching the Muse / Instinct for enterprise, backed by a16z Speedrun. We believe the next big shift in work isn't the model or the agent. It's the interface. General-purpose AI assistants are supercharging how consumers get things done. Employees still spend 60% of their day doing work about work, a third of their time finding things, and switch between apps 1,200 times a day. Contxt works where your team does. Text it in Slack or Teams, call it in Claude or ChatGPT, or forward it your emails. It has a dynamic memory across your tools, history, and people, and can do work directly in your browser, just like your team does. You get the 100x productivity of a consumer AI assistant with the controls, audit trail, and shared knowledge a company needs. We're a Princeton-educated team of ex-NASA engineers, repeat founders, and quants.
70
19
578
135,268
The scene is set folks! Bringing together folks to make sense of Multiplayer AI, IRL in San Francisco on the October 1st :) Excited to announce and celebrate the partners making this happen @AtlanHQ , @hydra_db and @UseCorgi. Bring your questions, curiosities, demos, learn and share notes.
2
4
7
2,942
Rishi Gaurav Bhatnagar retweeted
Jev is truly underrated (still). We have been testing its limits and reimagining how we can use AI. @rohanatlan just open sourced our internal bench. Hope it helps!
Made decisionbench.ai to test because I wanted to see how well Jev actually performs compared to other models. Tested 12 models across 35 tasks. Jev did really well for how little it costs. Wrote up what I found 👇
Article

Putting Jev to the test with Decision Bench

12 AI models tested on 1,071 cases across 35 tasks and 11 domains. With all the hype around Jev and the experiments people were sharing, I wanted to see how useful it actually was. I was interested in

2
2
11
2,022
Jev tested against a real world bench - fascinating to see all the use cases where is a performing well and at a cheap price. But also surprised to see all the cases where it is struggling!
Made decisionbench.ai to test because I wanted to see how well Jev actually performs compared to other models. Tested 12 models across 35 tasks. Jev did really well for how little it costs. Wrote up what I found 👇
Article

Putting Jev to the test with Decision Bench

12 AI models tested on 1,071 cases across 35 tasks and 11 domains. With all the hype around Jev and the experiments people were sharing, I wanted to see how useful it actually was. I was interested in

1
1
102
Everyone is talking about evals and judges as possibilities. We put one in production. It scores every MCP outcome for real user value and turns the failures into engineering work. I built it from 1,000 conversations read by hand. Here is the whole thing.
Article

Empathy at Scale: How an LLM Judge Improves Our Atlan MCP Server

TL;DR: An LLM judge scores every Atlan MCP outcome for real user value, and a loop turns the failures into shipped engineering fixes. I built it from 1,000 conversations read by hand. At Atlan, we are

45
2
37
4,321
Rishi Gaurav Bhatnagar retweeted
seeking those obsessed with evals if you: - are half researcher, half engineer - love measuring hard to measure things - are creative - have a particular interest in post-training ...you'll like it here. sf based, in person, obsession coming from a place of curiosity. apply 👇
19
25
310
19,119
I have been a big fan of @HamelHusain and @sh_reya 's public lectures, and blogs. Learnt so much their writings. Now, this list is a beaut.
This is the most valuable free resource we've created on AI Evals 🎉 (not exaggerating!) I organized our public materials into this guide. It allows you to find answers to your eval problems w/o searching aimlessly. Humans: pick the row that sounds like you. Agents: point it at the post The material draw on 60+ hours of office hours from our Evals course, where @sh_reya and I have taught 5k engineers and PMs Evals. We add new material often, with 15 FAQs added in the last two weeks. Recent additions are marked with a "New" badge. You can find all of these and more here: hamel.dev/blog/posts/evals-f…
77
Rishi Gaurav Bhatnagar retweeted
This is the most valuable free resource we've created on AI Evals 🎉 (not exaggerating!) I organized our public materials into this guide. It allows you to find answers to your eval problems w/o searching aimlessly. Humans: pick the row that sounds like you. Agents: point it at the post The material draw on 60+ hours of office hours from our Evals course, where @sh_reya and I have taught 5k engineers and PMs Evals. We add new material often, with 15 FAQs added in the last two weeks. Recent additions are marked with a "New" badge. You can find all of these and more here: hamel.dev/blog/posts/evals-f…
55
107
927
61,756
Rishi Gaurav Bhatnagar retweeted
Multiplayer AI is the next frontier. But what is it? Shared memory? Context? Org design? Something else? We're bringing together AI builders in San Francisco to explore what real multiplayer AI looks like w/ real prototypes & demos Register: luma.com/fj1r2uu7?tk=2msWbb
3
1
9
4,965
Rishi Gaurav Bhatnagar retweeted
Today, we're launching MCPJam as the first testing & evaluation platform for MCP servers. Most MCP servers FAIL to deliver value to their users. Over 106,000 developers and +300 enterprises use @mcpjams open-source solution to see how their servers behave locally across major AI clients. Now, MCPJam helps teams answer a hard question: what does "good" look like for MCP?
15
7
35
107,828
Suren is one of the strongest leaders - leading probably the most AI native org I know of and building the most agile group of agents and a marketing OS. Listen in to the conversation if you can :)
Looking forward to speaking at the LinkedIn Marketing Connect in Chennai tomorrow & share our learning / experience of building an AI Native Marketing org @atlanhq + how the customer journeys are evolving in the era of AI! Come by if you are around?
1
215
Multiplayer AI will be one of the biggest design problems of 2027. The phrase alone means something different to every person I talk to.  Is it shared memory? Context? Org design? Something else entirely? On Oct. 1, we’re bringing a group of sharp AI builders together in San Francisco to explore what real multiplayer AI could look like. Real prototypes and demos.
14
12
71
8,985
The lens everyone skips: change. I keep landing on the same uncomfortable thought, the hardest part of multiplayer AI might not be technical at all. Getting a whole team to adopt one shared way of working is more org design than engineering. Nobody wants that to be true.
1
1
326
If any of this is what you're building, or just what you can't stop thinking about, come. Oct 1, SF. A room of people figuring it out, real prototypes, hard questions. Apply to demo or RSVP here: luma.com/fj1r2uu7
1
2
357