The context layer making AI 5x more accurate

Context Conference, Oct 28 👉
Atlan retweeted
Added two models to Decision Bench—949 shared text cases: Sage: 92.1% accuracy · 733 ms median Tev1 4B: 85.4% · 387 ms median Results: decisionbench.ai Thanks @levantolabs @bigironchris @marco_derossi @togethercompute @nutlope
Made decisionbench.ai to test because I wanted to see how well Jev actually performs compared to other models. Tested 12 models across 35 tasks. Jev did really well for how little it costs. Wrote up what I found 👇
Article

Putting Jev to the test with Decision Bench

12 AI models tested on 1,071 cases across 35 tasks and 11 domains. With all the hype around Jev and the experiments people were sharing, I wanted to see how useful it actually was. I was interested in

4
8
172
The scene is set folks! Bringing together folks to make sense of Multiplayer AI, IRL in San Francisco on the October 1st :) Excited to announce and celebrate the partners making this happen @AtlanHQ , @hydra_db and @UseCorgi. Bring your questions, curiosities, demos, learn and share notes.
4
4
9
3,786
Read this as an architecture diagram, not a leaderboard!! @rohanatlan built decisionbench.ai/ to test Jev on the small decisions we keep handing to models: 1,071 real cases across 35 tasks. On the same text-only cases, Jev and Gemini 3.5 Flash were basically tied on accuracy (93.2% vs 93.5%)....but Jev was about 8x faster and about 70x cheaper in his setup! Next thing I'd want to know is whether Jev knows when it's wrong, aka how reliable is its confidence...
Made decisionbench.ai to test because I wanted to see how well Jev actually performs compared to other models. Tested 12 models across 35 tasks. Jev did really well for how little it costs. Wrote up what I found 👇
Article

Putting Jev to the test with Decision Bench

12 AI models tested on 1,071 cases across 35 tasks and 11 domains. With all the hype around Jev and the experiments people were sharing, I wanted to see how useful it actually was. I was interested in

1
5
209
How well does Jev hold up on the decisions a company actually makes? Our team tested it on 1,071 real cases across 11 domains, from finance to legal to support. The result: 93.8% accuracy, under half a second per answer, about four cents per 1,000.
Made decisionbench.ai to test because I wanted to see how well Jev actually performs compared to other models. Tested 12 models across 35 tasks. Jev did really well for how little it costs. Wrote up what I found 👇
Article

Putting Jev to the test with Decision Bench

12 AI models tested on 1,071 cases across 35 tasks and 11 domains. With all the hype around Jev and the experiments people were sharing, I wanted to see how useful it actually was. I was interested in

1
4
533
Atlan retweeted
The most underrated value of AI — in software, we had to guess what the user wanted. With conversational AI — they tell us. For a user obsessed company, this has been a gold mine. @rishigb expands how we have turned traces to empathy at scale — and how it is helping us compound user value.
Everyone is talking about evals and judges as possibilities. We put one in production. It scores every MCP outcome for real user value and turns the failures into engineering work. I built it from 1,000 conversations read by hand. Here is the whole thing.
Article

Empathy at Scale: How an LLM Judge Improves Our Atlan MCP Server

TL;DR: An LLM judge scores every Atlan MCP outcome for real user value, and a loop turns the failures into shipped engineering fixes. I built it from 1,000 conversations read by hand. At Atlan, we are

1
1
11
2,386
For most teams, an eval is a report card. Ours is a control loop. 👇
Everyone is talking about evals and judges as possibilities. We put one in production. It scores every MCP outcome for real user value and turns the failures into engineering work. I built it from 1,000 conversations read by hand. Here is the whole thing.
Article

Empathy at Scale: How an LLM Judge Improves Our Atlan MCP Server

TL;DR: An LLM judge scores every Atlan MCP outcome for real user value, and a loop turns the failures into shipped engineering fixes. I built it from 1,000 conversations read by hand. At Atlan, we are

3
408
If any of this is what you're building, or just what you can't stop thinking about, come. Oct 1, SF. A room of people figuring it out, real prototypes, hard questions. Apply to demo or RSVP here: luma.com/fj1r2uu7
1
2
358
Atlan retweeted
We built agents for our Customer Success team 3x times. They all failed. Each one got better at answering questions about our customers. None could act on the answer. The thing holding them back was context. The 4th one stuck because we onboarded it like a human. Before we gave Relay a single task, we treated it like a new hire: the best books in the field, our blog posts, our values, our org chart, the tools the team lives in. Then we built it on the same principle as a self-driving car. Sense each account, plan the next step, act on it. That's when Relay stopped answering questions and started doing the job. A deck that took a CSM two hours now takes 30 seconds. Giving the team their time back for what they do best: helping customers do their life's best work.
5
3
32
4,706
Our Customer Success team built the same AI agent 4 times before it stuck. Each version got smarter. But what mattered most was context. The 4th version, Relay, they onboarded like a new hire, customer success books and all. Now it drafts follow-ups, builds renewal decks, and flags the risk a human missed. The full build story via @dhruvsaharya
2
5
578
Atlan retweeted
I've joined @AtlanHQ as AI Builder Intern for the next 6 months. Hardest part was deciding between @AtlanHQ and @Razorpay (Had similar offers for similar role from both, infact @Razorpay was paying more than @AtlanHQ) Time to give 120% for the best work of my life🥳
8
1
12
825
Atlan retweeted
Looking forward to speaking at the LinkedIn Marketing Connect in Chennai tomorrow & share our learning / experience of building an AI Native Marketing org @atlanhq + how the customer journeys are evolving in the era of AI! Come by if you are around?
3
13
472
Every single enterprise is racing to figure out how to architect their Context Layer for AI Agents. This is your opportunity to dive into this topic FREE with me and @AKronzAnalytics from @AtlanHQ in the next few minutes. You can register here and also catch the recording later maven.com/p/015913/architect… #contextengineering #enterpriseai #aiagents #context
5
7
32
2,044
Multiplayer AI will be one of the biggest design problems of 2027. The phrase alone means something different to every person I talk to.  Is it shared memory? Context? Org design? Something else entirely? On Oct. 1, we’re bringing a group of sharp AI builders together in San Francisco to explore what real multiplayer AI could look like. Real prototypes and demos.
14
12
71
9,040
Registered to my own event... this option to create you're own virtual "badge" after registration was very cool. @AtlanHQ
1
1
59
Turns out “WTF is a context layer?” gets a lot of different answer. 👀 Hear the debate .. and pick your side.. at The Context Conference on Oct 28.
Five months. Countless experts. One deceptively simple question: WTF is a context layer? We found semantic layers, ontologies, knowledge graphs—and plenty of disagreement. So naturally, we made a documentary about it.* October 28. Context Conference. The Final Showdown. *Netflix hasn’t called back.
1
2
122
Atlan retweeted
We’ve been cooking with AI in the Context Layer side a lot!! Excited for this 💙
Five months. Countless experts. One deceptively simple question: WTF is a context layer? We found semantic layers, ontologies, knowledge graphs—and plenty of disagreement. So naturally, we made a documentary about it.* October 28. Context Conference. The Final Showdown. *Netflix hasn’t called back.
2
3
14
1,339
Atlan retweeted
We've been working on the context layer the last few months. Join Context Conference to see what we've been cooking! đź’Ą
Five months. Countless experts. One deceptively simple question: WTF is a context layer? We found semantic layers, ontologies, knowledge graphs—and plenty of disagreement. So naturally, we made a documentary about it.* October 28. Context Conference. The Final Showdown. *Netflix hasn’t called back.
1
1
6
173
Five months. Countless experts. One deceptively simple question: WTF is a context layer? We found semantic layers, ontologies, knowledge graphs—and plenty of disagreement. So naturally, we made a documentary about it.* October 28. Context Conference. The Final Showdown. *Netflix hasn’t called back.
2
1
10
2,806