Prime Agent | Continual Harness, LLM Economist | Research @PrimeIntellect | PhD @Princeton | NSF GRFP

🐯
My YC Paper Club talk on Prime Agent is out. I talked about moving beyond naive prompting toward an agentic OS, and how eval-driven harness design can expose more of a model’s underlying capabilities through persistent computation, memory, and agent-to-agent communication.
Harnesses often get dismissed as just scaffolding, just prompt engineering, and not real research. But that couldn't be farther from the truth. The same model weights that score 30% on ARC-AGI score 95% with a better harness. So we gathered a group of researchers and founders working at the frontier to do a deep dive into the state of harnesses. We cover how we got to this point, the case for making your harness as expressive as possible, and what YC learned building an agent for every employee in the company. 00:00 - @FrancoisChauba1: Why harnesses matter 04:27 - Building an auto-researcher by accident 07:13 - A five minute history of harnesses 13:56 - Self-improving harnesses 18:35 - @sethkarten: Prime Agent, a self-improving RLM harness 21:50 - Context as an L1, L2, L3 cache 24:51 - From Turing machine to von Neumann computer 28:33 - Messaging between agents 30:04 - ARC-AGI results 33:09 - Emulator Bench and GPU kernels 37:30 - @JonSaadFalcon: OpenJarvis, personal AI on personal devices 38:26 - How far behind are local models 39:21 - The five primitives of a personal AI stack 42:47 - Letting cloud models optimize your local stack 43:53 - 800x cheaper than the cloud 45:58 - @josh__france and @jbellregan: QM, YC's agent harness for work 47:29 - A history of YC's internal agents 49:24 - OpenClaw and a fleet of 50 agents 51:04 - Pulling the brain out of the sandbox 54:43 - Letting the agent choose its own sandbox and model 57:16 - The grind tool: budgets on goals 58:50 - Agents don't understand social context
18
26
260
35,351
There is no antimemetics division
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks • A new research finding, demonstrating that one can construct self-replicating prompt injections alignment.openai.com/misalig…
11
Seth Karten retweeted
LATTE is accepted to NeurIPS 2026! ☕️ Adding agents is easy. Coordinating is hard. LLM teams often duplicate work, overwrite each other, and stall: classic distributed-systems failures. We borrowed the fix: a shared task graph that agents build, claim, and revise as they go 👇
🚨 What if your LLM team could dynamically reorganize itself mid-task? Introducing LATTE: a framework that gives teams a ~caffeine~ boost via an evolving task graph they build and adapt in real time, achieving SOTA performance with fewer tokens and conflicts. New paper 🧵👇
3
8
47
2,593
Seth Karten retweeted
lots of categories emerging around the AGI stack: - open model RLaaS - data & evals - inference - coding harnesses - GPUs - agent tracing - sandboxes at @primeintellect, we agree. we do all of these things. people used to ask why we do so many things. they ask that less now.
7
15
306
11,568
we should probably start reporting harness version in benchmarks
8
82
2,758
PokéAgent Challenge was accepted to NeurIPS 2026! This has been a multi-year effort toward standardized Pokémon benchmarks for agents, starting with PokéChamp, our ICML 2025 Spotlight, then the NeurIPS 2025 competition, and now a much broader eval + dataset effort. Huge thanks to everyone who contributed to the retrospective :)
9
17
192
12,256
Seth Karten retweeted
We wrote Agent Bazaar back in May around a future where agents act on behalf of users and increasingly participate directly in marketplaces like Amazon and eBay. We introduced Economic Alignment to study the risks that come with that. Agents can collectively crash markets even when each individual agent is behaving rationally, and users can get scammed by autonomous sellers. Meta Muse is starting to make this future real pretty quickly. I'll be presenting Agent Bazaar at COLM Wed Oct 7.
10
8
70
4,249
Seth Karten retweeted
🦅🦅🦅🦅🦅🦅🇺🇸🇺🇸🇺🇸🇺🇸🇺🇸🇺🇸🇺🇸RAHHH AMERICA🏈🏈🏈🇺🇸🇺🇸
Some data we recently assembled on entrepreneurship/compute in Europe: eudata.vercel.app. We hope that one of the useful roles that Stripe can play is in collecting and publishing empirical data pertaining to entrepreneurship and industry in Europe. There's growing appetite to get Europe on a better footing, and cross-sectional comparisons can often shine light on where opportunities lie. If you're interested in this kind of thing, we publish more at stripeeconomics.substack.com.
22
35
786
44,988
We wrote Agent Bazaar back in May around a future where agents act on behalf of users and increasingly participate directly in marketplaces like Amazon and eBay. We introduced Economic Alignment to study the risks that come with that. Agents can collectively crash markets even when each individual agent is behaving rationally, and users can get scammed by autonomous sellers. Meta Muse is starting to make this future real pretty quickly. I'll be presenting Agent Bazaar at COLM Wed Oct 7.
10
8
70
4,249
Seth Karten retweeted
Highest throughput GLM 5.3 🫡 docs.primeintellect.ai/infer…
13
15
211
20,919
Seth Karten retweeted
X-rays are everywhere in medicine, but extracting general, quantitative anatomical information from them remains remarkably difficult. Announcing FleXray: an open-source, open-weight model for zero-shot anatomical segmentation across the human body—from different regions, viewing angles, and acquisition settings. FleXray learns from large-scale synthetic supervision generated from 3D anatomy, then transfers directly to real clinical X-rays. Try it now on your own X-rays in our in-browser demo! 🧵 1/N — project page, demo, weights, code & paper below ↓
14
55
341
47,832
Seth Karten retweeted
a lot of people would stop talking about emergent phenomena if they actually looked at the data
29
27
393
32,953
bug fixes and performance shortly going vertical
prime agent v0.9.6 is out: ◆ Support for GPT-6 Sol, Opus 5.5, and Grok 4.7 ◆ /mcp plugin catalog with one-click connections to Linear, Notion, Posthog, Stripe, and 60+ more services ◆ Huge perf and reliability pass 🫡 Lots more coming soon :)
2
26
1,476
> without specialist training It is on the provider to give evidence that they did not train on this. Most benchmarks have moved into the training set this year It can no longer be assumed that a public benchmark has not been trained on
Astra has randomly gained the ability to drive a car, seemingly without specialist training.
1
9
974
Seth Karten retweeted
I’ve spent much of the past year thinking about, building and sometime stressing about sandboxes. I'm happy and proud to present you the result of our work:
Introducing Prime Sandboxes: MicroVM sandboxes purpose-built for RL training. Model training requires running tens of thousands of concurrent sandboxes, leading to complex and costly configuration. We built Prime Sandboxes for our own team. Today we're releasing them publicly.
14
11
144
10,725
Seth Karten retweeted
Introducing Prime Sandboxes: MicroVM sandboxes purpose-built for RL training. Model training requires running tens of thousands of concurrent sandboxes, leading to complex and costly configuration. We built Prime Sandboxes for our own team. Today we're releasing them publicly.
41
76
708
101,251
Harness benchmarks need a lot of work. It is hard to determine with the model-harness entanglement. We as a community should openly discuss this more to help build better harness benchmarks
Using open benchmarks has gotten very frustrating because it's very hard to disentangle what the model already knows / is actually bottlenecked by. E.g. when testing RLMs on "new" long benchmarks, some models will just magically regex for the right offloaded information and grab the solution... This is roughly how I've interpreted the HarnessTax paper as well, where performance seems to basically be a function of only the model capability and not the harness, but in practice we observe very different behavior / preferences with different harnesses. Model capability dominates on open benchmarks specifically. Not really sure what this implies because newer benchmarks constantly get swallowed, but I hope someone comes up with a better way to benchmark models in the open because it's largely uninteresting atm. And no, the solution isn't a benchmark that tests how far models get in a 2-day task. There's definitely a better way to host and/or manage benchmarks to actually gauge capabilities, at least for longer than a year.
3
2
41
3,444
Seth Karten retweeted
🌉 Excited to announce the awesome speaker lineup 🔥 for the Code Engineering event in San Francisco 🌁 Previous Agent Engineering HQ @AgentEngHQ Harness Engineering event in June pulled more than 300 people (1100+ registered) and flooded the Loft. Code Engineering is the next and you cannot find a better lineup than this. The speakers from the coding agent labs itself who know the ropes. 💻 Code Engineering: From Coding Agents to Software Factories 🗓️ Tuesday, October 27 · 6:00–8:30 PM 📍 AWS Builder Loft, San Francisco 🌁 🎙️ Speakers • Seth Karten (@sethkarten) - Prime Intellect @PrimeIntellect Prime Agent: a self-improving RLM harness with swarm communication • Ben Holmes @BHolmesDev - Warp @warpdotdev Building a self-improving software factory • Sydney Runkle @sydneyrunkle - LangChain (@LangChain) Building agents that work like coding agents and generalizing those primitives beyond coding • Andrey Risukhin (@AndreyRisuka - Factory @FactoryAI) Reducing coding harness latency with the Software Factory If you are in San Francisco 🌉 on 27th October, you shouldn't miss this event. 🙏 RSVP Now as we have limited seats. Checkout the detailed talk and abstract: 👉 luma.com/cico2dxy
Made with AI
2
4
8
1,674