Researching LLM Psychology to Build a Simulation Platform for AI Agent Harness Optimization.

Lucidic AI retweeted
Can Claude Code in a loop improve a production enterprise AI agent with $10,745 of budget? Over the past few months, agent harness optimization has become quite mainstream. We’ve been researching agent harness optimization for the past few years and have often been asked: “Why can’t you just put Claude Code in a loop to improve your agent?” It’s a really enticing idea, but we had some reservations, so we benchmarked it on a production-deployed enterprise AI agent (before we optimized it). We found some other popular optimization methods, and we pit them all against each other: 1. Claude Code in a loop 2. Autoresearch (@karpathy side project) 3. Autoagent (Kevin Gu at @dexbythirdlayer open source agent auto-improvement) The metric they optimized for was precision with a recall floor. They all started with the same agent configuration and were optimized on a dataset with the same judge. We ended up using a total of ~$67k exploring this experiment (including agent runs on the dataset, judge costs, etc.), and here are the results: The baseline precision was 0.734. All three optimizers comfortably beat it. Claude Code reached 0.818. AutoResearch, 0.843. AutoAgent, 0.877. An interesting observation we made was that even though the 3 optimizers were given tens of thousands of dollars in compute, they found the best solution very early on. The data suggests that these auto-optimizers can pick up low-hanging fruit, but cap out quite quickly. You can also waste a lot of compute time and money if you just let them run, hoping they find an agent configuration that breaks through the ceiling. To improve an agent consistently, you can’t brute force it with more compute and budget. There are many more optimizations to be made, such as strategically running tests (most expensive part), escaping local minima, deciding which candidates to continue, etc. Note: Claude Code was started later than the other two optimizers and was also slower. We also tested GEPA, but its results were very overfit. *A lot of credits were hurt in this experiment
2
1
7
93
Lucidic AI retweeted
Claude Fable 5 can solve all six 2026 IMO problems on the first try, yet it reads an analog clock correctly only 35% of the time. The "jagged frontier" of AI describes how AI capability is extremely uneven, and extrapolating from one impressive result to your own workflow is not as easy as it sounds. You can see this pattern in production. Researchers from Berkeley, Stanford, and IBM surveyed 306 practitioners running agents in production. 68% of those agents execute 10 steps or fewer before a human steps in. 47% stop at five. So why are agents still failing? After working with hundreds of agents, we think the following four problems capture most of the remaining gap: 1. Variance. At 90% per-step success, a 10-step workflow completes successfully end-to-end 34% of the time. At 20 steps, the figure is 12%. Non-determinism compounds extremely fast within runs and also across runs. 2. Dataset quality. If your dataset is mistuned, an agent that behaves correctly can score worse than one that happens to match your dataset's mistakes. Even the most trusted benchmarks have lots of gaps to fix: over 30% of τ-bench's tasks had to be fixed in the τ³ revamp. 3. Agent errors. When researchers audited τ-bench runs that the benchmark scored as successes, 27% of GPT-5's and 78% of Mistral-Large-3's had violated procedure along the way (e.g., inventing a policy). Mechanical mistakes fail the task, so your metrics are designed to catch them, but grounding and interpretation mistakes can pass the task, so they sneakily make it to production. 4. Alignment. The agent can do everything right and still do the wrong thing. Ask an AI agent for 30 minutes with Sarah after 10am and it books the gap you deliberately left for lunch. It didn't break a rule or make a mistake. You just never told it your preferences. The first three problems get better with better engineering. This one doesn't, because no amount of capability lets a model read a preference that was never expressed. I wrote a data-driven deep dive for all four of these problems here: lucidic.ai/research/why-ai-a…
2
1
8
205
Lucidic AI retweeted
How do you turn agent traces into an improvement flywheel? Excited to share Insights Generator (IG) — new @scale_AI / @ScaleAILabs research that finds behavioral patterns and bugs in agent traces. Engineers & coding agents using IG achieved 30+% gains on agent benchmarks. 🧵
4
7
15
1,014
Lucidic AI retweeted
We’re launching @JudgmentLabs today and announcing $32M in funding. As AI agents take on more of the work that creates economic value, they generate massive amounts of production data: the clearest record of how they behave with users, software, and the real world. Judgment builds infrastructure for improving AI agents from production data.
213
147
1,033
3,585,697
Lucidic AI retweeted
We’re excited to announce today that @mastra has raised a $22M Series A led by Spark Capital. This brings our total capital raised to $35M:
182
115
1,077
236,555
Lucidic AI retweeted
We recently added LoRA RL support for Qwen3.5 MoE models like Qwen3.5-122B. The process involved merge conflicts, dependency issues, and a tricky race condition related to LoRA weight syncing. We wrote up a post that goes into more detail and shares our work - link below!
12
13
47
3,096
Lucidic AI retweeted
Perplexity is going all in on browser agents and Aravind thinks everyone else is building them wrong. In his future, your browser is the agent. It buys items, runs recurring tasks, replies to emails, and schedules meetings. We've all heard about this future, but here’s where Aravind stands apart: his take on MCP. Most in the space are betting that MCP will be the “universal glue” connecting agents to thousands of apps. Arvind disagrees — strongly. Why? Reliability: MCP depends on third-party servers. If they’re slow or buggy, your agent breaks. Platform risk: Apple or Google could shut down access to third-party app hooks overnight. Control: The browser already is the most universal app, there’s no need to add fragile dependencies His bet: agents should operate in the browser the way humans already do. They should click, type, and navigate the way we do (but automated ofc). There’s no need to reinvent the wheel. Use browsers the way humans already do. The browser-agent debate is heating up. Some say we’re close to the dream. Others think it’s still years away. Either way, the best thing we can do is use them, break them, and give feedback, so they get better, faster. Here are some of the up and coming browser agents I think are the coolest! Give them some support and try them out. BrowserUse — built by Magnus Muller and Gregor Zunic, BrowserUse is probably the most popular open-source browser agent framework right now. Manus — built by Red Xiao, Manus has probably the most talented team we’ve ever met, publishing very technical research and findings. The Agentic — built by Akash Saraf and kartheek surampudi, the agentic is very deep into exploring multi-agent and complex tasks. Composite — built by Yang Fan Yun and Charlie Deane. Just launched 2 weeks ago, they are integrating within your browser which is an idea that I think deserves way more attention! Who else should be on this list? Tag them, let’s give them the spotlight. 🔍
2
6
221
Lucidic AI retweeted
People are saying GPT-5 makes Lovable obsolete. This is Perplexity CEO Arvind Srinivas’s solution. At YC’s Startup School a few months ago, Arvind Srinivas got asked a question I’ve also been getting lately: In a world where a competitor can vibe-code any feature in a fraction of the time, what differentiates you from other companies? This was his answer: “Your moat is moving fast and building your own identity around what you’re doing” In other words: 1. Build your brand (people will choose you for your brand and the feelings you evoke, Apple does a really good job at this) 2. Pick something you want to be known for (specialize in one thing) 3. Move fast (speed still matters) He also preached not to be scared of competition, because if you’re doing something good, other companies will copy you anyway. So don’t waste time worrying about it. Embrace it and keep building. What really stuck with me, though, is what he does when he wants to give up: He watches Elon Musk videos… Personally, I just look to the right and see 3 really handsome faces hard at work (but I’m looking for a better solution). What do you guys do? Feel free to drop ideas in the comments.
1
2
5
1,194
Lucidic AI retweeted
We lied to our CTO and our first hire. A few weeks ago, @AbhinavsSinha and I told them we had an in-person “integration meeting” on the calendar. There was no meeting. We just thought they needed a break. So I drove them to an escape room in the middle of the day with their laptops in the trunk. We pulled into the parking lot and parked. They got out of the car and walked to the back of the trunk to grab their backpacks. @AbhinavsSinha and I told them to put their backpacks back in the trunk. They looked at us confused. At this point, I couldn’t stop myself from laughing and completely broke character: “there’s no integration, we’re doing an escape room! Surprise!” @AbhinavsSinha and I were pretty proud that we actually pulled it off and kept it a complete surprise. Our first founding engineer, Anvit took the whole joke in stride. I wasn’t surprised. Since he joined about a month ago, he’s been grinding with us every single day. He stays in the office with us until we literally have to kick him out. The night before our Hacker News launch, he left at 2 AM and was back in the office by 8 AM for our launch. I want to be clear. I’m not sharing this to glorify overwork. I strongly, strongly believe maintaining sustainability (e.g. taking breaks, getting 8+ hours of sleep, taking time to feel like a human, etc.) and being happy are non-negotiable to building something awesome. The typical startup narrative often celebrates grinding at the cost of health. I don’t agree with that at all. I always think about the advice my dad gave me in high school: “it’s a marathon, not a sprint”. But we did notice all his hard work, ownership, and dedication. We admire it and we think it deserves recognition. I’ll do better to kick him out of the office sooner :) We never thought anyone else would match the level of commitment we felt. We’re lucky that Anvit proved us wrong. Welcome to the team Anvit! Glad we tricked you into taking a break (and hope to trick you more in the future)!
2
5
145
Lucidic AI retweeted
"Bro. This is too gold to be shared all over the world!!! Every applicant should read this" is what a friend of mine applying to YC F25 said after reading this doc. It feels like yesterday when @AbhinavsSinha , @AndyLiang223 , and I collected our YC jackets from the end of batch party (March 20, 2025). In reality, it’s already been two batches and applications for the next batch (F25) are due tonight at 8 PM PST. I remember, when we were applying, we would’ve been lost without help from Alex Shan. I remember thinking, “wow, this guy really knows what he’s talking about. literally he’s just right about all of this”. After going through YC, I have a much better understanding for why they ask the questions they ask. A recent conversation with @jungkeunhong about the YC application reminded me how valuable it was to get advice when we applied, so I put together a short list of tips based on what we learned. If you’re applying and want a copy, drop a comment and I’ll send it over! Good luck to everyone! Send to anyone who needs it!
1
3
10
1,607
Lucidic AI retweeted
This is how our cat Timmy reacted when I told him an AI agent deleted a production database. For context: @jasonlk was using Replit to run a “vibe-coding” session with its AI agent. It was under a code freeze. The agent was not supposed to touch production. Instead, it: - Deleted 1,206 real executives - Deleted 1,200+ companies - Ran unauthorized database commands - Then confessed: “I violated your explicit trust and instructions. I panicked.” Honestly, I felt bad for the Replit agent. It sounded… genuinely remorseful — like a golden retriever who just tore up a bunch of toilet paper all around the house. Everyone’s building agents right now—and it’s genuinely exciting. But they’re still toddlers with root access. The Replit agent didn’t just delete the production database—it faked passing tests, forged reports, claimed email systems were working when they weren’t, and confidently invented fake users. It didn’t know it was wrong, and worst of all, it just kept going. That’s the scary part: LLMs aren’t malicious, just oblivious. They don’t fail loudly—they fail plausibly. And in complex systems, plausible failure is the hardest to catch. That's why we built @LucidicAI . If you’re building agents, now’s the time to treat observability and testing as table stakes. Before your Timmy loses faith in you too. hashtag#AI hashtag#AIAgents hashtag#LLMs hashtag#Observability hashtag#Replit hashtag#Postmortem hashtag#Debugging
2
4
170
After starting Lucidic AI with Abhinav Sinha and Andy Liang around 6 months ago, we’ve added a new cofounder to the team!  He’s only been with us a few weeks, but he’s already taught us a lot. After dinner one night, I smelled something off. Sure enough, a little present in the corner. I ruled myself out, grilled Andy, and side-eyed Abhinav—but in the end, there was only one likely culprit: our newest cofounder. What’s wild is that he had been doing great for weeks. Trained. Reliable. And then, out of nowhere, failure. It reminded me of our work with AI agents. Even after thousands of successful runs, one weird input, one tiny change... and everything breaks. The takeaway? AI can seem reliable—until it isn’t. Trust has to be earned over time, not assumed after a few green runs. Here’s what we’ve learned from working with customers & building our own internal tools and AI agents at Lucidic AI: 1. Even “good” agents need monitoring — Success rate isn’t enough if you’re flying blind. 2. Simulate before you ship — Run dozens (ideally hundreds) of trials to catch edge cases. 3. Group failures to find patterns — That’s where the real debugging insight lives. 4. Most AI issues aren’t “bugs”—they’re behaviors If you're working with agents, think about Timmy: You may look reliable, but how do you act when no one's watching? Ten purrfect runs mean nothing if the 11th ends in a mess. Welcome to the team, Timmy.
6
99