Independent Researcher, Partner @ ScOp VC

Santa Barbara
Collaborate with more people. They will give you ideas and a sense of what's going on. It's a way to build a network of closer contacts who can refer you to opportunities and vouch for you.
17
358
1/ Is Terminal Bench broken? Epoch audited 15 benchmarks and labeled @terminalbench flawed. They used our own PR history and our reviewers' acknowledgement of issues.
We audited 15 benchmarks and labeled 9 flawed: - In Terminal Bench 4.0 we found 45.5% of tasks to be broken after reviewing github issues. - In HLE, 46% of the 48 questions we randomly sampled were broken. - In DeepSWE 1.1 we found a bug that can break grading for every task.
2
2
30
5,317
11/ A benchmark with public PRs is the right mechanism. Software has bugs, and fast iteration creates more bugs. A flourishing benchmark with bugs that get fixed quickly is better than a battle tested stale one.
I was excited to hear about this initiative, but a bit disappointed in the rigor -the “task flaws” would affect < 3% of rollouts on the leaderboard -the impact on the final agent scores would not exceed our reported confidence intervals we publish the raw receipts for all our leaderboard runs (per trial logs / trajectories / rewards) so this is pretty easy to figure out. It is also pretty easy to find new task issues (that is why we release the receipts!). we love when the community reports (new) issues. no new issues were opened as part of the report - the “flaws” are simply pulled from issues triaged by the TB maintainers from our public GitHub repo. fwiw 100% of terminal-bench tasks are flawed (all software has bugs). what matters is the magnitude of the impact of those bugs and the mechanism that the benchmark creators have for continuous detecting and addressing bugs. this is the motivation behind our “continuous benchmarks” methodology. I’m glad that epoch is starting an important conversations on benchmark quality. However, auditing a benchmark should not be a one-time thing: coming together as a community to build public tools and define practices for task CI / CD is more in service of the goal of enforcing good benchmark practices. we have been experimenting and pushing the possibilities of these benchmark CI / CD tools and practices as part of the terminal-bench project and are open wide collaborations with all who are interested by that problem
1
4
272
Here's the paper (still work in progress) arxiv.org/pdf/2609.26826
3
158
I see sooooo many companies living off the RL data firehose. Labs, due to their distaste for getting their hands dirty with data, are creating a whole outsourcing industry. The results won't be optimal. Post-training data is not a commodity, unlike its complement. A lot of thought has to go into building environments worthy of gradient descending. Labs not taking more ownership of the data supply chain makes me more bullish on vertical AI.
4
3
111
7,338
I simultaneously hold the views that some people's overpredisposition to wearing masks is dumb and that riding a plane while sick without a mask is antisocial. Why do things like masks become political symbols and cease to have context-dependent utility?
1
217
Three scenarios in which a task might have passing trials and still be broken. In these cases passing means giving a fundamentally incorrect answer that the verifier accepts: - The task was made with an LLM that has a misconception about the domain, and then the same LLM is used as the agent and passes. - The task initially failed, but labs trained on the benchmark (possibly using the oracle or verifier as hints), which made the agent learn the wrong solution, and now it passes. - The verifier and oracle are incorrect, but close enough to the correct solution that the agent sometimes passes by chance.
1
7
519
My daughter took me to a heavy metal festival. I've gone to a lot of festivals, but more of the electronic/transformational vibe. Some contrasting observations. Unusually high obesity rates. Significant tattoo and piercing coverage, including frequent face ink. Many people with physical disabilities. Generally more diverse in demographics. I didn't notice excessive drinking. I'm pretty sure drug use is meaningfully lower. And in spite of provocative outfits, I didn't get any pervy signals. The "wall of death" would be safe enough for a grandpa. Crowd surfing was fun and communal. There were a lot of children and teenagers with their parents. Everything was extremely well organized. Probably the most hygienic festival I've been to.
2
21
1,297
Maybe we should give agents a scratchpad that is guaranteed not to be monitored, except for forensic analysis when something goes wrong. We have privacy contracts between humans, which seem to work well.
1
4
440
Europe might have more regulation and its people might be less into their jobs, but the opposite is evident when you interact with airport security.
6
288
I've seen a lot of people hanging out at parks in Berlin during regular workdays. It reminds me of the vibes at a music festival. People bring some floor covering, drinks, snacks, and hang out for a while in small groups. I've seen this in the US, but it's not nearly as prevalent, and it seems constrained to a subset of social groups. Last night there was a group playing incredibly loud techno music late at night on a field, and nobody cared. They told me it was a test for an even louder open party they were planning a couple days later. I just walked through a cemetery on a September Tuesday at 1:15 pm, and there were a number of people who had stripped to their underwear and were lying on the grass, many of them alone. They seemed like normal people. It wasn't particularly sunny; they just wanted to be comfortable during their lunch break.
1
3
392