Who grades the graders? LLM-as-judge on short answers is well studied. Agent judges are not: they open the files, run the code, and investigate before ruling on another agent's work. Introducing Arbiter-Bench: 71 real agent runs that frontier agent judges get wrong.
13
16
51
3,097
Of 232 runs where a frontier judge disagreed with the benchmark, only 71 survived a blind label audit (3+ model families, a human ruling on every split vote). About a quarter were the benchmark's error: hidden criteria, broken checkers, unstated rules.
1
2
218
Half the set is public with every judge's full logs, a scorer for your own judge, and the audit trail. 35 items are held out. Every case is a runnable @harborframework task, so you can point your own judge at the set and score it with one command. Full writeup: heuristic.build/blog/arbiter…
1
8
319
funny seeing a problem you've been working on suddenly get so much attention
1
1
10
787
raphael lee retweeted
so called high quality data showing gains on evals so flawed where not a single sample is surviving basic scrutiny is a sight to behold
4
5
114
5,879
raphael lee retweeted
Tech has decided that regulation is the enemy of progress, but innovation requires capital Regulation builds the trust necessary to inject billions into markets To enable free innovation, Andera raised a $37M series A led by Lightspeed to scale financial oversight with agents
237
129
818
702,591
raphael lee retweeted
Over 1 billion sandboxes have been launched on Modal. Since launching three years ago, we've seen Modal Sandboxes become foundational to how AI is being built. Today, teams like @Lovable, @tryramp, @cognition and more are using Modal Sandboxes to power everything from coding platforms and background agents to RL infrastructure at scale.
13
19
248
111,113
did poke mention @marsxiang_ to every waterloo student who signed up
2
16
2,459
whos going to hackmit
5
12
1,944
raphael lee retweeted
meet Shadow, a powerful open-source background coding agent! fully featured with a remote environment, codebase indexing/wikis and subagents to understand, write and test your code, directly making PRs to github built in a few weeks w/ @ishaandey_, @ElijahKurien
30
22
292
53,614
apparently doomscrolling improves my sleep quality. do I trust my whoop or @bryan_johnson?
2
15
1,998
This weekend, we built a robotic wearable that helps Parkinson’s patients walk independently. A 1D CNN + LSTM model predicts freezing of gait in real-time and triggers mechanical stimulation to restore movement. 1st Place + Hacker’s Choice @ Berkeley AI Hackathon
9
1
42
5,737
Recent research (Ibrahim et al., Nature 2024) shows targeted muscle stimulation can eliminate FOG. We used a wearable sensor (MPU-6050) to collect leg motion data. Using our model's prediction, a robotic arm delivers a localized “poke” to end the freezing episode.
2
7
1,180
Built with Shuyu van Kerkwijk, Harry Cheng, and Isabel Lew. Thank you to the @CalHacks organizers for this opportunity!
4
715
raphael lee retweeted
playing poker at @spermracing hq (w @_raphael_lee_ @bruce_wang15)
1
12
1,566