Member of the load-bearing staff

Will appear at Neurips. Its impact so far tho is the real validation.
Once again, SlopCodeBench is at the top of hacker news:
2
15
1,389
Aws Albarghouthi retweeted
Will be at NeurIPS '26! Congratulations to @jiayuwang111 and @ming5_alvin
Agent skills are powerful. But can they help us with agent orchestration and routing? We show they can! 🎼 Introducing SkillOrchestra - a skill-aware alternative to RL-based agent orchestration: 💰 Higher efficiency 📉 Orders-of-magnitude lower training cost 🧠 More interpretable decisions TL;DR: Stop training bigger orchestrators and start modeling skills. Paper: arxiv.org/abs/2602.19672 [1/n]
6
46
2,113
Aws Albarghouthi retweeted
My team at @QuantinuumQC is hiring a Principal Quantum Compiler Engineer! We're tackling some fascinating compiler problems on the path to fault-tolerant quantum computing. Come work with us or please share! jobs.eu.lever.co/quantinuum/…
3
5
856
Aws Albarghouthi retweeted
An unforgettable few days at Xanadu HQ! We had the pleasure of hosting the International Workshop on Quantum Computing (IWQC) over the weekend, followed by Quantum Resource Estimation (QRE) on Monday. Bringing researchers, developers, and speakers under one roof to collaborate on everything from compiler infrastructure to photonic hardware is what this community is all about. Thank you to everyone who made both events so impactful!
3
4
33
1,794
Cherny is using formal verification exactly how most people have used it for decades: model a system’s dynamics and discover issues in the process. Some people got Turing awards for such ideas.
I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs and race conditions. Video attached. TLA+ also works well. I sometimes combine Lean and TLA+ to look for issues around data flow, concurrency, and state mgmt. I don't know either language well, but Claude is excellent at both. This approach is super useful for formally modeling your code and finding bugs that a human probably wouldn't have spotted. Is formal verification the future of coding (or at least, bug finding)?
4
5
89
17,308
Aws Albarghouthi retweeted
I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs and race conditions. Video attached. TLA+ also works well. I sometimes combine Lean and TLA+ to look for issues around data flow, concurrency, and state mgmt. I don't know either language well, but Claude is excellent at both. This approach is super useful for formally modeling your code and finding bugs that a human probably wouldn't have spotted. Is formal verification the future of coding (or at least, bug finding)?
478
320
5,639
1,950,676
Aws Albarghouthi retweeted
I'm hiring a postdoc! Please share broadly with potential candidates. I'm looking for someone with expertise in PL, FM, and/or networked systems who is excited about contributing to a research and teaching in a growing lab. careers.epfl.ch/job/Lausanne…
1
17
54
3,777
🔥
Here’s my first attempt to grapple with AI companies’ conquest of mathematics and our new mathematical condition. argmin.net/p/applied-pure-ma…
3
1,122
So Claude has read all of Phil Wadler’s papers but still can’t write. Can an ML person explain?
9
308
Aws Albarghouthi retweeted
“Once you free your mind about the concept of harmony and of music being 'correct,' you can do whatever you want.” - Giovanni Giorgio Moroder, but his friends call him Giorgio
6
2
34
4,670
I've been crystalizing some advice for first-year college students: "Advice to the New College Student in the AI Era". Thesis: "the purpose of going to college is to emerge as close to middle-aged as possible". docs.google.com/document/d/e…
14
41
257
16,803
Aws Albarghouthi retweeted
finally finished a full SlopCodeBench run against GLM 5.3, Sol 5.6, and Astra (Fable 5.1 results coming soon) This is different from all our previous research on SlopCodeBench (from @GOrlanski @ UW) in that we ran the full benchmark, every single scenario here. Previous runs did a small subset of the challenges. Asterisks: - this was run over ~1 week and had to be resumed a few times due to various provider outages - I still think a more scientific approach here would be to do what's common with other benchmarks, which is run multiple evaluations and aggregate the scores Looks like Astra scores a few points higher than gpt-5.5 here. Not as many as I'd think. I'm surprised Sol got lower than gpt 5.5 because in my experience I like working with Sol a bit more. My vibes-best guess is that the newer models are likely to go more off the rails on higher thinking modes (e.g. I almost always use Sol in medium or low effort for most work on @humanlayer_dev)
18
5
146
16,318
We should rebrand Formal Verification to Proof Against the Machine
5
48
1,737
Aws Albarghouthi retweeted
When code is abundant quality becomes the differentiator, but what is good code? New blog post from Earendil engineer @SebastianBaye on why quantifying the ‘sloppiness’ of code is difficult, which eval methods don’t work well, and what metrics matter. Link to the post below
22
37
716
49,581
Once again, SlopCodeBench is at the top of hacker news:
10
2,315
Our promo video for the College of AI is basically @EthanCecchetti designing a type system in the dark :)
4
1
32
3,039
West coast: AI Safety Midwest: AI Progress and Preservation
3
200
I’ve managed to port my quantum circuit optimizer (Rust) to Lean, along with a proof of correctness. Still significantly slower than Rust, but provably correct :) Pretty wild that such projects used to take years and now they take just a few days. github.com/qqq-wisc/tzap/tre…
3
2
54
3,261
Aws Albarghouthi retweeted
SlopCodeBench Update: Fable 5.1 Medium - 11% - I think we need to re-run this at higher effort GLM 5.3 Medium - 7% GLM 5.3 High - 19%, still in progress GPT-5.6 Sol medium - 11% GPT-5.6 Sol XHigh - 18%, still in progress GPT-6 Astra Xhigh - 21%, still in progress This is the first time i've done a run with multiple models at different thinking settings, since I don't thinking effort translates cleanly across providers. My next fable reset is in 4 days so unless @trq212 can to find me some credits ❤️ I will have to hold off on the fable 5.1 xhigh until next week -
24
10
208
29,397