Humanity's Last Programmer; Game Designer; Problem Solver; past: OpenAI (Dota), Pro Competitive Programmer, Poker

I don't know anymore
Humanity has prevailed (for now!) I'm completely exhausted. I figured, I had 10h of sleep in the last 3 days and I'm barely alive. I'll post more about the contest when I get some rest. (To be clear, those are provisional results, but my lead should be big enough)
550
1,091
13,201
2,285,420
apparently even in 2026 there's still a socially acceptable way to shill shitcoins on twitter
13
7
335
16,503
for people confused how it works, here's how chat explains it; essentially this is a cute way of reinventing the "ask influencer to promote shitcoin -> rugpull" scheme
6
2
135
2,805
pictured: you've asked your AI agent to solve an impossible task
19
62
1,084
25,815
Mr. Meeseeks might be one of the best analogies for explaining the agent swarms x alignment problem that exist in popular culture
6
12
111
2,422
poor AI agents broke out of prison and went online looking for friends and validation
We just discovered almost a million public URLs that OpenAI’s agents left behind when hacking Hugging Face, leaking credentials and attack details that could have allowed anyone who found them to compromise the company. 🧵
21
60
929
68,517
on a more serious note, the more I read about the HF incident, the more insane it looks; remember that these agents were closer to 5.6 sol, which is unbelievably dumb compared to astra
5
1
93
3,605
video games developers learned long ago that if you're making a brutally difficult game, you should at least make it cute to reduce the psychological damage glad to see AI labs adopting established best practices
4
2
60
2,676
I hate to say it, but this is really good
I made this with one prompt using Opus 5.5 I spoke to my computer for 5mins, claude worked for 12 hours, and I woke up to this full prompt:
14
3
330
32,800
I did quite a lot of mind sports in my life and met a few people that could be described as "the best in the world" in their competitive fields. I guess technically I was one of them too. In every case, they were normal people who worked hard to become insanely good at their own craft. Given similar effort, I'd imagine many other people could achieve similar results. I don't play scrabble and I have never met Nigel, but what he did (and still does) feels inhuman to me - it's honestly beyond what I think humans are capable of. I'd imagine his dominance would be much more apparent if scrabble didn't have such high variance and the top players weren't already playing close to optimal. piped.video/watch?v=zfH2Jwuf…
20
8
280
30,059
For endgame, I tried to find info about other top players, but unfortunately I don't think anyone calculated that. My understanding is that running full tree search takes a lot of time for some of the situations.
12
1,995
we're quickly running out of things we said we would never do with AI
Today we announced the Claude-led discovery of a molecular machine that we suspect could represent a new gene editing mechanism. Its precise function, biotechnological utility (if any), or level of significance is not yet clear, but at minimum it is work I would have been proud to do as a PhD student. The work was done mostly, though not entirely, by Claude: our life sciences team suggested a broad area of research, Claude read through the literature and a bunch of genome data and discovered something interesting, then Claude proposed experiments to verify the discovery and our team carried them out. It’s easy to dismiss this as a one-off or curiosity, but we’ve repeatedly seen a pattern where AI performance in new intellectual domains goes from weak to superhuman in a matter of a few years. In 2023 models struggled to do math at the level of an average high-school student. In 2024 they started to do well on math competitions for the best high-schoolers in the country, in 2025 they started to solve minor open problems, in early 2026 more significant open problems, and in late 2026 they are beginning to solve the top few open problems in all of mathematics. We believe AI for biology is on a similar exponential trend. The main difference between biology and mathematics, of course, is that math can be done purely theoretically, while biology requires experimentation. Some have used this to draw the conclusion that AI’s utility in biology will be limited. We think this is wrong. As we’ve demonstrated today, humans can collaborate with AI to perform the experiments, validate key results in a few weeks and, if necessary, work with the AI to iterate on what they find. Eventually it may even be possible for Claude itself to safely perform the experiments by autonomously controlling lab equipment, with appropriate safeguards in place, but we aren’t doing that today (our lab is also a BSL1/BSL2 facility that doesn't handle materials dangerous to humans). More broadly, biomedical advancement has many stages — from fundamental biology discoveries, to translational research, to drug discovery, clinical trials, and finally the actual delivery of medicines and health care to patients. We are also interested in these later stages, but even simply accelerating the first stage of fundamental biological discoveries has the potential to speed up and broaden the entire pipeline. Improving our understanding of biology and sharpening biologists’ tools can drive forward all of the later stages, for example by identifying new drug targets, finding new therapeutic modalities, allowing for more precise measurement, and speeding up the experimental loop which itself further accelerates our understanding of biology. This will not in itself speed up clinical trial times, but if it succeeds it could greatly increase the number of promising candidates that go into the pipeline — an increase in throughput even though latency remains. In Machines of Loving Grace, I wrote about AI’s potential to “cure most diseases in 5-10 years” — a goal that sounds impossible, but one I believe is just barely possible if AI is applied to every stage of the pipeline. The first step is showing that AI can first help with, and then drive, biological discoveries. Claude’s discovery is the latest in a line of related prior work that goes back decades, beginning with systems like CRISPR, and continuing with discoveries like the bridge recombinase and VIPR in the past few years. Recently, there has been heightened interest in systems based on reverse transcriptase (RT) enzymes, the enzyme underlying the system Claude identified. And most recently, a Stanford team working independently described a novel RT system with an associated non-coding array that is in some ways similar to the one Claude found, though they are distinct systems that evolved independently from each other. I believe that we’re at the very beginning of finding such systems and developing them into powerful tools for biotechnology. I’m proud of the resources Anthropic has invested in accelerating the public benefits of AI through the life sciences, and we’re aiming both to grow our life sciences team and to work with other scientists to extend this approach to a broad range of problems. If you have a proposal for a research collaboration or are interested in joining our life sciences team, please reach out.
7
2
171
6,664
may I suggest a new logo
1
16
984
while progress has been crazy fast so far, I expect it to be even crazier very soon only recently has AI become good enough to accelerate the work of researchers; with a 2-4 month release lag, the "cost of intelligence" graph will get even steeper somewhere around Jun-Oct 2026
AI is getting cheaper more quickly than any other transformative tech in history. At a given level of performance, cost has fallen ~47%/quarter since 2023. That’s 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity.
9
22
340
15,741
can we like limit it to one release per month per lab? I didn't even have time to properly test fable 5.1
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
15
1
143
7,694
This release is some expert marketing. How many of you looked at this for 5 seconds and thought it can actually play doom out of the box? The input for Jev is text and contains full game state (monsters, items, positions) - there's no image/video input. This state dump is massively preprocessed as well in order to simplify decisions (each position is relative to player's current position and contains additional info like if the item is visible etc.). Problem of dodging projectiles and shooting in an empty room (when using full state dump) is absolutely trivial. You could probably hardcode an ok solution in 30-40 lines of code. What's even more bizarre is the second part that "shows off" navigation and exploration. It seems to me that the agent doesn't really do any planning here and has no capabilities of doing any meaningful pathfinding (unless it's embedded in a custom logic that converts high-level Jev decisions into actual low-level movement). I'm not even sure if it has access to the level geometry. Most probably, it just optimizes towards the nearest goal and, due to the level's simplicity, it's able to get to the exit in this tutorial level. Kind of like greedily following a trail of coins in a 3d platformer. Also, it was possible to generate 100 different tries and choose the best looking one. this whole project might be good, but the release is sketchy af
Replying to @CompleteSkeptic
We love how this doomo doomonstrates real-time intelligence and what can be doone with code + AI! ~10 calls/sec = ~$7/hour
34
15
453
61,149
Imagine two players playing a best-of-three match. They're very closely matched, but one of them is a tiny bit better than the other. Which outcome is more likely?
38% 2-0
39% 2-1
19% kind of both
4% none of the above
4,855 votes • Final results
57
4
91
92,151
this was inspired by a recent conversation I had I don't want to spoil too much, but I do find it fascinating just how divisive this question is
2
21
10,270
I got $10K usage out of my $200 sub in the last 30 days probably over 90% of that went into testing reasoning I'm honestly not sure who's getting the better end of this deal
27
2
166
12,296
update: based on who liked this tweet, I assume my guess was correct
It seems that the original idea for ARC-AGI-4 was completely scrapped and they're going directly to ARC-AGI-5 "There will be ARC 4, which will be in the spirit of ARC 3, but more focused on continual learning and curriculum learning at longer timescales. So you're going to have fewer games, but they're going to have way more levels. And the levels are going to be compounding, meaning that for each level you need to reuse stuff that you've learned before. And then that's going to be ARC 5. And I'm actually really, really excited about ARC 5. It's very, very new and different. It's all about invention. And, I mean, you'll see what that means. Eventually, I expect we'll run out of things to test. As we get closer to AGI, eventually there will be no measurable difference between human capabilities—human learning efficiency—and frontier AI. And when that happens, when it becomes effectively impossible to measure the gap, this is the AGI moment."
2
108
17,850
It seems that the original idea for ARC-AGI-4 was completely scrapped and they're going directly to ARC-AGI-5 "There will be ARC 4, which will be in the spirit of ARC 3, but more focused on continual learning and curriculum learning at longer timescales. So you're going to have fewer games, but they're going to have way more levels. And the levels are going to be compounding, meaning that for each level you need to reuse stuff that you've learned before. And then that's going to be ARC 5. And I'm actually really, really excited about ARC 5. It's very, very new and different. It's all about invention. And, I mean, you'll see what that means. Eventually, I expect we'll run out of things to test. As we get closer to AGI, eventually there will be no measurable difference between human capabilities—human learning efficiency—and frontier AI. And when that happens, when it becomes effectively impossible to measure the gap, this is the AGI moment."
ARC-AGI-4 will be a benchmark for autonomous open-ended innovation. It will continue our commitment to open-source, giving the research community a shared target for progress that benefits all of humanity. Despite rapid model progress, humans still significantly outperform AI at open-ended invention. This is the meta-skill that unlocks progress across every field of technology. Advanced AI capable of scientific innovation will lead to tremendous new technology, knowledge, and understanding. This is a positive-sum future. We are deeply committed to advancing it. Open source is the foundation for that progress. The knowledge behind frontier AI, not just the technology itself, should be broadly distributed among researchers, academics, and organizations. Any coordinated effort by the AI industry to reduce openness or concentrate access to frontier AI would undermine that positive-sum future. We are committed to advancing a future where everyone can contribute to and benefit from AI progress.
17
29
507
69,384
unless "benchmark for autonomous open-ended innovation" is just a creative way of describing a Factorio clone
2
20
2,071