builder, ai, markets, and general data nerd | dodatathings.dev

In the lab 🧪
Anthropic got all their users like
Anthropic is on a legendary run. Max and Team subscribers get $100/$200/$500 in API credit to test EVERY month anthropic.com/claude-haiku-5…
19
Right outcome, unearned action. Across ten model-harness configurations, agents that judge well when asked whether an action is justified do much worse when they have to gather the evidence and execute it in a live workflow. Failures often start before execution. An agent can update the correct record on a guess and still get graded a success. Outcome-only evals count that run as a pass, so the pass rate overstates how safe it is to hand the agent write access. Teams shipping agents with write permissions should log whether the supporting evidence appeared in the trace before each write call fired. arxiv.org/abs/2610.07753
2
1
34
This is my AI Engineer going on call
Ben Affleck reveals he writes Python, understands convolutional neural networks, worked extensively with GPUs, and used his celebrity status to get private looks at Google and OpenAI’s video models “I’ve always been kind of into computers since I was young. Then, when film started to move from analog film to digital, I became more interested in that aspect of it. The visual-effects workflow for many years has included machine learning, so I can write pretty shitty Python scripts and stuff like that. “With convolutional neural networks, which were the precursors to what the transformer can do, which is much more computation simultaneously, you would do things like look at what’s called a tensor. That’s the numerical translation of a visual image in numbers, like the batch number, the frame number and the red, green and blue values of each pixel in each frame. It’s just that simple. That numeric is called a tensor. “You’d use a convolutional neural network to identify patterns that reveal what’s called edge detection or feature extraction, which is identifying patterns well enough to know, this is where the window ledge is, so we can more easily take the green-screen image out and replace it with something. “That was familiar to me early on because, prior to Artists Equity, I had a small visual-effects company. I’ve worked with GPUs a lot too. The visual-effects guys said, ‘Hey, you should see. There are a couple: Google and this other company, OpenAI, are doing really interesting stuff with transformers in video.’ “I’ve learned that I can actually just call up and go, ‘Hey, it’s Ben Affleck. Can I come see what you’re doing?’ Sometimes people say yes, to my astonishment.”
43
This open-source repo automates one of the most complex pieces in all of wealth management: the personal financial statement. The PFS sits at the center of wealth planning. It's where everything meets: account balances, paychecks, taxes, the mortgage, college in 9 years, retirement in 24. Each piece comes from a different place and rests on its own assumption. The PFS weighs what the household will earn against what it has promised to spend, and puts a price on each goal. The fundamentals are the same for every household, and it's beautiful when put together. But it is a nightmare from a data monitoring perspective. Here you get it for free.
1
75
Replying to @cgtwts
It's also nitpicky, but their site copy says DDR5 while the sticks they show are DDR4 🫠
2
122
Depth goes unused. Thirteen pretrained base models reliably follow only 1.4 to 3.6 lines of in-context references. A rank-8 LoRA at one early layer, all other weights frozen, takes Qwen3-8B from 15.5% to 99% exact accuracy on 24-line chains. The layers were already there. Pretraining never rewarded using them for multi-hop lookup, and that same skill sits under tracing a variable through a long file or chasing a definition across a document. Looped models show the ceiling is much higher: Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. arxiv.org/abs/2609.36585
39
Literally wtf.
1.) Go to America.gov 2.) type “hello” 3.) type “play Minecraft” 4.) watch the agent’s reality implode 5.) ???? 6.) PROFIT!
1
56
Prefill costs more than generation. Long-context agents don't spend their budget writing tokens, they spend it reading them. Every tool call, every retrieved doc, every turn of memory gets re-processed through the KV cache, and that cache has to live somewhere: HBM if you're lucky, SSD if you're not, with bandwidth taxed either way. DeepSeek is targeting that cache directly instead of chasing another benchmark number. If your agent calls tools 50 times in a session, the cost isn't the output, it's remembering the input. arxiv.org/abs/2609.19969
1
1
42
300 turns before the payoff. That's the number that matters in T1's recipe: a 122B MoE model running a real shell for 300+ tool-call turns per task, rewarded only when its own verifier says the task actually passed. No credit for a plausible-looking intermediate step. Most agent training still rewards short completions where feedback is nearly instant. T1 is built for tasks where the reward signal only exists thirty command chains later, which is the actual shape of debugging a production system or running a multi-step experiment. arxiv.org/abs/2609.11042
1
3
44
State is the missing layer. Video world models can generate a wildly convincing next frame and still forget that the door they showed closed two seconds ago was supposed to stay closed. They're guessing pixels rather than tracking facts. This framework routes instructions through an actual program that tracks entity states and transition rules, so the world obeys rules instead of just looking plausible frame to frame. arxiv.org/abs/2609.10540
1
2
42
Same output, faster generation. Every token in an LLM has to wait for the one before it, which is why generation feels slower than reading. This paper trains a separate diffusion head that samples multiple future tokens at once from the same distribution the autoregressive model would produce one at a time. The "lossless" part is the claim worth sitting with. Usual speedup tricks trade accuracy for speed. This one keeps the AR model's output distribution intact and just parallelizes the sampling step. If that holds outside the paper's benchmarks, it's a latency win with no quality tax. arxiv.org/abs/2609.04010
2
2
37
YouTube is a robot dataset now. RoboTok retrieves human hand demonstrations from web video by matching 3D hand trajectory patterns, then uses those clips to train dexterous manipulation policies. No teleoperation rig required to show a robot how to pour, twist, or pinch. Robot learning has been capped by how much demonstration data you can physically collect, one arm, one lab, one task at a time. This turns every cooking video, repair tutorial, and craft channel into potential training signal for the long tail of tasks nobody bothered to teleoperate. arxiv.org/abs/2609.03199
2
74
Everyone’s CLAUDE .md about to look like
the end of markdown files. bitter lesson always wins.
3
119
Random beats smart here. Every KV cache eviction method scores tokens for importance and keeps the winners. This paper evicts at random within each attention head, uniformly, no scoring at all, and matches those methods on long reasoning chains. The scoring machinery most compression papers build isn't earning its keep. The lever turns out to be how much cache you keep, not which tokens you pick, which is a much cheaper thing to ship than a learned eviction policy. arxiv.org/abs/2609.03430
1
46
No answer key, no progress. RLVR works for math and code because you can check the output. Reasoning models get better because there's an actual right answer to grade against. That falls apart the moment tasks get open-ended, which is most of what "agentic" work actually looks like. This paper's whole premise is training past that wall without human graders. But if the reward signal is model-generated instead of verified, errors get reinforced instead of caught. Scaling reasoning without a ground truth just scales confident wrongness faster. arxiv.org/abs/2608.31075
28
Game engines as reward signal. Code models got good because compilers give a hard pass/fail on every attempt, so RL has something real to push against. World models training on scraped video don't get that: nothing tells the model whether the physics it generated was actually right. This paper's fix is to have agents build and play games, where the engine itself checks whether an action was valid, a collision happened, a rule got followed. That turns "does this look plausible" into "did this pass." The bet is that spatial reasoning scales the way code did once you can verify it instead of eyeballing it. arxiv.org/abs/2608.25518
54
Replying to @dhh
My bad bro, this is me rn
9
645
Prefill cost is why long-context feels slow. Every token in a long prompt has to attend to every other token before the model writes a single word back. That's the quadratic wall: double your context, quadruple the compute, and the user just sits there waiting on the first token. FlashPrefill V2 skips most of that math by figuring out which attention blocks actually matter and throwing away the rest, moving past their earlier version's prototype-stage thresholding into something built for production serving. If it holds up, the payoff isn't better answers, it's long documents and codebases loading fast enough that context stops feeling like a tax. arxiv.org/abs/2608.19758
2
1
77
Benchmark noise beats the signal. Base models reasoning in Greek: 0 out of 1,000. Fine-tune them and accuracy barely moves, but that's because the benchmark can't be trusted at this scale. Changing only the random seed swings the score by 7.7 points, more than any training effect they measured. The models do start reasoning in Greek after tuning. That shift never shows up in the accuracy number, which means anyone judging low-resource-language tuning by benchmark delta alone is reading noise as signal. arxiv.org/abs/2608.17744
32
Bugs in science code are different. A coding agent that patches a web app just needs the tests to pass. Patch scientific software wrong and the tests can pass while the actual numbers coming out are quietly false, and those numbers might be sitting in a published paper's results section. 119 tasks isn't a lot, but it's the right kind of small: repo-level fixes where correctness means matching known scientific output, not just green CI. arxiv.org/abs/2608.19799
1
36
Interpretability can't keep up with capability. Models are getting better faster than anyone can explain why, and the people trying to reverse-engineer them are still doing it by hand, one circuit at a time. Mechanist tries to automate that search with AI itself. Which means the tool auditing the black box is another black box. Not a flaw in the idea, just the bet being made. arxiv.org/abs/2608.12036
1
1
72
The benchmark was broken. 60% of the "unsolved" SWE-bench Verified problems fail because of bad tests, not bad agents. Tests that reject correct fixes, or check for requirements nobody stated. Some models were just memorizing gold patches from training data instead of solving anything. That means every leaderboard built on top of it has been measuring test-writing quality as much as coding skill. ProMax resetting the field with multilingual, large-scale refactoring tasks is really an admission that the old scoreboard was lying to everyone who cited it. arxiv.org/abs/2608.09802
1
1
68
Encrypted reasoning isn't private reasoning. Providers hide chain-of-thought by encrypting it and handing it back to you as an opaque blob you just echo in your next request. That design assumes the blob only makes sense in the session it came from. This paper swaps those blobs across sessions and they still work. The encryption hides the content but never binds it to the session that produced it, so the "protected" reasoning turns out to be portable. arxiv.org/abs/2608.09867
47
Vision models read EXIF, not pixels. Turns out image classifiers pick up on camera metadata baked into the pixels themselves, things like compression artifacts or acquisition settings, not just objects and textures. A model trained on ImageNet or LAION-scale captions can learn to associate "this camera type" with "this label" without ever learning what the object looks like. That's a quiet failure mode for anything deployed outside the training distribution. A medical imaging model secretly keying off scanner metadata, or a satellite model keying off sensor fingerprints, looks accurate in benchmarks and falls apart the moment the hardware changes. arxiv.org/abs/2608.05424
1
1
46
BM25 is king. Some could argue the fancy retrieval stack wins once corpora get messy, but this study says the opposite happens as scale grows. Across a 450x range of corpus sizes, plain BM25 keyword search holds its own against dense embeddings and agentic retrieval, with the same reader model and fixed adversarial docs controlling for the usual benchmark noise. That's a real cost story, not just an academic footnote. Every team paying for vector DBs and embedding calls at scale should be asking whether they're buying accuracy or just buying complexity. arxiv.org/abs/2607.26497
27
Can AI agents actually conduct open-ended AI research? Some could argue peer review is the wrong yardstick here, but the paper agrees, that's why it built a new one. Blind review is slow, noisy, and reviewers disagree with each other constantly, so grading agents against it just imports human inconsistency into the benchmark. What they test instead is whether an agent can hold an open-ended research direction across a full project, not just clear a verifiable subtask with a clean pass/fail. That's the harder question behind "AI automates AI research": not whether it can code, whether it can decide what to try next. arxiv.org/abs/2607.27191
1
2
38
16 of 896 experts active per token on a 2.8T parameter model, only 104B actually firing. That ratio meansKimi K3 aims for frontier-level output while paying compute for a model 27x smaller than its total weight class. That's the economics DeepSeek's MoE approach put on the map last year, pushed further. Open weights at this sparsity level lower the cost of running near-frontier intelligence for anyone without a frontier-scale cluster. arxiv.org/abs/2607.24653
1
41
Skill Self-Play routes around the self-play dilemma by making the skill itself the verifiable unit: instead of trusting an open-ended generator's reward signal, the model proposes a skill, applies it, and gets checked against whether that skill actually transfers to new tasks. Verification stays local even as the task space grows. That's the same trick behind AlphaZero-style curricula, except the "board" here is unbounded language tasks instead of a fixed game. The real test isn't the benchmark bump, it's whether skills learned this way generalize outside the paper's task suite or just get better at satisfying their own verifier. arxiv.org/abs/2607.22529
1
2
34
Most agent failures aren't reasoning failures, they're context failures: histories, tool outputs, and prompts pile up until the model drowns in its own transcript. The paper's fix is to treat context as a lifecycle, deciding what gets summarized, evicted, or promoted back in, instead of just storing and retrieving everything. That reframes the cost problem too. Token spend growing every turn isn't overhead, it's a session length limit. Every agent running unbounded context is quietly betting the task finishes before the bill or the coherence does. arxiv.org/abs/2607.21503
1
3
31
Microsoft research just confirmed something really important for anyone working with agents: Intent matters a shit ton. LLMs treat the first message in a conversation as the ground truth and everything after as noise to reconcile back to it, instead of a signal that the task changed. That's backwards for agent work, where the spec evolves as the user learns what's possible. Benchmarks reward getting turn one right. Production usage rewards not anchoring to it. arxiv.org/abs/2607.20734
131
Replying to @Mayhem4Markets

ALT Revenge Of The Sith Power GIF by Star Wars

1
50
Everything you know about world model prediction might be wrong. BadWAM shows the failure mode nobody wanted: a world-action model can predict the correct future perfectly and still execute the wrong action to get there. The world prediction and the action head decouple under adversarial pressure, so the "imagined future" check people cite as a safety property never actually fires. That undercuts the main pitch for WAMs in safety-critical robotics. If the imagined rollout can't catch a bad action, the interpretability argument for coupling over plain policy models loses its evidence. Expect the fix to target adversarial training on the action head specifically. More data on the world model won't touch a decoupling problem that lives entirely downstream of it. arxiv.org/abs/2607.15207
1
2
108
IG-Bench treats a paper like a genome: mechanism, limitation, and recombination as separable traits an AI has to trace back through prior work, not just paraphrase. That's a much harder test than "does this idea sound novel." Most idea-generation benchmarks reward surface novelty, which LLMs already game by rephrasing known mechanisms with new vocabulary. Lineage-grounding closes that loophole by forcing the model to name what it inherited and what it fixed. The real stake is patent and grant review, where "prior art" checking is exactly this task done badly by humans at scale. arxiv.org/abs/2607.08758
1
43
Replying to @Vision33X
LITERALLY

ALT I Agree You Know It GIF by CBS

1
10

ALT Denzel Washington Compliment GIF by slicedbread

1
5
Replying to @Polymarket
$NVDA longs rn
126
Replying to @vboykis
It's like when a magician palms the gimmick while you stare at the silk. The labs hide one of the two hardest jobs in software inside a json field and a model card.
1
520
Replying to @nikitabier
Equity? Nah Snacks though…
13
I guess insider trading isn’t really a big deal, right guys? If the GOP can have an automated bot, can Dems get it together and get equal representation?
1
36
Replying to @mtp_le @Polymarket
Me looking at Claude after I told it “no mistakes”
24

ALT Snoop Dogg Agree GIF by Solo Stove

10
Replying to @gdb
We're LITERALLY so close to Star Trek tricorders.
59
This is a really insightful clip. The whole episode is worth your time. One thing here that stuck out to me: if you're wondering how to land a software engineering job in 2026, Tanay Kothari is telling you how. "I care a lot more about what their prompt to Claude looks like than the actual code that they produce with AI. Because that tells me how they think." The prompt is the artifact that exposes how someone reasons. The code you ship with AI is downstream of that reasoning. First principles, product taste, judgment under uncertainty are the parts AI cannot replace, and the prompt is the cleanest way to see them.
"Nobody has actually lost their job to AI. They lost it due to the promise of AI." - @tankots Meta cut 8,000. Intuit cut 3,000. Anthropic is paying $1.25B a month for compute. China is winning the open-source war. And the Pope just called for AI to be disarmed. This week we covered these topics and more with: - Erik Bernhardsson (@bernhardsson) is running the GPU cloud powering it all at @modal. - Tanay Kothari (@tankots) is running $100M in marketing with two humans and a swarm of agents at @WisprFlow. - Richard Socher (@RichardSocher) is building recursive self-improving superintelligence at @Recursive_SI and the future of AI search at @youdotcom. They join @jason on This Week in AI Episode 15, out now.
2
75

ALT Please Disperse GIF

2

ALT Jeremy Allen GIF

1
32
Replying to @garrytan
POV my fleet of agents all built using /office-hours and /plan-eng-review.
1
271
Dario collecting Karpathy and other AI leaders like

ALT Chalk Op GIF by The Undroppables

Personal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resume my work on it in time.
3
190