Co-Founder at rekursiv.ai. Prev. @LumaLabsAI (Realtime Video World Models), @GoogleAI (VideoPoet). Let's automate research!

Mountain View, CA
We sent a swarm of AI agents to solve Karpathy’s NanoChat benchmark, and they crushed SoTA in 3 days. We built our own harness on top of a graph database to do auto-autoresearch, each iteration learning from mistakes to do research better. The team wrote >15,000 entries. 🧵
61
135
1,291
312,728
We sent a swarm of AI agents to solve Karpathy’s NanoChat benchmark, and they crushed SoTA in 3 days. We built our own harness on top of a graph database to do auto-autoresearch, each iteration learning from mistakes to do research better. The team wrote >15,000 entries. 🧵
61
135
1,291
312,728
Our guiding question: "What should we test next, and what would make us change our mind?" So we built trackinizer, priml, and sagent to trace the connections between ideas, experiments, and evidence. github.com/rekursiv-ai/track… github.com/rekursiv-ai/priml github.com/rekursiv-ai/sagen…
3
4
55
5,230
Dan Kondratyuk retweeted
Today we are launching  EdotEnv, a Quant Neolab building toward RSI. RSI needs a loop of increasingly difficult tasks, which markets naturally are: Trading well means markets become more efficient, this makes successful trading harder. Reach out if you are interested!
32
17
351
86,720
This is a core problem we found when conducting autonomous research: without a clear benchmark or baseline, it has a hard time making progress. However, we found these issues can be substantially improved with the right harness design. More to come 👀
Can AI agents conduct open-ended AI research? Most evaluations of agents conducting AI research focus on narrow, verifiable tasks. But AI research is often open ended. Researchers pick hypotheses, decide what evidence is appropriate, and recognize a failing approach. We gave agents research questions from two unpublished papers, six days, and thousands of dollars of API credits and compute. The authors of the original papers then reviewed the AI-generated papers. They unambiguously rejected agents' outputs. arxiv.org/pdf/2607.27191 We call these "shadow evaluations", since the agents are shadowing the original research effort by the authors. Agents were fluent at most *engineering* tasks They conducted serious literature reviews, debugged GPU environments, ran hundreds of experiments, and turned in camera-ready LaTeX without human help. We also found no evidence of reward hacking. If anything, we found the opposite: the agents started with marketable claims and walked them back to negative results as the evidence came in. Neither agent output was close to the bar of a top conference paper Both papers suffered from similar failures: poor judgment about the bar for an AI paper submitted to a top conference, the lack of creative problem solving and ineffective backtracking, poor awareness of resources, and instruction drift. 1) Lack of judgment about the bar for a top conference. The agents had a poor model of the bar for an AI paper submitted to a top conference. We allowed agents to self review their papers. Despite the poor paper quality, their reviews predominantly labeled the papers "weak rejects". 2) Lack of creative problem-solving to address feedback. When they received negative reviews, the agents typically narrowed their hypothesis and claims, rather than working out creative ways to address these concerns. 3) Ineffective backtracking. The agents dropped their most ambitious hypotheses within the first fifteen hours of carrying out the experiment and never changed course afterwards. 4) Poor resource awareness. Both runs ended with over half the API budget unspent. One agent declared itself done seven hours before the deadline, right after its own self-reviewer returned another reject. 5) Instruction drift. They did not follow explicit instructions on minimum exploration time, incorporating feedback for reviews, and on paper length (the outputs exceeded the page limits in both cases). This research design has many limitations Limitations include the small sample size, non-blind reviews, and the reviewers knowing that the work was AI-generated. We also couldn't test Anthropic's strongest model, because Fable 5 is deliberately limited on frontier AI research tasks, so ended up using OpenClaw with Opus 4.8 (extra-high) for our main experiments and Codex with Sol 5.6 (ultra) for a robustness check. But we think the research design is still helpful in assessing AI agents' ability to conduct research, and it is complementary to evaluations on verifiable tasks, as well as blinded reviews of AI outputs. Our results show early evidence that even though agents are proficient on verifiable research tasks, they do not make genuine progress on open-ended ones. It is worth understanding if this is a fundamental limit, or if better models, scaffolds, and more compute could help close it. As the evidence for the gap between open-ended and verifiable tasks firms up, it is also worth understanding how much progress in AI depends on open-ended research rather than hill-climbing on well-specified objectives. In follow-up studies, we are expanding the set of non-public papers we evaluate. If you are an AI researcher with unpublished papers, we would love to collaborate with you on our next shadow evaluation. Expression of interest: forms.gle/CEcA4JmYhDWXQGot8 Conducting shadow evaluations involves a lot of researcher degrees of freedom. In many places, our coauthors disagreed with our interpretation of the findings, and we have surfaced those disagreements in the paper. (This is one reason why having a group of coauthors with different priors is important for open-ended research.) We also release the agent logs, one of the AI-generated papers (the other original paper is still not public), and all the code and data, so that others can conduct their own analyses of our results: cruxevals.com/crux/can-ai-ag… Finally, we plan to conduct shadow evaluations regularly, and are hiring a senior researcher to help lead these efforts. Apply here: cruxevals.com/careers/senior… I'm grateful for the core team leading this effort: @PKirgis, Andrew Schwartz, @steverab, and @random_walker, and to our collaborators who reviewed AI papers, analyzed agents logs, and gave feedback on the paper: @DavidDAfrica, @KozzyVoudouris, Viet Nguyen, Toby Pilditch, @DubMagda, @HarryCoppock, @CUdudec, @nityndg, Matilda Orona, @tilmanbayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, @hlntnr, @ghadfield, @sethlazar, @snewmanpv, @shostekofsky, @RishiBommasani
2
15
1,215
Dan Kondratyuk retweeted
At our latest YC Paper Club, researchers and builders presented on multi-GPU kernels, intelligence per watt, heterogeneous inference, and more. Thank you to our presenters: 0:00 – @FrancoisChauba1: The case for chip and kernel specialization 7:16 – @stuart_sul: Parallel Kittens - Systematic and Practical Simplification of Multi-GPU Al Kernels (arxiv.org/abs/2511.13940) 21:29 – @JonSaadFalcon: Intelligence per Watt - Measuring the Intelligence Efficiency of Local and Cloud AI (arxiv.org/abs/2511.07885) 31:05 – @MarkSaroufim: When Al Starts Writing Systems Code 47:04 – Misha Smelyanskiy: Why AI Inference Needs Heterogeneous Hardware 1:04:33 – @shacklettbp: Building a High-Throughput Game Engine that Runs ENTIRELY on the GPU (madrona-engine.github.io/sha…)
23
33
289
66,886
Dan Kondratyuk retweeted
LLMs can't trade & higher reasoning doesn't help. we ran SOTA models for a 2y period. TL;DR: they suck & reasoning doesn't help. - no model comes close to simple static baseline - more reasoning ≠ better trading - when losing money Sol trades less instead of better details 👇
16
22
329
187,758
Reading documentation written by Claude hurts my brain, while GPT doesn't. The cognitive load just seems way too high. Anyone else feel the same?
4
10
888
GPT's thinking traces are kind of adorable.
8
477
Great to see training next-token-prediction on video unlocks emergent understanding. We tried something similar with VideoPoet in 2023 but it failed. Key unlock seems to be higher VAE compression, more semantic latent space, and cleaner data. Well done!
We’re introducing imagination models: a new foundation model architecture that unlocks learning from internet-scale video. Our first imagination model, Photon-1, learned to use a computer by watching 18 years of screen recording video without action labels.
1
3
42
9,080
We raised $5M from Y Combinator to scale AI scientists whose own breakthroughs accelerate the next. In days, they matched top models on ARC-AGI at up to 10,000× lower cost, every gain from methods they invented themselves. How it works below:
7
8
49
5,200
On ARC-AGI-1 it went from 44.9% → 75.5%, from solutions it invented itself. It transferred the recipe to ARC-AGI-2, and on Sudoku invented a new method (Hypothesis-Pinning Search) hitting 100% on Sudoku-Extreme, writing 125k+ lines of code and running 684 experiments.
1
1
201
We're two ML researchers with 21 years of experience (ex-Google, DeepMind, Luma AI, worked on TensorFlow Probability, Gemini, VideoPoet). We're hiring Founding Scientists. And if you've got a hard benchmark you'd love an autonomous team to attack: contact@rekursiv.ai
4
192
Perhaps the marketing behind "Mythos is too powerful" was a little too convincing
The US government, citing national security authorities, has issued an export control directive to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States, including foreign national Anthropic employees. The net effect of this order is that we must abruptly disable Fable 5 and Mythos 5 for all our customers to ensure compliance. Access to all other Claude models is not affected. We apologize for this disruption to our customers. We believe this is a misunderstanding and are working to restore access as soon as possible. Read our full statement: anthropic.com/news/fable-myt…
1
7
362
Dan Kondratyuk retweeted
Thanks all for coming and circling with us!
Diffusion Circle happening today @CVPR #CVPR2026 Upper Lobby F at 3:30pm, come join us!
3
4
30
15,710
This is incredibly cool! An OS that is entirely simulated by a model. This shows glimpse of how world models might look like in the future. Any idea can transform into any app. piped.video/z3pV6FHvcgM
1
2
330
If we push this thought to the extreme, the ultimate form of a world model might resemble something like an operating system. It can simulate any type of computable program, running its own set of programmable neural processes to accurately represent any kind of environment.
4
327