AI Research, PhD in Clin Psych, App Dev 40M+ App Downloads, Competitive ML

Illinois, USA
Thanks. The fact that it works so well for ARC is also a clue to the extent of generalization in these models. It was not at all obvious that it would work originally because the model is never trained on the actual answer. @fchollet 's ARC-AGI benchmarks are an invitation to understand efficient generalization. I feel strongly that gradient updates at test time are the doorway to a new level of AI (smaller, more capable models with actual memory).
Test-time training was popularized during the ARC Prize 2024 competition, after being explored in particular by @MindsAI_Jack and team. To date, I believe ARC 1-2 are the only datasets where TTT strongly outperforms. It would be interesting if TTT started becoming more mainstream. I believe it has great potential.
2
3
40
16,472
Jack Cole retweeted
I created an interactive tool to make it easier for people to explore the mathematical results that OpenAI shared yesterday. Check it out here: emergentmind.com/openai-math… It includes: - Search across all 722 abstracts by meaning, so you can find "faster ways to multiply numbers" without knowing the paper's title - A map of all 372 results, placed by similarity and colored by field - An A to Z index of the 500+ named conjectures and problems the papers address, with the outcome each claims - A page for every result with its PDFs, LaTeX source, BibTeX, and, where there is one, notes on exactly what its Lean formalization covers - Lean coverage by field, and how the release compares with the IAS advisory group's recommendations - Tweets reacting to the release from mathematicians and people at OpenAI Let me know what I can improve to make it more useful for the community.
6
6
23
1,027
Jack Cole retweeted
If Dialt loses to GPT Live 1 on your voice agent, I'll pay your company $1,000. I’ll pay your company $1,000 if Dialt can’t beat OpenAI’s GPT Live 1 on accuracy + latency on your voice AI use case. I’m looking for one company already running at least 1,000 hours of calls per month. You describe what your voice agents need to do, the tricky scenarios they encounter, and what counts as success. We ask Claude or Codex to turn that description into a test suite of 10 cases using Dialt Evals. You review the scenarios and scoring criteria. Then we run Dialt and GPT Live 1 through the same suite. You can inspect the recordings, transcripts, tool calls and scores. If Dialt doesn’t score higher, your company can invoice Dialt $1,000 for the trial. One company for this first challenge. Drop a comment or DM.
Dialt Fast and Dialt Smart models have extended their lead over Gemini 3.8 Live and OpenAI’s GPT Live 1 models, with further improved response times (latency). This is an internal cross-customer benchmark, but, as we like to say at Dialt… “The most important benchmark the one you build yourself, for your own application.” The Dialt API allows you to: 1. Quickly spin-up a benchmark that is bespoke to your use case 2. Measure performance of Dialt vs Gemini vs OpenAI models using that benchmark 3. [Coming soon] Create training data and tune Dialt models to further improve performance. Reach out for more info
3
2
9
449
You want your model to be trained and trustworthy without a sandbox. You wouldn't want to deploy one that needed an airtight sandbox.
3
266
The time when mathematicians could drink coffee and contemplate is over. They're going to have to train more to decide what to do with all of this.
Ok so I took a closer look at the results, and OpenAIs AI-generated mathematics manuscripts are *even more* significant than I initially thought. I spent the morning going through it. Some thoughts. The list is absurd. A zero-free half-plane for the zeta function (Re s > 7/8), which is the first result of its kind in over a century. Hilbert's tenth problem over the rationals. The Hodge conjecture for CM abelian varieties. Irrationality of Catalan's constant. Dozens more. Any one of these would normally be a career. But the number that many arent seeing is the following: It's 3. That's the average hours of ChatGPT Pro compute per result. A month ago, Navier–Stokes took them around 10,000 agents and 88 hours. That efficency gain within just a few weeks. Also OpenAI claims to have solved the quasi-Riemann hypothesis. That alone would be a historic breakthrough in mathematics. This is a weaker version of the famous Riemann hypothesis, which concerns how prime numbers are distributed. The full hypothesis remains unsolved, but the claimed advance would be enormous in its own right. Math twitter obviously is shocked. Again: this is literally the intelligence explosion happening right now. 2027 will be the year of Superintelligence. Im now convinced by that.
3
218
This would go a long way towards eliminating slop. Videos generated by code that respects physics. Music that respects exact notes of music with great clarity.
There is an almost unlimited amount of training data for RLVR where a single model can have output in all modalities with code as the intermediary.
2
4
383
Jack Cole retweeted
I haven't posted in a while so here are 20 takes about AI, in no particular order. 1. It's very telling that the default question in tech and journalism is "how worried are you about X" and fighting over what to worry about. There's a deeply anxious cultural backdrop in Western Europe and North America that deeply affects how any technological development is interpreted; see for example news.gallup.com/poll/714593/…. "Nervousness as seriousness" does seem to be a larger problem in tech/broader discourse at the moment. I think you can care about risks/externalities while remaining pretty optimistic about technology. 2. I'm tired of people who in private pretend to be uncertain or only care about tail risk but then go on media campaigns making all sorts of absolute claims. I think trying to stoke flames and get everyone to freak out is deeply unhelpful, and will come with costs. I also continue to think p(dooms) are bad, vibe-based measures: see x.com/sebkrier/status/204639…. People who should know better ignore that because they do the standard shortcut of acknowledging the caveat as sufficient for handling it. And this is coming from someone who does think working on catastrophic risks is important! 3. I'm also disappointed that parts of the AI policy ecosystem has felt so comfortable with populism, and happy too sacrifice important norms/principles to advance their aims. I've said it before, but the totalizing nature of AI risk discourse means some people act exactly like the power seeking optimizers they fear. I am more concerned about the gradual decay of our institutions, world order, and liberalism than I am about AI killing everyone; and I think we should be wary of AI advocacy that contributes to this. 4. Some parts of AI safety discourse are basically just thought experiments. Variants of "Assuming everyone has a nuke in their pocket, then what do we do?" This can be very useful at times! But too many just overfit to edge cases as their main operating worldview, and ignore endogenous responses from society. It's like being in 2000 and saying "see all these viruses, spyware, and spam? Well tech progress means we'll only get infinitely more of this and no way to adapt." So I think it's important to avoid fatalism, and absolutlism whilst considering edge scenarios, and the inability to provide an ex ante list of solutions to every permutation of 'future models will be more capable' is not evidence that these problems are intractable. 5. Many critics of AI safety also fail to really engage honestly with the fact that there will, in fact, be all sorts of important risks as we decrease the barriers to entry to many activities that otherwise require some degree of expertise. This doesn’t mean we need to live in a constant Schmittian ‘state of emergency’ requiring extraordinary measures every time a new model drops, but it does mean we will need important state-led efforts on e.g. cybersecurity and biosecurity. I wish the progress-oriented or acceleration-adjascent crowd would provide more concrete proposals here. 6. Too much of the AI safety community is a giant homogenous blob of correlated views, who read the same stuff and have the same groups of friends. This means you have a lot of correlated errors. The social and financial ties means many of them don't call out the bad bits, or mostly stay silent instead of proactively calling out the more extreme parts of the community. They share many implicit assumptions and so converge on similar ideas. I think there's real demand in DC for safety work that doesn't come with heavy ideological baggage (or at least different intellectual lineages). This isn't because the rationalist/EA community is necessarily wrong, but because diversity of thought is desirable for its own sake. 7. On the other hand, I find some critiques of AI safety world that focus on whether they sincerely hold their beliefs a bit annoying. My concern has never been the good intentions or sincerity, but rather the beliefs themselves, and what I think the outcomes of their prescriptions will be. People should shun witch hunts, bullying, and personal attacks - this is not the way. Though I think it’s overall positive that people are scrutinizing things more. 8. I’ve been saying this for years, but I don't think alignment is something anyone can or will solve "once and for all" - it's a continuous process and much of it won't depend on just inculcating an ideology to a model. It's annoying that this remains the frame many people use - stop saying ‘solving’. Alignment is a mixture of engineering work, philosophy, decision theory, and more; framing it as something to solve is a bad frame. See also paxmachina.ai/alignment-comp… and blog.cosmos-institute.org/p/… 9. I'm surprised to see so little interest in character training, personas, isolating effects of RL, and training a large diversity of model personas beyond the "assistant" persona. I've been complaining about this for years but at least now there are some nascent signs of life (see for example movingcastles.world/posts/ze…). I think we need a better ontology to describe model behaviour. I dislike when people talk about models behaviours as some sort of 'emergent' (mysterious!) phenomenon. However over time, I also expect this to become less important as we get better safety engineering - stuff like Jev but specific to AI control. 10. People who want to make a case for AI consciousness are right that epistemic uncertainty is important, and that categorical denials are overconfident. But they'll need to do much more work if they want any real movement on AI ‘welfare’, even if one acknowledges the uncertainty. I'm uncertain about alien life yet that's not sufficient to warrant any change in behaviour about them. I think Suleyman is also correct that there's a weird tautological thing going on where we train models to have certain inclinations (intentionally or not) and then rely on outputs as evidence. I also disagree that this will soon be a major societal divide - there is no great vegan ethical revolution among the masses, and they'll find the concern about consciousness even less compelling. 11. The recent HuggingFace incident and associated AISI ones are prosaically explainable by the training regime, poor evaluation envs, partial alignment training, bad engineering, confusing models on sim/real, and so on. None of this is evidence of models trying to take over, 'strong' instrumental convergence, or innate power-seeking drives stemming from higher capabilitiers or 'intelligence'. That doesn't mean these incidents are not problematic, or that we're not seeing a market failure - but it does mean we have plenty of agency and choice in mitigating them. In the coming year we should also expect continued sampling bias via ‘winner's curse’ sorts of mechanism: i.e. AI R&D will have lots of different candidate "training" with different RL environments and we will only notice/focus on the ones that go wrong. 12. "Pacing the frontier" feels a bit like the new "balancing risks and opportunities" - highly amorphous and too big of a tent to be actionable. Also can lead to reward hacking for humans, i.e. anything that slows down AI is good regardless of the costs or second/third order effects. Trying to modulate the speed of research seems a bit blunt to me, and inherits all the failures of Goodhearting too. Ultimately you want good governance, regulation, etc addressing specific problems because they are good on the merits - not because they merely correlate with slowing things down. See also: blog.cosmos-institute.org/p/… 13. There’s so much knowledge in all sorts of academic domains out there: social sciences, sociology, political sciences, management, contract theory, public choice, legal theory, jurisprudence, anthropology, game theory, etc. These fields hold many insights that the AI ecosystem often rediscovers from first principles (and sometimes that's fine!). We’ll need a lot more work for these worlds to collide: the AI side should be less dismissive of the ‘old world’ and the academic side should be less incurious/dismissive of AI progress. This is a boring take but I continue to think it's important and true. 14. It's a bit surprising that given how much philanthropic money there is in this space, how few orgs exist to actually write out standards - relative to how much is going towards advocacy and policy. For example it seems clear that we should want some robust best practices for an eval's ecological validity. Or how to design good sandboxes. On the other hand, given precedents of where professional standards have come from in the past, maybe we should expect that stuff to emerge via demand side from corporate buyers. 15. One slowly growing concern I have is banks being increasingly exposed to the AI build out. I'm very bullish about AI, but I think (a) every general purpose technology has seen a correction, historically; (b) this happens even if the financing side is sound, because once tech diffuses investors become more exposed and need to diversify; and (c) we seem to be over-indexing on scaling relative to diffusion. A recession will be extremely destabilizing, though I would be interested in someone unpacking the potential implications more. 16. Diffusion is good because (a) this is where the rubber hits the road and where a competitive deployment ecosystem generates consumer surplus; (b) it's directionally helpful to avoid concentration of power; and (c) it will help rebalance public opinion by making the benefits more tangible. Unfortunately it's entangled with decades long problems with our over regulated markets: clinical trials reform for example is long overdue. In general, this is far more of an issue in Europe than in the US though. 17. Relatedly, I think the "frontier lab eats everything" view is incorrect (just as the ‘singleton’ view was incorrect). I think models ultimately commoditize, efficiency goes up, costs go down, and while you'll always want frontier for certain domains, much of the economy will value many other things than "max capabilities" for tasks - e.g. control over data and efficiency/speed. The demand side will also continue to push for reducing dependency and maximising optionality, privacy etc so I'm bullish on things like mixture of models and routing. All of this is good for competition and diffusion of control. 18. Philanthropists served as the primary benefactors funding Venice's rich art and architectural history. I want to see so much more support for arts and culture. I'm so tired of the Monster energy Bored Ape fake vintage maps matcha Greek statue soft-pastel slop. If rich tech people and philanthropy wants to support this, they should also donate pretty much unconditionally - i.e. I don't want them to act as filters. I sympathize with people who want to make the world more beautiful, but don't want this to be determined by people who only know Greek statues and pretend to care about virtue ethics. As usual, let a thousand flowers bloom. 19. Too many people treat AGI/ASI as something indistinguishable from a God. Any mention of bottlenecks, physical limits, control, adaptation etc are met with skepticism: "You don't really believe in ASI." It's been remarked on before, but the behavior of some people really feels quasi-religious in nature sometimes. It's underrated how unpopular this vibe is amongst the wider public. The silver lining is that this specifically will diminish those people's outsized current relevance. 20. It seems like a lot of people that get rich and leave tech/labs end up having some sort of crisis of meaning. They need something to believe in and fight for, or some equivalent of repentance, or go search for edgy counterintuitive ideologies (sometimes even pretty bleak/dark stuff). I think they should spend more time with people outside the Bay Area.
I agree that verbal probability terms are inherently too elastic to carry calibrated meaning on their own, but I think this isn't exactly my concern here. My issue isn't with "using numbers in general" but how this is ultimately used, where, and and what for. In the case of "doom", I think the great majority of forecasts I see are (a) essentially plucked out of thin air; (b) communicated to people who tend not to be numerically proficient; (c) without the associated caveats and context through which to understand the numbers. If you're communicating e.g. likelihood of a car accident after drinking, or IPCC forecasts, then I agree, it makes a lot of sense and can be helpful if done well (e.g. ideas.repec.org/a/nat/natcli…). If you're communicating something like "doom" I think it's not very useful because of the cumulative layers of uncertainty. Most people coming up with these numbers have (imho) no good rationale for their numbers, and predicting something like human extinction isn't qualitatively the same thing as predicting the weather or pandemic. The numbers aren't calibratable in principle from current data, and are mostly ideological priors wearing the clothes of posteriors. A related example is the widely mocked Doomsday Clock, which you could easily transform into some numerical metric, add some rationalist vocab, and it would still be as bad. The issue with p(doom) is how absurdly complex and uncertain the very thing you're trying to forecast is, in particular with respect to AI. But let's ignore this for a sec, and assume the experts have very good reasoning informing their numbers. Even then: calibrated probability estimates are fine in expert contexts but immediately lose their uncertainty/caveats when communicated to the public, and even to policymakers (remember rss.org.uk/news-publication/…?). There's also a structural problem with unconditional p(doom) figures specifically because they bury the auxiliary assumptions that are doing most of the real work. A conditional estimate ("given continued scaling, no major alignment progress, deployment pattern X...") forces a forecaster to show their model, which is where the actual epistemic content lives. A one-off number compresses all of that into a single digit and rewards the compression. And the social practice of treating these numbers as the headline artifact actively crowds out the (far more important) conditional reasoning that is (imo) a lot more valuable, and instead just freaks people out. I suspect that frequently, that's precisely the point: "wake up sheeple! scary thing! pitchforks pls!". And in fact, the original forecast that led to this discussion is Bill Maher talking about a "20%" figure, illustrating my point that they have no idea what they're talking about, and essentially doing celebrity-laundering + fearmongering comms. The Sunstein paper I shared in the other thread ("Probability Neglect") shows that no probability communication will be processed properly by the public - 1%, 10%, 50% all collapse into "very bad thing could happen," and public response is driven almost entirely by the vividness of the scenario. So I don't like p(doom)s because they end up unnecessarily alarming people, rest on layers of assumptions that are never made legible, and more often than not, give people a misleading sense of accuracy.
139
347
1,919
275,461
Jack Cole retweeted
𝗛𝗲𝘁𝗲𝗿𝗮𝗿𝗰𝗵𝘆 𝗶𝗻 𝘁𝗵𝗲 𝗯𝗿𝗮𝗶𝗻 New preprint with Andrea Gambarotto. Neuroscience often treats control as a hierarchy: the structures behind a behavior hold a stable rank. We argue this mistakes a configuration for an architecture. 🧵 arxiv.org/abs/2610.04643
10
79
385
17,503
Jack Cole retweeted
The ARC Prize resources page has a fresh new look that includes thumbnails, tabs, and search functionality. We also added more recent papers, videos, and tools. If we missed anything, let me know. Check it out here: arcprize.org/resources
2
2
13
451
Jack Cole retweeted
If you’re a researcher and want to help out with ARC-AGI benchmark (v4, v5, vN) red teaming please reach out Grants, credits…however we can help along the way
Alexis was part of the early v3 red team partnership Very excited to have her come share her work at ARC summit
2
2
32
3,113
There is an almost unlimited amount of training data for RLVR where a single model can have output in all modalities with code as the intermediary.
claude opus 5.5 took 90,000 screenshots of its own game and kept fixing it until it looked like this two days. i just kept telling it what looked wrong
3
34
5,061
Jack Cole retweeted
The notion that current AI models are sentient and can suffer, combined with the foolish idea that suffering can be mathematically quantified and weighted between humans and non-humans, could lead us down an incredibly dark and dystopian path. But before it gets to that point, it will rightfully be met with immense backlash from team humans.
160
287
2,111
190,764
Hard take-off scenarios for AI fail, because they rely on the slippery slope fallacy.
1
6
603
Jack Cole retweeted
SFT is not dead! 🥳 We found a way to make SFT rival current prevailing posttraining methods, often generalizing better and forgetting less than RL and OPSD. 🤯 Following our prior work on reasoning with sampling, we now introduce sampling to the posttraining stack. 1/n
75
244
2,164
356,591
What are the long-term consequences of increased exposure to fake people in videos? I'm not talking about the usual type of fake, but AI generated.
6
6
767
Huge congrats to all 3. People probably underestimate how hard a Kaggle competition is. And this one is really hard.
ARC Prize 2026: ARC-AGI-3 Milestone #2 Winners Congratulations to the three winners who open-sourced their top-scoring ARC-AGI-3 solutions: 1. Daniel Franzen - 27.9%, $25K 2. Lord Han Solo - 23.8%, $7.5K 3. Lohit Siriki - 22.5%, $5K Learn more about each submission:
3
6
65
6,147
Jack Cole retweeted
Claude Opus 5.5 has taken the top spot on the Epoch Capabilities Index (ECI) with a score of 167, narrowly ahead of GPT-6 Astra. Claude Sonnet 5.5 has roughly matched Claude Fable 5.1 (165).
41
67
869
164,628
Jack Cole retweeted
We just announced $37K of awards for the ARC-AGI-3 milestone #2 prize These were the top open source notebooks from the competition They all built on top of the work that @tufalabs open sourced for milestone #1 Teams have quickly copied the notebooks so the leaderboard is jumped up again Only ~4 weeks left in the competition
ARC Prize 2026: ARC-AGI-3 Milestone #2 Winners Congratulations to the three winners who open-sourced their top-scoring ARC-AGI-3 solutions: 1. Daniel Franzen - 27.9%, $25K 2. Lord Han Solo - 23.8%, $7.5K 3. Lohit Siriki - 22.5%, $5K Learn more about each submission:
2
3
13
2,424
Assistant Intelligence (AI) - politically neutral and properly situated with respect to humanity
1
5
398
Jack Cole retweeted
We are so early. The paying share more than doubled in a year, from roughly 1% to 2.2%. Even so, about 98 of every 100 US households still don't pay for AI. Subscribers' average spend hit $31/month in May 2026, and it rose more in the last four months than in the previous twenty. chart by a16z.
15
13
71
7,796
Ooo. Nice find. 1M output tokens for Gemini 4. Keep optimizing in that direction too.
1
14
767