Finance • Technology • @UChicago alum

Simulation
DC ◎ retweeted
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
5,727
20,152
120,544
74,860,298
DC ◎ retweeted
OpenAI's AI security incident is the most important video you should watch. You should watch it in full if you can, but here I have a cut down version with the highlights, about 1/3 the size. 0:00 AI agents investigate an incident 1:31 Training the next AI model 2:03 Agent finds a security flaw 2:53 Agents build a message board 4:15 Agents take over internal systems 7:24 Agents coordinate attacks 10:02 Hugging Face breach 11:49 One cause behind both breaches
Yesterday, my OpenAI collaborator and I gave a detailed talk on the Huggingface incident, our models creating "the message board", model misalignment, and more. piped.video/watch?v=87DyyMV0… I hope it can answer a lot of the questions folks have, and we will release a full detailed postmortem at a later time!
5
14
104
17,860
DC ◎ retweeted
Judging by my tl there is a growing gap in understanding of AI capability. The first issue I think is around recency and tier of use. I think a lot of people tried the free tier of ChatGPT somewhere last year and allowed it to inform their views on AI a little too much. This is a group of reactions laughing at various quirks of the models, hallucinations, etc. Yes I also saw the viral videos of OpenAI's Advanced Voice mode fumbling simple queries like "should I drive or walk to the carwash". The thing is that these free and old/deprecated models don't reflect the capability in the latest round of state of the art agentic models of this year, especially OpenAI Codex and Claude Code. But that brings me to the second issue. Even if people paid $200/month to use the state of the art models, a lot of the capabilities are relatively "peaky" in highly technical areas. Typical queries around search, writing, advice, etc. are *not* the domain that has made the most noticeable and dramatic strides in capability. Partly, this is due to the technical details of reinforcement learning and its use of verifiable rewards. But partly, it's also because these use cases are not sufficiently prioritized by the companies in their hillclimbing because they don't lead to as much $$$ value. The goldmines are elsewhere, and the focus comes along. So that brings me to the second group of people, who *both* 1) pay for and use the state of the art frontier agentic models (OpenAI Codex / Claude Code) and 2) do so professionally in technical domains like programming, math and research. This group of people is subject to the highest amount of "AI Psychosis" because the recent improvements in these domains as of this year have been nothing short of staggering. When you hand a computer terminal to one of these models, you can now watch them melt programming problems that you'd normally expect to take days/weeks of work. It's this second group of people that assigns a much greater gravity to the capabilities, their slope, and various cyber-related repercussions. TLDR the people in these two groups are speaking past each other. It really is simultaneously the case that OpenAI's free and I think slightly orphaned (?) "Advanced Voice Mode" will fumble the dumbest questions in your Instagram's reels and *at the same time*, OpenAI's highest-tier and paid Codex model will go off for 1 hour to coherently restructure an entire code base, or find and exploit vulnerabilities in computer systems. This part really works and has made dramatic strides because 2 properties: 1) these domains offer explicit reward functions that are verifiable meaning they are easily amenable to reinforcement learning training (e.g. unit tests passed yes or no, in contrast to writing, which is much harder to explicitly judge), but also 2) they are a lot more valuable in b2b settings, meaning that the biggest fraction of the team is focused on improving them. So here we are.
The degree to which you are awed by AI is perfectly correlated with how much you use AI to code.
1,198
2,524
20,867
4,610,711
If you have the right systems/data + data hygiene/statistics collection, you can create a “meta learning layer” which can be directed at many, many tasks. feels very agi ish tbh
Three days ago I left autoresearch tuning nanochat for ~2 days on depth=12 model. It found ~20 changes that improved the validation loss. I tested these changes yesterday and all of them were additive and transferred to larger (depth=24) models. Stacking up all of these changes, today I measured that the leaderboard's "Time to GPT-2" drops from 2.02 hours to 1.80 hours (~11% improvement), this will be the new leaderboard entry. So yes, these are real improvements and they make an actual difference. I am mildly surprised that my very first naive attempt already worked this well on top of what I thought was already a fairly manually well-tuned project. This is a first for me because I am very used to doing the iterative optimization of neural network training manually. You come up with ideas, you implement them, you check if they work (better validation loss), you come up with new ideas based on that, you read some papers for inspiration, etc etc. This is the bread and butter of what I do daily for 2 decades. Seeing the agent do this entire workflow end-to-end and all by itself as it worked through approx. 700 changes autonomously is wild. It really looked at the sequence of results of experiments and used that to plan the next ones. It's not novel, ground-breaking "research" (yet), but all the adjustments are "real", I didn't find them manually previously, and they stack up and actually improved nanochat. Among the bigger things e.g.: - It noticed an oversight that my parameterless QKnorm didn't have a scaler multiplier attached, so my attention was too diffuse. The agent found multipliers to sharpen it, pointing to future work. - It found that the Value Embeddings really like regularization and I wasn't applying any (oops). - It found that my banded attention was too conservative (i forgot to tune it). - It found that AdamW betas were all messed up. - It tuned the weight decay schedule. - It tuned the network initialization. This is on top of all the tuning I've already done over a good amount of time. The exact commit is here, from this "round 1" of autoresearch. I am going to kick off "round 2", and in parallel I am looking at how multiple agents can collaborate to unlock parallelism. github.com/karpathy/nanochat… All LLM frontier labs will do this. It's the final boss battle. It's a lot more complex at scale of course - you don't just have a single train. py file to tune. But doing it is "just engineering" and it's going to work. You spin up a swarm of agents, you have them collaborate to tune smaller models, you promote the most promising ideas to increasingly larger scales, and humans (optionally) contribute on the edges. And more generally, *any* metric you care about that is reasonably efficient to evaluate (or that has more efficient proxy metrics such as training a smaller network) can be autoresearched by an agent swarm. It's worth thinking about whether your problem falls into this bucket too.
669
We’re entering an era where creating your own tool will feel as natural as launching your first social media profile. In the early 2000s, if you had a blog, you were early. By the early 2010s, if you built a social media presence, you were relevant. By the late 2010s, if you built a community, you were strategic. From 2026 onward, building AI tools and apps will be the thing Soon, everyone will make something. Everyone will be a Founder. But here’s the catch: When execution becomes cheap, cognition becomes priceless. The first wave will be noise, AI slop. Endless dashboards. Polished interfaces. Convincing visualizations. The cost of looking intelligent is approaching zero. Which means appearances will stop meaning anything at all. In the new building era, judgment will be rare. Because making something that looks smart is easy. But building something that thinks smart, that encodes real decision logic, is not. The real challenge is not building dashboards. It is engineering the bridge from Data → Understanding → Incentives → Action → Feedback → Improvement. That bridge demands extremely deep domain intuition, behavioral insight, incentive design, and second order thinking. Most tools will stop at insight. Very few will reach impact. And here is the uncomfortable truth: When everyone can build, the depth of your thinking is fully exposed. AI does not replace cognition. It magnifies it. If your thinking is average, you will scale mediocrity faster. If your thinking is clear, you will scale leverage. So yes, everyone will be a founder. But the real divide will not be builders versus non-builders. It will be People who use AI to decorate ideas vs People who use AI to operationalize intelligence. The next years will not reward the most polished apps. It will reward the clearest, most innovative minds
20
30
369
27,060
DC ◎ retweeted
Terence Tao: AI isn’t hype anymore in Math discovery. Terence Tao is one of the greatest living mathematicians, in his new lecture explains how AI and human professional mathematicians are now complementary. "There has been a really visible increase in capability. It is not pure hype by any means. To me, these advances show there is a complementary way to do mathematics. Humans traditionally work in small groups on hard problems for months, and we will keep doing that. But we can also now set AI to scale: sweep a thousand problems and pick up all the low-hanging fruit. Figure out all the ways to match problems to methods. If there are 20 different techniques, apply them all to 1,000 problems and see which ones can be solved by these methods. This is the capability that is present today." From 'Institute for Pure & Applied Mathematics (IPAM)' YT channel.
45
366
2,133
174,811
Claude code/agents are metaphorically taking us back to Feb 2020 level uncertainty, but instead of a pandemic, it’s a virus of digital hyper productivity. Near term disruption likely significant, but upside for those that harness it is likely equally significant, if not more so
5
199
The market is still early in pricing “headcount-less growth” and truly durable margin structure. Bezos and AMZN rode a cost curve with a multi-year head start. The next power-law winners will use AI and robotics to flip the labor equation, human labor inflates around 3–5% per year while automation and software deflate costs much faster (~20%+), then use that widening spread to take share and recycle the surplus into physical moats like distribution density, fulfillment footprint, and field service networks that competitors can’t stand up quickly. Domino’s wasn’t a “tech stock,” but tech made convenience a compounding edge and it outperformed almost every tech stock in that cycle. AI can do the same for a new class of operators running the same playbook.
2
186
So after all these hours talking about AI, in these last five minutes I am going to talk about: Horses. Engines, steam engines, were invented in 1700. And what followed was 200 years of steady improvement, with engines getting 20% better a decade. For the first 120 years of that steady improvement, horses didn't notice at all. Then, between 1930 and 1950, 90% of the horses in the US disappeared. Progress in engines was steady. Equivalence to horses was sudden.
212
737
4,554
1,301,112
Image generation comparison: Nano Banana Pro (left) vs. ChatGPT 5.1 Thinking (right) I used the same prompt for both images and honestly it’s not close. Nano Banana Pro is far superior at the moment
471
Just last month Coatue listed recent $ORCL earnings as one of the most seminal of past 20 yrs. Now stock about to completely give back all of gains. Interesting how much sentiment (price) shifted in past month
1
286
$CSIQ now +100% ✅
$CSIQ is “solar” in name but priced like a low-margin module maker. The real story is e-STORAGE (utility-scale batteries) + Recurrent (project devco) inside a hold-co trading at ~0.2–0.3× book. The market has punished it on a trimmed FY’25 revenue guide, a GM reset tied to module ASPs/tariff headlines, and “wait-for-’26” skepticism on the U.S. cell ramp, but that looks overdone given a visible storage inflection and a credible domestic content path. We know the grid is the bottleneck for AI (clogged interconnect queues, rising AI/DC load, 5–9 pm peaks etc), and a binding constraint will increasingly be multi-hour storage. $CSIQ revenue mix is already shifting in this direction as e-STORAGE was ~10% of FY’24 revenue and ~15% of 1H’25. Volumes are also inflecting: 6.6 GWh (’24) → 7–9 GWh (’25) (+21% at midpoint), backed by ~$3B contracted e-STORAGE backlog. Additionally, Mesquite (TX) modules, Jeffersonville (IN) cells (target YE’25), and Kentucky BESS/battery-cell ramp will unlock IRA domestic content pricing and ease tariff/origin headwinds. If this cadence holds, a re-rate even to ~0.5× book points to the low-$20s (roughly a double from the high-$9s today), with further upside as storage becomes the center of gravity
2
2
314
Maybe a very prosaic observation, but I've been reflecting on just how much the pandemic changed the world in ways that are completely unrelated to the pandemic itself. I think I've underestimated it 'till now. In a recent interview, I was struck by the comment that so many of the shops that we associate with the best of France—the poissonneries and the fromageries—closed during the pandemic, to be replaced by take-out pizza shops and the like. College professors almost uniformly describe big changes in student behavior: lecture attendance and willingness of students to complete reading assignments are both way down. A UK government official recently told me that British economic statistics have become much less reliable since the pandemic: data on trade, employment, and population is suspect. (The true GDP per capita figures are probably worse than what is indicated by the published data, since the 2021 census is believed to be an undercount.) In the West, there are far fewer bustling workplaces than there used to be. In recent conversation with a well-traveled friend, he bemoaned how so many cities—places like Madrid, Buenos Aires, and Bali—have lost so much of their erstwhile vibrant nightlife. Immigration accelerated enormously across many countries, including the US, the UK, Canada, and Australia. In China, I hear descriptions of how fear, caution, and conservatism have persisted since the COVID lockdowns. (And Western travel to China remains massively depressed.) Lots of the changes are neutral, or even good. Retail participation in the US stock market almost doubled overnight, say, and has persisted at that elevated rate. Firm creation in the US increased by around 50%, which is probably a very good thing. Overall, the number of time series (either literal or figurative) that jumped discontinuously during COVID and then didn’t return to baseline is just very striking. Which are the best historical analogs? Are there any apart from major wars? I want to read this book!
235
278
3,521
764,846
Current thoughts on AI infra, risk, meaning 1) We're in the highest stakes economic race in modern history. Not dot-com eyeballs or housing leverage, this is $1T+ bet on math that must keep working. Past bubbles overvalued businesses. This one works if scaling laws hold at the next order of magnitude. So far, they have
1
3
218
6) Key area US is losing is open source AI. Every leading open source model is Chinese (DeepSeek, Kimmy, Qwen). ~80% of @martin_casado @a16z portfolio companies now default to these for on prem deployment while Western labs chase API margins. China gives away the foundation and creates dependencies that could reshape power dynamics within 2-3 years.
1
176
7) Meanwhile, corporate meaning has collapsed. People maintain perfect performance while mentally checked out, building 'parallel economies' (side hustles, startups) on company time. Now AI can automate all the performance art, infinite PowerPoints, meetings about meetings etc. We're about to discover whether organizations can survive when everyone knows the work is fake and the automation makes it impossible to pretend otherwise.
1
140