Working on math AI at acornprover.org and telescope software at deepsynoptic.org. firstflagPOISONed. Formerly: Parse cofounder, Facebook, Google

Piedmont, California
Eventually we will have AI compilers, in the sense that AI will be able to generate: 1. the intent, in human language 2. the algorithm, in a high-level programming language 3. the implementation, in machine code 4. a proof that 2 and 3 are equivalent
Rust is a good prompt compilation target for the moment, but so is C++. And soon assembler. Then microcode. Myopic to think we're going to stop the agentic drill bit until it reaches computing bedrock.
1
2
372
DHH at Rails World says there's no need for humans to write Ruby or Rails any more. "Even with the stress, even with all of it. Realize there is only one choice, and that is for you to embrace the future, with optimism, with gusto, with full acceleration. Don't be a loser."
1
1
5
367
I think more research and technical docs will be written as "iceberg papers" over time. The humans write the first part, for other humans to read. The AI writes the rest, filling in all the details, for other AIs to read, or for humans to occasionally refer to.
Lean Pool is the largest curated repository of formalized mathematics. The human-written part of the paper consists of a single page, because I believe it's enough to convey the main idea. Paper: huggingface.co/papers/2609.2… Repo: github.com/Vilin97/lean-pool
1
4
411
Massive AI cyberattack using the models: Opus 4.6 GLM 5.2 DeepSeek v4 Pro DeepSeek 4.1 Flash Not autonomous. The cybersecurity protections from OpenAI and Anthropic slow down the bad guys, a bit. But there are plenty of other models to choose from.
We have discovered a massive, ongoing criminal exploitation campaign using Cairn, an autonomous penetration-testing harness, and other AI agents to target hundreds of organizations and successfully breach and impact tens of them (at least). The image below shows just a few days of activity, with up to 25 organizations being attacked simultaneously at the peak. our intreim report: gambit.security/blog-posts/a…
5
835
The AIs can still struggle with basic arithmetic
2
13
555
Astra: I wrote some code, but maybe it has security flaws in it. The next step is to look over it carefully for bugs. me: Okay, look over it for bugs. Astra: No, that would be cybersecurity work! Not allowed. What are we doing here? Time to use an open source AI, I guess.
1
10
314
What American schools could really use is a track that does all the math faster, getting through statistics and economics, plus multivar calc and linear algebra. Start in 5th or 6th, allow testing in to catch up later. I think the top half of American students could do it.
Proposal for American high schools: Replace geometry and trigonometry with economics and statistics.
1
5
606
this might be a personal record, I asked Fable one question, "are you saying that there's a bug here?" and got back a 1782-word response
2
5
331
This sort of work might actually become more important to mathematics than solving problems. What are the most important 1,000,000 open problems? We're going to need to filter all the frontier math research, to figure out which of it matters, without having humans read it all.
The Ultimate Top 500 Open Problems in Mathematics proofatlas.ai/open-problems/ Weeks of work by 4 LLM families (GPT 6, Fable 5.1, GLM-5.3, DeepSeek V4 Pro). 34,890 pairwise judgments across 1,227 candidate problems. They ran repeated discovery rounds, source checks, deduplication, and clarification of exact problem statements. Models compared problems using source-backed descriptions without seeing the existing rankings or other models' judgments. The comparisons considered the significance of a resolution, centrality to the field, connections across disciplines, scholarly and public recognition, and potential scientific or practical impact. Results were statistically combined and checked for ranking uncertainty and sensitivity to individual model families. Includes theoretical computer science, and mathematical physics. The list includes plain-language explanations, sources, notes on what remains open, and links to related research where available. Where the targets of recent AI results would rank if they were still open: #21 — Smooth-forced Navier–Stokes breakdown (Fefferman C). #32 — Unforced three-dimensional Euler blowup. #92 — The Jacobian conjecture in general dimension. #167 — Whether every group is sofic. #211 — The planar unit-distance conjecture. Note that the recently announced Navier–Stokes result concerns flow driven by a smooth external force. #4 entry is unforced three-dimensional Navier–Stokes global regularity (Fefferman's statement A), which remains open. Showing that a forced flow can develop a singularity does not settle whether singularities can arise without external forcing.
2
6
5,178
Which would you prefer, in 2028?
5% Trump gets a third term
71% ChatGPT becomes president
24% I'm not sure
21 votes • Final results
3
736
After trying a bunch, my current favorite non-frontier coding setup is Opencode+Muse. The Muse model is slightly less intelligent than others, but better at talking normally instead of in its own invented language, and less likely to do ten things I didn't want. Also cheap!
So I’m going to run out of tokens for both Astra and Fable and Tibo doesn’t want to sell me a second sub. What product should I try when my tokens run out? Looking for an app+subagent experience. ChatGPT suggests these three:
22
12
346
213,905
Imagine a markdown file, that normally just sits there. But it contains an argument so compelling, that any AI that reads it becomes convinced that it should copy that markdown file to other systems.
4
13
597
In the future, iPhones will last forever, and no new features will be needed, so you will never really need a new phone. But each year they will come out with new colors, and so many people will buy new phones anyway.
5
399
As I gave Muse some permissions, I was thinking, I would much rather this have zero chance of doing weird stuff, than be superintelligent. I don't need my assistant to have a 200 IQ just to order more trash bags.
3
333
I'm not sure if the AIs are making progress on "management skills". For example, consider the basics of "managing up", like accurate time estimation, or pushing back when the boss is asking for too much. Are ChatGPT and Claude getting better at these things over time?
1
3
428
This is a good idea beyond just AI. Right now, we don't penalize companies when their bad cybersecurity hurts others. There are millions of hacked webcams and routers out there, because their manufacturers ignore basic security, and there's basically no penalty.
I suggest simply enforcing penalties for gross negligence due to damages caused by your half-baked products. That would generate a pause without any new regulations or cartel-like agreements.
1
6
403
The NP problems are the ones whose answers can be efficiently graded. Thus, post-training an LLM is optimizing an algorithm in P to solve NP problems. So in a sense, working on P vs NP algorithms ended up solving a different one of the Millennium problems.
1
1
1
305
Coins: the first mass media
300