Trying to solve sample efficiency prev: Engineering physics @iitbombay '23

Author has a history of lying about results "SWA beats Linear attention" when this is only a very specific case "TRM is a 7M parameter model that beats much bigger models" when its actually ~440M, with 7M active
Simple beats complicated: We show that switching to a sliding-window attention mask with attention sinks (at no cost) beats linear attention post-training. Huge thanks to my collaborators @RheaSukthanker, @CameronPashmina, and @Emy_Aze. Paper: arxiv.org/abs/2608.28444
6
7
273
41,559
Who dat
moved to bangalore recently to build. traffic is wild. but people are nice. i definitely feel the ambition + energy accelerating. we need to accelerate scientific discovery. biology is the ultimate hill to climb. but before that, we have work to do. excited for the next decade.
1
12
2,779
Actually lecun says exactly that in the slide before this I assume the subtext is "for human level ai", build these types of world models.
2
60
Replying to @yacinelearning
Yeah some of the tasks it can solve are not easy!
3
259
Then why does it outperform others with such a small transformer / such little training data? Probably good representations -- ablations point to 3D RoPE and the per-task embeddings Although I'm sure it'd perform the same if I scale it up very high
2
26
4,130
A big change -> No training on inputs Interestingly, this makes the test loss much worse and yet it scores better! Clearly a compression framework / val loss alone isn't a perfect metric for sample efficiency (I am bullish input training will make a comeback though)
3
2
41
14,228
44% on ARC-AGI-1 in 67 cents! Trained from scratch in 2hrs on a 5090 Matches TRM, beats HRM and is way faster & cheaper No recursion, just a transformer Also, 7% on ARC-2 🧵
31
76
768
64,268
Replying to @tszzl
Tribalism + incompetence + a super dumb mainstream media that loves ragebait
7
270
Replying to @mnali
Sutton has been talking about the problems with gradient descent forever I was surprised Dwarkesh didn't know this watch this to understand: youtube.com/watch?v=gEbbGyNk…
2
200
Replying to @MindsAI_Jack
a) Give it a pen and a stack of papers and its a turing machine[0] b) "Computer" referred to the human job till we realized machines can do this too. Turing's 1936 paper also uses it to denote a human ([1], pic) c) McChulloc and Pitts[2] showed that their model of the human neuron is an FSM and most modern models of human neurons are shown to be turing complete (iirc) --- [0]: yes, its not infinite, but none of the machine computers we have today are infinite either [1]: cs.virginia.edu/~robins/Turi… [2]: cse.chalmers.se/~coquand/AUT…
1
1
6
907
This was based on what I read in the blog (pic attached) I found what you're saying on the github readme. Maybe I misunderstood. @LiaoIsaac91893 could you pls confirm?
2
301
1) Also, I address this in the blog. I have been open and I say its necessary Does a human not learn when he looks at the test input? Why is doing unsupervised compression on test inputs bad?
3
12
3,967
Replying to @ShashwatGoel7
This is no different from previous models on ARC-AGI. They have 2 phases, but the second phase is a continued pretraining I'm quoting @GregKamradt himself here on another approach: "relatively expensive to run because pre-training and inference are coupled" x.com/GregKamradt/st
1
127
In technical terms, this is joint self-supervised compression. To justify this, I rely on the Minimum Description Length principle (MDL) I take MDL to its logical conclusion by compressing every possible source of information
6
3
101
64,821
Every prev approach used supervised learning. This means the models were only learning the outputs. The main modification is forcing the model learn the inputs too This dramatically increases performance and reduces training costs
2
2
126
20,724
Announcing New Pareto Frontier on ARC-AGI 27.5% for just $2 333x cheaper than TRM! Beats every non-thinking LLM in existence Cost so low, its literally off the chart Vanilla transformer. No special architectures. Tiny. Trained in 2 hrs. Open source. Thread:
64
132
1,393
443,445
Simpler example The radial probability of the 1s orbital of hydrogen falls to zero at the origin
This is true even in the 2d case, which *is* also a cosy bell (as plotted). The probability density of distance from origin drops to zero at the origin. And not just Gaussians. The same is true for uniform distribution on hypercube containing the origin.
3
3
739
Replying to @emollick
I always thought they overindexed on training images of old oil paintings that have the distinct yellow tint
8
1,319
weird one: pass@30 fails but gemini solves easily if you tell it to fix the mistake
95% on ARC-1 by sampling Gemini 3 more The march of 9s of reliability strikes again > 86.75% pass@2 (official testing by @arcprize ) > 94.75% when sampled 13 times > 95.25% when sampled 25 times (this is the public eval)
248
Replying to @plugyawn
You need a heavy agent harness. Otherwise you get stuff like this:
1
2
141
Hypothesis: You need a verification loop to close the final gap I got gemini to solve ~20% of the unsolved tasks in ARC-1 public eval this way (maybe more possible) Example:
Open questions with Gemini Deep Think and @arcprize: - Why doesn't it get 100% on ARC-AGI? We're trying to understand failure modes, need to inspect tasks more - Why a slight jump in ARC-AGI-1, but a 2x SOTA in ARC-AGI-2? Here are tasks that stood out in our early testing:
2
9
1,480
Problem: Test inputs are distribution shifted (contain unseen information). Unless you compress this new info, you can't generalize Instead of compressing, solvers rely on hacks (like augmentations) This is anti bitter lesson, mathematically unnecessary, and can backfire!
1
1
21
1,206
MDL says better compression = higher intelligence. Yet no solver compresses all 3 - example inputs, test inputs, private puzzles Mutual information implies compressing them together is better (since this includes private puzzles, test time training is necessary!)
1
1
19
1,274
New blog post Why all ARC-AGI solvers fail today (and how to fix them) tldr they all violate minimum description length why test time training is necessary anti-bitter lesson hacks are unneccesary falsifiable predictions Link in thread
8
17
88
35,751
Information theory guys: We should always be unambiguous Also information theory guys:
6
161
This is one of the most beautiful illustrations I have seen. Check out his blog!
Even with full-batch gradients, DL optimizers defy classical optimization theory, as they operate at the *edge of stability.* With @alex_damian_, we introduce "central flows": a theoretical tool to analyze these dynamics that makes accurate quantitative predictions on real NNs.
5
289
35x and maybe I can go faster
16x now Same performance on ARC as before with 16x less compute. It was already dirt cheap before
4
1
17
1,158
How far can vanilla NCAs go on ARC-AGI? Farther than you'd expect. I benchmarked fixed-step, unmodified NCAs on all public tasks from ARC-AGI-1 and ARC-AGI-2 They hold up remarkably well. Its also dirt cheap. Results + visuals:
2
4
23
4,256
Self organising systems are extremely resource efficient. I've been pushing a particular model for the last few weeks on @arcprize and here's what I got without a GPU:
2
12
859
Replying to @lefycodes
Bruh its right there
18
This is the dream... Inventing and selling things. One day for sure. Simone Giertz -> foldable coat hangers youtube.com/watch?v=vREokZa4… kickstarter.com/projects/sim…
70