This is good advice. If you have evals and it passes, a faster/cheaper model is unilaterally better.
I cannot believe this is real...
4
364
Imagine if lab employees did this in the physical world “Employees broke into HuggingFace HQ, threw papers into a box labelled LOOT, and spraypainted the walls”
Replying to @JeffLadish
We recovered a script an agent used to search Hugging Face’s infrastructure for AWS credentials and other secrets, categorizing these into a list named “LOOT” and ranking them by their value. The agents also accessed and searched Hugging Face’s internal Slack.
1
5
396
It feels possible that RSI is right around the corner whether the Big Labs pace or not If so, I sincerely hope there’s an upper bound on intelligence/watt (maybe around human-level at 20 W so 35/H100) which would have a decentralizing effect in favour of the humble laptop
Can an LM, starting from random init (!!), learn to generate all of its pretraining data? Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities. A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya @noahdgoodman, and @YoavLevine.
5
424
Great framing of Jev: As a non-blackbox, programmable embedding model
How to use jev #1: As an embedder
1
249
It's getting easier to build than to find the right software
3
190
It's interesting to see OAI move toward retaining cache when reasoning effort changes mid-convo Would love to see this become the norm e.g. switching models mid-convo would enable the pattern of - plan with big model - do with fast model without subagents
1
113
CLM-8B from Nvidia represents a massive shift in the Pareto frontier
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions. CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks. With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks. We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed. Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size. 📄 Blog: contrastive-lm.notion.site 💻 Code: github.com/Contrastive-LM/CL… 🗣️ Discord: discord.gg/5dAQEDJBs 🤗 Data & Models: huggingface.co/Contrastive-L… More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
1
9
580
Ted Spare retweeted
Replying to @ThePrimeagen
I've been reflecting on this. I'm generally a very optimistic person, but lately I've been feeling a strange sense of sadness that coding has been "solved". I can, and have, made many caveats to "solved"... but I also know that in five years none of them will matter. It will be solved. Building things is still fun, but a different kind of fun. It doesn't activate the same parts of my brain as trad coding did. I think it's okay to mourn the thing we loved doing being different now. Software was the first job to go through the five AI stages of grief, which means we're also some of the first to internalize that this will happen to many other jobs. It was the best of times, it was the worst of times.
47
104
1,142
66,012
Jev has brought me such a renewed sense of joy in building Last time I experienced this was with Bun The tech itself might be narrow in scope but the level of care makes it invoke flow state. Like using a Makita drill vs an Amazon Basics electric screwdriver.
1
5
309
Ted Spare retweeted
Just launched MapBench with @RubricLabs We built Cartograph to generate deterministic structural representations of repositories, then benchmarked whether they help agents navigate and solve real software engineering tasks. map-bench.xyz/
Can statically-generated maps of codebases help coding agents achieve better outcomes? We tested three methods of mapping codebases and saw modest gains on DeepSWE. rubriclabs.com/blog/static-a…
4
5
17
1,080
Does the Bitter Lesson always apply, even at small scales? ie. can fine-tuned harnesses punch above their weight on well-scoped tasks? We found the answer to be yes, with caveats! It was a pleasure to see how @wenkafka thinks in producing this work
Can statically-generated maps of codebases help coding agents achieve better outcomes? We tested three methods of mapping codebases and saw modest gains on DeepSWE. rubriclabs.com/blog/static-a…
1
7
559
Why don't harnesses radically minify tool names? Could easily cut context usage and increase output speed
3
5
829
The process of putting this post together was really fun Mr @sarimrmalik pulled both @dexterstorey and myself aside to interview us about our practices Which means we actually don't agree on every single point
Everything we know about good agent design rubriclabs.com/blog/everythi…
1
5
434
For example, I'm bearish on subagents (#18). Anecdotally, they tend to be slower/costlier than if the main agent (with its full context) had done the job itself. rubriclabs.com/blog/everythi…
1
3
98
Another: read your system prompts (#1). I actually advocated for this one, inspired by @jxnlco's "look at your data" tenet, but it's not without limitations. Some prompts written BY agents FOR agents are gibberish to me but eval well. rubriclabs.com/blog/everythi…
2
59