CEO @bespokelabsai. RL/Envs/Posttraining/Agents. Created Generative Retrieval at @GoogleDeepMind.

Inside an RL Env
Introducing Bespoke Nimble: an open data, open model, open recipe for an open Jev. Code and info: github.com/bespokelabsai/nim… Model: huggingface.co/bespokelabs/B… Data: * A new data curation recipe called contrastive data curation. * Slightly change facts to generate negative data. This pushes the model to discriminate better and become a better decision maker. The calibration is implicit. * Didn't do ablations but I think this is a critical piece! * This also means training data doesn't need probabilities. * Data covered 10 categories, and is fully synthetic. * This data is split into train and eval. Training * LoRA finetune of Qwen3.5-9B. * Distillation-free: we use Jev to only evaluate. * No RL yet! Serving * Parallel constrained decoding as suggested by @NielsRogge and @harshagundal. Results: * The post-trained Qwen (Nimble) became substantially better on our curated eval: 66% for Qwen to 90% for Nimble. Jev is at 93%. * 100ms on H100 and free to use on your macbook! Feel the AGI for free. * 2 days of building in public. :) Big caveat is that there is no standard benchmark to measure performance, and it's possible Nimble is much worse on other benchmarks compared to Jev. But it should be better than Qwen! We thank @typesafeai for making Jev and the inspiring discussions in the community. Hope this release lifts all the boats and encourages more research and activity in this space.
Should I build an OpenJev?
82
178
1,384
550,293
Mahesh Sathiamoorthy @ COLM2026 retweeted
Launching our AutoResearchExam benchmark (again) with a video. Never stop asking. Never stop launching.
1
6
22
910
If you don't have creativity and agency today with access to all these agi-esque tools, you are ngmi
4
1
18
1,746
This is my biggest fear btw
248
That sinking feeling when you are on the phone and your app can't reach the remote session..
4
12
1,595
This is actually hilarious
overheard in sf: “i ran out of data i need more data so bad” “i’m actually selling data rn what kinda data you need? coding data? healthcare data?” “bruh i need mint mobile my plan ran out” “oh that’s a shame. so you don’t need any coding data?” “no i do, what kind of data do you have?”
13
4,780
I will start a company to sell coffee that has, well, coffee and creatine and lions mane mushrooms and matcha and what did I miss?
People who drank more coffee had less body fat, less visceral fat, and more lean mass than less-frequent or non-drinkers. The biggest benefit was seen at 5 or more cups per day. And for men, every extra cup per day was associated with a ~10 ng/dL higher total testosterone. Higher coffee intake was also associated with better glucose and insulin regulation in men. This was an observational study, but mechanistic and experimental evidence actually supports coffee's benefits on metabolic and hormonal health. I'm willing to believe there's a real link here rather than draw a conclusion such as "coffee drinkers live healthier lives in general."
1
11
1,721
If you didn't know noul is from Bernoulli
1
9
959
Mahesh Sathiamoorthy @ COLM2026 retweeted
589497851845378089281879567916917641682963759302900576502946102 ± 1 are both prime Happy weekend
67
114
4,294
387,948
Hear me out. Decision models are small and non auto regressive. Let's make them large and autoregressive 😅
9
1
57
5,345
What was old becomes new, what was new becomes old
1
420
Who's going to COLM? I will be there and excited to meet people!
1
34
1,986
We are so back to the 2018 era, and this is so beautiful!
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions. CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks. With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks. We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed. Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size. 📄 Blog: contrastive-lm.notion.site 💻 Code: github.com/Contrastive-LM/CL… 🗣️ Discord: discord.gg/5dAQEDJBs 🤗 Data & Models: huggingface.co/Contrastive-L… More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
7
9
196
28,211
This is harder than training a model. :D PS: Planning to release a better version of Nimble soon
2
10
2,012
Nimble is already powering research where access to the model is needed!
Replying to @yufan_zhuang
4. Adversarial training maybe a remedy, we show that training on the injected prompts can significantly improve Nimble's robustness, and reduce the flipping rate
2
21
1,711
Should I host Nimble?
7
17
3,600
"If you supervise the chain of thought, then you could lead the model into hiding it's intentions in a way that's unobservable". This, combined with the fact that models know when they are being evaluated should cause you to sit up and think.
5
1
21
1,453
Nimble is able to process images! This is a confusing floor plan where there are three bedrooms marked as Bedrm 1, Bedrm 3, Bedrm 4. Nimble is able to count the bedrooms and bathrooms without a hiccup, while the base Qwen struggles. Jev doesn't have multimodal support so i didn't check it. Oh damn, what have I done! :D
Introducing Bespoke Nimble: an open data, open model, open recipe for an open Jev. Code and info: github.com/bespokelabsai/nim… Model: huggingface.co/bespokelabs/B… Data: * A new data curation recipe called contrastive data curation. * Slightly change facts to generate negative data. This pushes the model to discriminate better and become a better decision maker. The calibration is implicit. * Didn't do ablations but I think this is a critical piece! * This also means training data doesn't need probabilities. * Data covered 10 categories, and is fully synthetic. * This data is split into train and eval. Training * LoRA finetune of Qwen3.5-9B. * Distillation-free: we use Jev to only evaluate. * No RL yet! Serving * Parallel constrained decoding as suggested by @NielsRogge and @harshagundal. Results: * The post-trained Qwen (Nimble) became substantially better on our curated eval: 66% for Qwen to 90% for Nimble. Jev is at 93%. * 100ms on H100 and free to use on your macbook! Feel the AGI for free. * 2 days of building in public. :) Big caveat is that there is no standard benchmark to measure performance, and it's possible Nimble is much worse on other benchmarks compared to Jev. But it should be better than Qwen! We thank @typesafeai for making Jev and the inspiring discussions in the community. Hope this release lifts all the boats and encourages more research and activity in this space.
22
17
181
15,654