A research lab by @baseten working to advance and democratize open-source intelligence.

SF
Distillation is a common way to prepare a model for RL where learning occurs from a stronger model, and then improves through trial and error. But how much does that better starting point help after RL? We conduct preliminary research on this question across model sizes and reasoning tasks. 🧵
9
23
257
47,560
We further dissected how the type of reasoning task impacted these findings. Larger models (4B, 8B) did better with RL alone on tasks involving calculation (e.g. a simple integral) while distillation retained an advantage on structured tasks (e.g. reversing spelling). Most importantly, a stronger starting model didn’t always become a stronger final model.
1
11
1,897
We also investigated different amounts of distillation and the rate of accuracy improvement that distillation confers, amongst other results. This is a first step toward understanding how different forms of distillation affect RL, and how to warm-start model training more effectively. Full work: labs.baseten.co/articles/whe…
1
1
27
7,456
We believe openness to be an advantage for AI safety. Openness provides more visibility into the behavior of models and, most importantly, greater means of turning safety research into actionable and transparent controls than closed-source. This is why Baseten and Base Labs are building a stronger safety and security standard for open models, with the launch of our safety infrastructure. Base Labs will develop and publish methods for training and monitoring open models, and Baseten will integrate that work into its deployment infrastructure, live at runtime, and offer this work as a managed service. This will be a standard that is transparent and built into how our models are trained and deployed. We invite the open-source community to contribute, and are proud to partner with @huggingface and @GoodfireAI to bring this vision to fruition. Together, we are building an ecosystem of open models that are safe and accessible to all.
19
40
304
58,531
We'd love to have more people involved. If you're interested in joining us, shoot us a DM.
1
9
797
A flavour of our research interests on the Dwarkesh Podcast this week, discussed by our very own @oneill_c. Just the start of a longer conversation about long-horizon RL and frontier open-source training here at Base Labs.
New episode with @johnschulman2, @oneill_c and @BerenMillidge. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. 0:00:00 – Steelmanning the case against RSI 0:18:39 – What’s driving the Chinese labs’ progress 0:28:06 – How will automated AI researchers be trained 0:33:51 – Will long-horizon RL elicit AGI? 0:45:24 – The sim-to-real gap 1:00:33 – How much progress is explained by data? 1:18:03 – Why is RL working so well? 1:24:54 – Move 37 and entropy collapse 1:28:31 – Rapid-fire timelines
1
3
24
3,780
RT @oneill_c: We RLd the equivalent of a controlled agent swarm to get qwen 122ba10b ot 63% on LAB Diligence (for context, 5.6 sol gets 12%…
2
40
Today we're announcing Base Labs, a dedicated research organization focused on advancing open-source AI. We believe in a healthy, open frontier model ecosystem. To enable this, we are working on: - Blue-sky research on continual learning, the science of RL, and how models learn across their full lifecycle, with every experiment and recipe shared openly. - The BaseHub Data Foundry: the highest-quality open RL environments, training data, and real-world benchmarks, built for anyone to train and benchmark on. - Post-post training: taking open-source models and making them better, safer, and more aligned through continual post-training, built on our research, and deployed with our frontier safety stack so organizations can use open-source models with confidence. - Making models cheaper and more performant through our model performance research. This is a mission-driven research effort, not a commercial product. We believe the health of the open-source AI ecosystem matters and that the best way to advance it is to do science in the open. We’re hiring engineers, researchers, and research fellows to advance this mission. labs.baseten.co/
55
118
995
238,813