New research! Some AI capabilities are both helpful and dangerous. E.g., knowledge of virology can be used to create life-saving vaccines or deadly pathogens. We introduce GRAM, a training method that puts dual-use capabilities (like virology) into removable modules.
24
54
377
449,962
The ideal: train two models. One with virology data, for trusted virologists, and one with virology data removed, for everyone else. But pretraining multiple models from scratch is very expensive.
1
1
58
6,834
Our method, GRAM, is a way to approximate multiple models in a single training run. We add small extra modules to the network, one per dual-use topic. When the model trains on virology text, the learning is isolated into a virology module.
1
2
57
5,309
To remove a capability, we delete its module. The model then performs like one that never learned from virology data in the first place. Using multiple modules, one model can be configured into many different versions, each knowing different things.
1
1
46
3,609
We tested this on real data covering general capabilities plus four dual-use subjects: virology, cybersecurity, nuclear physics, and niche code (a proxy). Deleting a module removes the knowledge as if the data was removed from training, while preserving general ability.
1
1
42
2,984
Compared to unlearning methods that modify existing models, GRAM removes capabilities far more robustly. A model trained with GRAM doesn’t regain the capability after a small amount of fine-tuning, whereas a model modified by unlearning does.

Jul 8, 2026 · 11:25 PM UTC

1
1
39
2,125
We tested models from 50M to 5B parameters. Knowledge separation gets better with scale: larger models forget the removed subjects more thoroughly and resist retraining better.
1
2
37
1,839
Real training data isn't neatly labeled. So we tested what happens when half of it has no topic labels at all. Versus alternatives, GRAM better separates dangerous knowledge from the rest of the model.
1
1
31
1,596
Our research draws on insights from prior work. GRAM is based on DEMix layers (Gururangan et al., 2021) and uses insights from extensions to gradient routing (Shilov et al., 2025). We are excited to see similar, concurrent work on NULLs (Ghosal et al., 2026).
1
1
31
1,458
This work is a proof of concept. More work is needed to translate it to production systems and understand whether it performs well there. We discuss limitations in the paper.
1
27
1,436
We're also hiring for alignment researchers and engineers! Come join our team: grnh.se/khoq17ou4us
2
23
1,510
Sort replies: Relevant Recent Liked