New research! Some AI capabilities are both helpful and dangerous. E.g., knowledge of virology can be used to create life-saving vaccines or deadly pathogens. We introduce GRAM, a training method that puts dual-use capabilities (like virology) into removable modules.
24
54
377
449,962
Compared to unlearning methods that modify existing models, GRAM removes capabilities far more robustly. A model trained with GRAM doesn’t regain the capability after a small amount of fine-tuning, whereas a model modified by unlearning does.
Jul 8, 2026 · 11:25 PM UTC
1
1
39
2,125
Paper (ICML 2026, Spotlight): ae.studio/research/modular-p…
Code: github.com/agencyenterprise/…
Blog post: alignment.anthropic.com/2026…
Announcement: anthropic.com/research/off-s…
Paper Site: modularpretraining.com
1
1
26
1,547
We're also hiring for alignment researchers and engineers! Come join our team:
grnh.se/khoq17ou4us
2
23
1,510





