We create specialized large language models for a sovereign Europe. Join us: jobs.ashbyhq.com/AlephAlpha #artificialintelligence, #writtenbyahuman

Heidelberg, Germany
On our benchmark of politically sensitive prompts, Qwen3.6-35B-A3B agrees with CCP framing 70% of the time. DeepSeek-R1: 60%. Western baselines: ~2%. NVIDIA's Nemotron Cascade 2: 18%. But Nemotron is not under Chinese jurisdiction. Here is what happened. 🧵
1
4
46
3,094
We assume this alignment wasn't a deliberate choice, but came with the training data. For sovereign AI, pre-training from scratch isn't enough. Some of the best sources of synthetic SFT data fall under CCP jurisdiction, making political alignment in effect a data-hygiene problem.
1
2
254
Reasoning models think in English, even on German prompts. We asked what it costs to make one think in German. Answer: a valley. Small doses of German reasoning data hurt, large doses mostly recover.
13
26
229
98,652
Climbing out takes the right data, not more data. Total German share predicts nothing (|ρ| ≤ 0.34), while domain-matched share does: German math → German AIME ρ = +0.86, German chat → German IFEval +0.66. At ×16 math, AIME is back to 67.3. Loops persist at ~15%.
1
6
2,642
English is untouched: pooled score 60.3 → 60.1. Language consistency is cheap but reliable termination is the open problem. Data came from ~800k traces made by prefilling a teacher's reasoning with German openers. Full post here: aleph-alpha.com/en/blog/thro…
1
9
2,315
Aleph Alpha retweeted
Hot off the press: We are becoming the first transatlantic sovereign AI solution together with our partner @cohere. More talent, more compute, and more innovation power to offer trustworthy AI at the security level that governments and enterprises need – across the globe.
3
15
78
3,401
Hot off the press: We are becoming the first transatlantic sovereign AI solution together with our partner @cohere. More talent, more compute, and more innovation power to offer trustworthy AI at the security level that governments and enterprises need – across the globe.
3
15
78
3,401
This is a big day for sovereign AI.
1
1
596
Training and evaluating LLMs is hard: it’s a multi-stage process, with different methods, data, and evals at each stage. In our new post, @SohirMaskey and @_wsascha_ study when intermediate evals become predictive of final model quality in a 30B MoE. 🧵tinyurl.com/8zh8hvjp
1
5
26
5,323
Evaluating later helps, but timing matters. Mid-training evals weakly correlate with final post-SFT ranking. After long-context adaptation, rankings within our LR sweep got much more predictive — but still missed the reversal between pretraining checkpoint sources.
1
2
190
Aggregate scores may miss how a checkpoint will respond to later training. We explored solution density: how robust performance is to nearby weight perturbations. The UMAP shows a much more fragile neighbourhood for Cooldown. A promising clue—not a proven selection rule.
3
181