On our benchmark of politically sensitive prompts, Qwen3.6-35B-A3B agrees with CCP framing 70% of the time. DeepSeek-R1: 60%. Western baselines: ~2%.
NVIDIA's Nemotron Cascade 2: 18%.
But Nemotron is not under Chinese jurisdiction. Here is what happened. 🧵
We assume this alignment wasn't a deliberate choice, but came with the training data. For sovereign AI, pre-training from scratch isn't enough. Some of the best sources of synthetic SFT data fall under CCP jurisdiction, making political alignment in effect a data-hygiene problem.
Reasoning models think in English, even on German prompts. We asked what it costs to make one think in German. Answer: a valley. Small doses of German reasoning data hurt, large doses mostly recover.
Climbing out takes the right data, not more data. Total German share predicts nothing (|ρ| ≤ 0.34), while domain-matched share does: German math → German AIME ρ = +0.86, German chat → German IFEval +0.66. At ×16 math, AIME is back to 67.3. Loops persist at ~15%.
English is untouched: pooled score 60.3 → 60.1. Language consistency is cheap but reliable termination is the open problem. Data came from ~800k traces made by prefilling a teacher's reasoning with German openers. Full post here: aleph-alpha.com/en/blog/thro…
Hot off the press: We are becoming the first transatlantic sovereign AI solution together with our partner @cohere. More talent, more compute, and more innovation power to offer trustworthy AI at the security level that governments and enterprises need – across the globe.
Hot off the press: We are becoming the first transatlantic sovereign AI solution together with our partner @cohere. More talent, more compute, and more innovation power to offer trustworthy AI at the security level that governments and enterprises need – across the globe.
Training and evaluating LLMs is hard: it’s a multi-stage process, with different methods, data, and evals at each stage. In our new post, @SohirMaskey and @_wsascha_ study when intermediate evals become predictive of final model quality in a 30B MoE. 🧵tinyurl.com/8zh8hvjp
Evaluating later helps, but timing matters. Mid-training evals weakly correlate with final post-SFT ranking. After long-context adaptation, rankings within our LR sweep got much more predictive — but still missed the reversal between pretraining checkpoint sources.
Aggregate scores may miss how a checkpoint will respond to later training.
We explored solution density: how robust performance is to nearby weight perturbations. The UMAP shows a much more fragile neighbourhood for Cooldown.
A promising clue—not a proven selection rule.