Product @Aleph__Alpha | Design Thinking, Design Engineering, AI Product Management. All views my own.

Berlin, Germany
Michael Hofmann retweeted
Benchmark text is all over training data. A model that has read the test set will nail it. That score measures memory, not skill. We scanned our mid-training pool and removed 1.5M documents. Our finding: trustworthy scores need decontaminated data.
1
14
66
5,851
Michael Hofmann retweeted
Does a model abstain when it doesn't know? We evaluated Kolibri Base and comparison models using the distributional correctness score to find out 🧵
1
2
13
1,704
Michael Hofmann retweeted
My favorite reinforcement learning plot from our Kolibri report: Inference v.s. training log probs, background is the logit gradient norm of importance-sampled Reinforce. Numerically unstable for small inference log-probs, but you can do exact token-level gradient clipping in RL.
Small bird, fast wings, Kolibri is here. 78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe. Now the weights are yours. Run it on your own hardware, under Apache 2.0.
2
5
27
1,064
Michael Hofmann retweeted
Love the reactions to our release today! Here are some of my favorite details covering architecture, load balancing and hyperparameter transfer from developing and training Kolibri, our 78B-total, 3.5B-active MoE model:
1
10
49
4,777
Really cool work from our SFT team
Here are two neat speed tricks we found while SFT'ing Kolibri 🧵 1: We use standard Next-k-fit online bin packing with document masking maximises useful tokens per sequence. Whats easy to miss: less padding means a larger effective batchsize, so the learning rate can go up.
10
1,768
Michael Hofmann retweeted
Our model, Kolibri, is out! 🐦 Incredibly proud of how much we've scaled in just a few months RL is as much an infrastructure challenge as a modeling one. We've scaled our RL environments and their infra to support 20K+ concurrent sandboxes per training run and 100K+ sandboxes overall. More details in the report A massive engineering effort behind the scenes. Proud of what we've built!
Small bird, fast wings, Kolibri is here. 78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe. Now the weights are yours. Run it on your own hardware, under Apache 2.0.
4
3
56
94,227
Michael Hofmann retweeted
germany just dropped a sovereign open weight model kolibri by @Aleph__Alpha runs 3.5b of its 78b parameters per word and its math is kinda ridiculous: 96.9% on aime beats every mixture-of-experts model they tested, even 3x bigger ones. only a dense model doing 8x the work wins anyone can run it on their own servers, it thinks in german, and in their evals it tops every open model its size in english and german models read text in chunks called tokens, and i ran kolibri's chunker (its tokenizer) on the german constitution: it needed 15% fewer tokens than gpt-5's for the same text. "bundesverfassungsgericht" is 6 tokens for gpt-5 and 2 for kolibri. fewer tokens means cheaper, faster german, and more of it fits in what the model can read at once how it works, simply: 1. every layer has 384 tiny specialists, and a router sends each word to 6 of them. so it thinks like a 3.5b model, but it needs the memory of a 78b one: about 78 gb, which means 2 big nvidia gpus (h100s) or 1 h200 2. most layers only look at the last 512 tokens, and every 5th layer looks at everything. that's how it can read 1 million tokens (a few thick books) without it costing a fortune 3. it reasons in german. their team found that a little german reasoning data is worse than none: the model's german thoughts go in circles and never finish. so they made about 800k german reasoning examples and gave it a lot 4. it's trained to say "i don't know". they play a game with it where parts of the documents are hidden, sometimes to help it and sometimes to hide the evidence, and it has to tell which. when it didn't know an answer it admitted it 44% of the time. qwen3.5 did 11% where it's weaker: answering from memory, using tools over a long back and forth, and coding agents, where qwen models are ahead. and to run it you need aleph alpha's add-on for vllm, a popular open source server for running models if you have german documents and need to keep them on your own hardware, this is a big deal. huge congrats to everyone at aleph alpha, my good friend @MichaelLHofmann included!! i wrote up how it works, the benchmarks, how to run it and when to use it: tej.as/blog/aleph-alpha-koli…
28
60
499
33,000
Today we launched Kolibri. On German National Day. A new LLM aleph-alpha.com/en/kolibri/ from Aleph Alpha available under Apache 2.0. Its been an intense few months across pre and post training to make this real. Look forward to getting adoption and feedback!
85
89
1,221
43,406
Michael Hofmann retweeted
Small bird, fast wings, Kolibri is here. 78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe. Now the weights are yours. Run it on your own hardware, under Apache 2.0.
431
1,014
9,835
2,656,175
Michael Hofmann retweeted
Hot off the press: We are becoming the first transatlantic sovereign AI solution together with our partner @cohere. More talent, more compute, and more innovation power to offer trustworthy AI at the security level that governments and enterprises need – across the globe.
3
15
89
5,670
Michael Hofmann retweeted
Today, we announce a landmark agreement with @cohere. By uniting our European research depth with global AI scale, we are building a transatlantic AI powerhouse to give enterprises control over their AI. Learn more about our shared vision here: businesswire.com/news/home/2…
4
11
125
12,874
Join us for the #DEsummershow featuring live online events and exhibition showcasing the graduating students and projects from our world-class courses. Register now at desummershow.london
7
3
Join us in our new building on 22nd March for an immersive celebration of the Imperial Design Engineering community and culture of innovation. eventbrite.co.uk/e/dyson-sch…
1
2
Replying to @leilathinks
@leilathinks and @MichaelLHofmann at the @TheManufacturer Smart Factory Expo in Liverpool talking about inspiring girls in STEM through Design Engineering
1
2
Great to be at @TheManufacturer @TMTop100 and Smart Expo in Liverpool as part of Digital Manufacturing Week representing @ImperialDyson and @ImperialDesSoc
4
9
Our library is packed for the first school evening guest lecture of this academic year. Tonight @assaashuach is talking on his vision of bio inspired semi autonomous design technology. #designengineering @imperialcollege
2
6
A pleasure to welcome @AnneMilton and @educationgovuk to show our Design Engineering research and student projects today. #STEM #Education @imperialcollege @RCA
4
5
Our doors are open for #LDF2018 come visit us on exhibition road. And For more pictures follow us on Instagram instagram.com/imperialdyson/ @imperialcollege @V_and_A @L_D_F #DysonLDF2018
2
7
Great to be back in Nottingham for #YSJconf2018
We're at @UniofNottingham for @YSJournal #YSJconf2018 come say hello at our stand in the Engineering and Learning Center.
1
3