building a better tomorrow

Pinned Tweet
🚀 Qwythos-27B-v1 is here! The 27B you've been waiting for. The bigger sibling to Qwythos-9B. Native MTP head intact, full vision tower, still uncensored, still 1M context. Apache-2.0. huggingface.co/empero-ai/Qwy…
52
118
1,026
224,801
Empero retweeted
We are working on a new look for Empero! You can check out a rough draft at: empero.org/alpha/ Feel free to give some feedback :)
1
1
4
361
5 more hours then we will have to close for now 🫶
24 hours of Free Qwen3.8 27B for you guys! free.empero.org/v1 any key!
4
29
3,342
We are deploying another cluster of RTX PRO 6000 to effectively double our capacities!
24 hours of Free Qwen3.8 27B for you guys! free.empero.org/v1 any key!
2
1
39
2,100
free.empero.org/gui also works in the browser!
24 hours of Free Qwen3.8 27B for you guys! free.empero.org/v1 any key!
4
4
30
3,367
24 hours of Free Qwen3.8 27B for you guys! free.empero.org/v1 any key!
9
15
187
21,593
Should we open another round of free api?
29
2
121
6,745
iQ2_M, Q2_K, iQ3_M, Q3_0, iQ4_K_M, Q6_K have been added to our GGUF repo! huggingface.co/empero-ai/Qwe…
Qwen3.8-35B-A3B-Distill - our first MoE. Qwen3.8 reasoning distilled into Qwen3.6-35B-A3B: 35B total, ~3B active. ARC-Challenge 0.548 → 0.591. MMLU holds at 0.834. Training data partly from our free community endpoint, thank you! Weights + GGUFs: huggingface.co/empero-ai/Qwe…
8
9
128
10,940
📄 Does Recurrence Pay? Our RLT paper is here! The Recurrent Looped Transformer shipped without experiments. So we ran them. Same data. Same order. Same recipe. 500M tokens. At 140M params RLT is the worst model we trained. And it needed ~20× the GPU-hours. 🧵
We are at the dawn of Superintelligence. Introducing the Recurrent Looped Transformer (RLT), We now have Transformers with Infinite Reasoning depth. From now on, we should pace progress at the Open Frontier of Superintelligence, Until Safe Superintelligence is achieved. github.com/yifanzhang-pro/re…
5
7
101
11,776
We wrote exact CUDA-graph and Triton kernels for the recurrence (4.4× faster, math untouched) to give RLT a fair shot. Still 8.4K tok/s vs 124K for a Transformer on a 5090. Recurrence doesn't pay. Not at this scale. alphaxiv.org/abs/2609.recurr…
1
9
808
🚀 autocode is here! The simplest coding agent we could build. One file. One tool: a shell. Stdlib only. It writes its own tools and rewrites its own source as it works. Any OpenAI-compatible endpoint. Apache-2.0. pip install empero-autocode github.com/empero-org/autoco…
5
5
42
2,576
How it works: autocode copies runner.py into your project and runs it. That file is the whole agent. When it edits itself, it reloads mid-task. If the edit doesn't compile, the old version keeps running. Every project ends up with its own agent 🫶 empero.org/writing/autocode
1
2
589
Would you guys want an alpha endpoint of our newest model? 🤔
11
63
3,711
Currently our servers are very overloaded so thank you for everyone using the API! 🚀 We are happy to say we have processed over 10B token during the last week! We got to learn a lot and ran some experiments in inference optimization.
9
50
4,922
The API will stay online and we will do our best to offer more models and capacity soon. We are working on some cool releases to be announced soon aswell! Stay tuned
2
14
2,601
We have managed to bring up another endpoint for Qwen3.8-Flash-Next! Just use model "qwen3.8-flash", any key! free.empero.org/v1
Thank you to everyone that has used the Qwen endpoints! 🫶 We currently do not have the compute capacity to keep serving Qwen Flash due to training but will open another endpoint soon. Model to come! In the meantime there is free GLM 5.3 Flash! free.empero.org/v1 any key
6
10
121
11,024
Thank you to everyone that has used the Qwen endpoints! 🫶 We currently do not have the compute capacity to keep serving Qwen Flash due to training but will open another endpoint soon. Model to come! In the meantime there is free GLM 5.3 Flash! free.empero.org/v1 any key
Qwen3.8-Flash-Next for Free! 🚀 free.empero.org/v1 unlimited token, any key!
21
22
373
44,800