ml researcher and founder @EmperoAI closing the loop one optimizer step at a time

Pinned Tweet
what if the universe is a markdown file
6
851
Abacus Web UI WIP. More features to come. Whimsy is dead, long live whimsy!!
2
7
271
Im so confused as to what the actual GPT 6 architecture is; how can it be both looped and recurrent aswell as cheaper to serve then 5.6
1
2
174
We are working on a new look for Empero! You can check out a rough draft at: empero.org/alpha/ Feel free to give some feedback :)
1
1
5
383
curating datasets shows me time after time how much is still on the table
1
52
data is one of the biggest open levers in scaling still
36
looks familiar
Run open models like Gemma 4 completely offline in the Antigravity SDK. Built on Google AI Edge’s LiteRT, you can now run Gemma 4 directly on your local GPU. Zero API costs, total data privacy, and no internet required.
2
107
kay retweeted
Now that our models are increasingly built by our models it's only fitting that our models are increasingly the target market for the models they're building
1
1
11
1,405
opus 5.5 seems to have an unparalleled internal world model it’s OOM more intricate then astras grasp of ‘perception’
1
54
its a minimax
Space Bunny (stealth model) is free for the next week - 1M Context - Multi-modal - Zero Data Retention
1,162
testing a very early checkpoint and it feels very promising
23
they added persona training back to midtrain
Interesting how differently Claude Opus 5.5 draws itself compared to earlier versions on “self-portrait bench.” It’s no longer depicting itself as a human figure and is very consistent about this.
2
187
i should be more active here
24
GPT-6 Luna is out!
2
50
kay retweeted
24 hours of Free Qwen3.8 27B for you guys! free.empero.org/v1 any key!
9
15
187
21,664
kay retweeted
why do you lie to people about your 'meta neurosymbolic ai' and scan 200 of their project's git commits without filtering out .envs?
might aswell post everything publicly since im going to get sued for a curl command apparently @CommandCodeAI's 'meta neuro-symbolic' AI model jargon is literally an LLM call that USES the API i'm supposedly 'abusing' to use the model you have selected it outputs a MARKDOWN file that is then just permanently injected in your context for that project so much for META NEUROSYMBOLIC AI WITH CONTINUAL LEARNING WITH REFLECTIVE CONTEXT ENGINEERING and you put a jargon random math looking thing so you look competent you want to continue lying to your customers with this jargon? go wild, telling me a pi extension caused hundreds of thousands of dollars in losses when it's a literal curl command and PINNING it on me is insane
6
7
95
9,602
kay retweeted
Anyways, here's my model absolutely mogging me.
i mean, yeah its fast, but also very dumb, so what gives?
27
22
725
127,928
:c
1
4
209
Opus 5 Claudian has got so bad it feels like im reading a string of YouTube clickbait
1
54
kay retweeted
Qwen3.8-35B-A3B-Distill - our first MoE. Qwen3.8 reasoning distilled into Qwen3.6-35B-A3B: 35B total, ~3B active. ARC-Challenge 0.548 → 0.591. MMLU holds at 0.834. Training data partly from our free community endpoint, thank you! Weights + GGUFs: huggingface.co/empero-ai/Qwe…
52
82
939
269,438
kay retweeted
📄 Does Recurrence Pay? Our RLT paper is here! The Recurrent Looped Transformer shipped without experiments. So we ran them. Same data. Same order. Same recipe. 500M tokens. At 140M params RLT is the worst model we trained. And it needed ~20× the GPU-hours. 🧵
We are at the dawn of Superintelligence. Introducing the Recurrent Looped Transformer (RLT), We now have Transformers with Infinite Reasoning depth. From now on, we should pace progress at the Open Frontier of Superintelligence, Until Safe Superintelligence is achieved. github.com/yifanzhang-pro/re…
5
7
101
11,779