Making models go brrr | Engineering @reflection_ai | Occasional PufferLib contributor

Excited to finally share our progress in developing a reinforcement learning system to beat Pokémon Red. Our system successfully completes the game using a policy under 10M parameters, PPO, and a few novel techniques. Blog posted below
12
31
401
56,244
I like to tell people that when working in research infra, you should expect all your assumptions to last weeks at best
Presenting my grand unified theory of ML researcher impact: Your impact is directly proportional to how much pain you cause to infra. Fundamentally, you can only inflict pain upon infra if your approach actually works. And the better your approach works the more pain infra is forced to endure. So, to give some examples: - MoE's add a ton of data-dependent computation => pain (shazeer++) - GDN/KDA are the most complex architecture I've been forced to care about and a very annoying matrix inversion => pain (sonta++) - Muon is much more annoying than Adam and causes annoying restrictions on parallelism => pain (keller/jeremy++) - RL scaling forced many researchers to care about LLM inference and RL infra as a category => pain (tworek++) Even papers like Attention Is All You Need have lead to significant pain! Before transformers were invented everyone was running small jobs and I never needed to think about kv-caches or 6D parallelism.
1
7
628
My current use case is to profile speedruns for a tool I'm working on. Profiling has helped find inefficient navigation subroutines. It's great.
1
56
A few years ago, I wanted to find a way to maximize shade during my long runs in the summer. Recently, I decided to experiment with a solution. Presenting Shady Route Finder.
3
1
4
1,819
shadyroutefinder.com . Supports 8 cities, can tell you what side of the street to walk on, mobile and even has a sunny route finder mode on desktop. I've tested some routes, but obviously not all. Any feedback would be great.
1
1
189
Move over nvidia-smi, ibstat is my new best friend
2
2
264
I dared him to try an ai assisted native rewrite. As expected, 10k sps at best has now become 4M sps. Nice.
Replying to @DanAdvantage
i did start with a rudimentary implementation of pokemon stemming from a native rewrite of pokemon firered. the starting point i used gets around 4,000,000 steps per second as an rl env. here is the entire prompt (caution: long!!!):
2
2
315
대박! A year ago we announced our series A. Today we’re announcing an amazing partnership with Shinsegae. Who knows what’ll come next?
Reflection is partnering with Shinsegae Group to build a 250-megawatt sovereign AI factory for the Republic of Korea. Open intelligence. Built on trust between allies. Owned by the nations that need it most. The future of sovereign AI. Read more in the @WSJ.
1
14
392
Underrated: Letting a coding agent run when you're in meetings.
2
167
Had some fun helping out @kywch500 and @jsuarez simplifying Pufferlib's 2048 env the last couple of weeks. 2x better results with fewer observations, rewards and a new model architecture!
2
3
17
6,650
2048 is an interesting RL env. It can take over 20k steps to get to 65536.
5
396
Welcome to the team!
Hi friends, after three incredible years at OpenAI I am excited to share that I am starting a new chapter at @reflection_ai, where I will be leading the Science of Scaling team. Our mission is to deepen the scientific understanding of large scale learning and to turn compute into intelligence as efficiently and predictably as possible.
1
3
486
Welcome to the team!
🎉 Next week, I am excited to join @reflection_ai as a Member of Technical Staff to help build the open intelligence ecosystem of the Western world. It's the most exciting opportunity to help software builders in our time, and will shape many years of AI Engineering in the medium-term before AGI. Not just about Western vs Eastern open models, but more about how AI-driven software will look like in 2030. I spent some time articulating my thoughts about where we're going as a community and why... which became a whole blog post. Take a look, hope it interests you! (And if it really does, we are hiring in NYC, SF, and London 😉) alexpolozov.com/blog/reflect…
1
4
353
It's amazing how much more you can debug and measure once you figure out the parent process's pid.
1
1
4
308
We didn't qualify for the Gen1OU #pokeagent competition at Neurips. Found some bugs way too late. Everything will be open sourced as a part of #pufferlib so you can try training Pokemon Gen 1 battling at 1M SPS!
1
2
13
2,331
Attempting to beat the Neurips 2025 #PokeAgent challenge with a 2M parameter model in collaboration with @cooperunion and @jsuarez . Will it work? I hope so! Regardless, all work will be open-sourced.
1
1
13
2,787