I used to watch PewDiePie play Happy Wheels when I was 13, now I’m 25 and watching him train his own AI models
PewDiePie trained a frontier model at home and beat OpenAI and Gemini
> be me PewDiePie
> play games on YouTube and scream 24/7
> become a meme reviewer
> get unfathomably famous
> "fuck that" I'm a family guy now
> retire
> move to Japan with my beautiful wife
> mfw I'm a dad now
> "as a dad I must do dad things"
> scratch that
> "as a dad I must do frontier AI research"
> goal is to beat GPT-4o at coding (16% on Aider)
> buys $30000 GPU setup
> reads DeepSeek paper
> decides to start massive GitHub scraping and data augmenting run
> not good enough
> read Magicoder paper
> generate tons of synthetic coding data
> train a new model
> guuuuuh. the data made the model worse
> mfw I just wasted months for nothing
> decides to lock in and try again
> makes model worse again ffs
> try again
> finally beating GPT-4o on data (16.1%)
> not satisfied
> "I should simply train a reasoning model"
> reads more papers
> start experimenting with more synthetic data
> "Mhhh something doesn't smell right"
> house almost burned down due to power connector
> shrug
> just buy a new one
> mfw computer is now crashing 24/7 generating synthetic data
> new plan: just call DeepSeek API for high quality synthetic data
> train model again
> 17.2%
> performance fluctuates slightly on each eval run
> big brain idea: repeat eval until we randomly reach >18%
> sike actually got 19.6%
> feelsgoodman.png
> nvm the benchmark was contaminated and I was training the wrong base model the whole time
> rerun everything again
> new score: 4.4%
> you read that right REEEEEEEE
> almost get a heart attack
> "have you tried plugging the device off and back on again?"
> change nothing and just retrain again
> 25.3 %
> LETS FUCKING GOOO
> realize that 1/3rd of the benchmark was not running. guuuh
> scared shitless it would score below 10% again
> run yet again. the whole thing this time
> Thirty fucking six percent
> accidentally beat Gemini 2.0 Pro Exp and GPT-4.1 mini
> pops the AI bubble
> "I want moaaaar"
> finds some more post-training data
> 39% babyyyy
> realize at the end that I was just benchmaxxing Aider polyglot
> next quest: run SWE-Bench and other coding benchmarks
> "I failed a thousand times, but prevailed in the end"
> just a little sad side-quest
> probably going to train GPT-6 myself by next month