CTO & co-founder @v7labs. Reliable AI agents at work, strange small models and weird interfaces for fun. I build everything I write about.

London, England
I trained an AI model to autocomplete my piano prompts. Here’s what happens when I give it the first few notes:
5
6
32
2,555
I think I built the world’s most overengineered TAS. I trained a small CNN+GRU on a handful of recorded DuckTales runs (Gemma didn’t work out). It plays those runs perfectly at ~1ms/decision. Deviate more than a few steps and it’s completely lost.
2
3
148
It used DuckTales TAS by Aglar and MESHUGGAH as training data
18
Nice to finally share this piece we worked on with OpenAI about V7 Go and our Context Graph. One interesting detail: GPT-5.6 Sol had saturated many of our existing evals, so we built a harder set with messier, real-world data. GPT-6 Astra scored 89% on the hardest tier.
We're proud to be building the most robust AI operating system for high-stakes financial work with @OpenAI Read more here: openai.com/index/v7/
3
134
Gemma 4 E2B went from 1% to 87.5% at locating Scrooge McDuck in DuckTales. After letting it play Doom yesterday, I fine-tuned it on 2,716 NES screenshots, using positions read directly from RAM as labels. Training and evaluation took 18 minutes on one RTX 4090. It still didn’t play any better, so perception wasn’t the only bottleneck. What would you try next?
1
3
150
The Jev Doom demo by @CompleteSkeptic is really cool, so while I’m waiting for access I tried the same thing with Gemma 4 E2B from @GoogleGemma. Mine takes 640×480 frames directly, with no intermediate text description of the scene. It shares the image prefill across batched action queries, then reads the controls directly from logits. It sort of works.
3
6
10
604
Think of the screenshot and shared instructions as the prompt. Gemma processes that once, and we cache the resulting keys and values. For each JSON field, we append a short question to an independent continuation of that cached prompt. All those questions run together in one batch. Instead of generating the answers autoregressively, we read the logits for a fixed set of answer tokens and map them to the allowed enum values. A boolean is simply an enum with two values.
1
2
111
The schema here is very basic: { move: "none" | "forward" | "backward"; strafe: "none" | "left" | "right"; turn: "none" | "left" | "right"; fire: boolean; use: boolean; }
2
75
Training the next RollTab to play the piano bass line live while you play the melody, so you can actually jam together instead of waiting for it to autocomplete after you stop. I used Gemini to score the last version because validation loss wasn’t a very good proxy, but it doesn’t seem consistent enough for bass lines. So I’m back to blind listening: the same 25-second melody, three bass lines, over and over and over again.
88
Say "sshhh" to drop your Mac volume 20%. 🔈 on! Tried a tiny trained model and a dumb heuristic that looks for sustained treble energy. With headphones, and very limited training data, the heuristic actually wins.
4
1
7
215
RollTab now supports MIDI out! 🎹 Send the AI’s performances to external MIDI devices, or straight into iOS apps like GarageBand.
3
103
Play a few notes on your piano → AI continues what you played in real time, on your iPhone. 🔊 Unmute! New RollTab release: five new sampling methods alongside top-k. Mirostat v2 has been a surprisingly good fit for music.
5
135
Been testing different sampling methods for my MIDI generation model. Top-k still works really well, but Mirostat v2 is surprisingly competitive.
1
3
133
Then I realised: all my DPO training samples were generated with top-k. So I may have accidentally post-trained the model to be particularly good at top-k sampling. Needs more testing!
1
50
5 days and my first App Store submission is still just “Waiting for Review” 🫠 Anyone else seeing long waits on first submissions lately? How long did yours take?
2
2
205
10 days now. I have already finished version 2 😅
38
gitops is amazing until github is down. then it’s just git oops
1
4
243