if you tried hyperframes, you would know the agentic video stack is being built on code-gen but SWE-bench code v.s. code-to-video that feels alive are two different things we built the Code2Video Bench, in collab w/ Google DeepMind & @Kaggle frontier labs can finally get good at agentic video tasks

Sep 21, 2026 · 5:17 PM UTC

473
361
1,402
1,665,013
Sort replies: Relevant Recent Liked
Wait what? GPT-5.5 is the best?
1
4
935
Yes - by a hair and within confidence interval
5
745
very excited to see Code2Video Bench come to life, awesome work by the Hyperframes / Heygen team!!!
5
74
5,621
Awesome Work!
4
406
The benchmark we’ve needed
1
4
301
Lets go Hyperframes 🚢🚢🚢
2
108
Exactly Code2Video Bench bridging code gen to video that feels alive is the missing eval we needed!
43
Kinda surprised this didn’t exist already
101
This is where things start moving fast
80
feedback loops win
61
public leaderboard or it didn't happen
22
kinda love how simple that is
52
The fact labs can train against this is the interesting bit
108
Love a public leaderboard
25
long overdue!
1
88
There’s probably a ton hiding in that data
167
This is how the quality starts compounding
23
can't wait to see people arguing over this leaderboard lol
33
I really like that video needs its own way to test these agents.
45
video finally has receipts
39
okay i’d actually browse through these matchups
76
feels like ai video just got a scoreboard
64
no more guessing from cherry-picked demos
47
this feels like a real step forward
85
Replying to @HyperFrames_
168 briefs is commitment
1
112
this space moves ridiculously fast
1
119
ai video has a report card now lol
1
1
98
we need more benchmarks like this
1
101
the training angle is the sleeper here
24
finally not an eval you need a phd to understand
138
this could push video quality forward fast
103
ai video has needed this badly
38
Interested to see what builders do with it
111