Real-time AI model APIs that run together, for voice agents, live video, and avatars.

San Francisco, CA
almost everything in ai is still optimized for stateless request and response. that falls apart the second you want video or audio a user can interrupt and steer. @keeganmccallum3 on tank talks, on why the state belongs on the server πŸ‘‡οΈ
How do you scale 500 to 9,000 H100s in six hours!! That's what it took to keep @DreamLabLA Dream Machine standing when it went viral. Half a million videos in 12 hours. @keeganmccallum3 is now building @urunml for the interactive AI wave. New Tank Talks [links below]
4
172
uRun retweeted
How do you scale 500 to 9,000 H100s in six hours!! That's what it took to keep @DreamLabLA Dream Machine standing when it went viral. Half a million videos in 12 hours. @keeganmccallum3 is now building @urunml for the interactive AI wave. New Tank Talks [links below]
1
4
4
867
$50 buys a day of continuously generated video. plenty of engineers spend that on coding tokens in an hour. our ceo @keeganmccallum3's @aiDotEngineer talk on what that changes is live. thanks @swyx + team πŸ™ new design partners: our dms open πŸ‘‹ piped.video/Xln-On3syJk
5
99
fast mode costs 2.5x and it's still waitlisted. the market is paying multiples for presence, not batch jobs. @keeganmccallum3 on why we make that the default πŸ‘‡
OpenAI and Anthropic charge ~2.5x for fast mode. Even then, OpenAI downgrades priority traffic and Anthropic runs a waitlist. That's the tell: a continuous, responsive session is worth multiples of a job you fire off and check on later. At @urunml build for that by default.
1
92
everyone with GPUs is being told to build a token factory. same tokens, same price war, margin competed away. the interesting bet is the workloads nobody else can serve. @keeganmccallum3 on why πŸ‘‡
Everyone with GPUs is being told to build a token factory and price on output. Reasonable advice. Also a race to the bottom. At @urunml we went the other way: serve the hardest workloads first, the ones nobody else can. That's where the margin actually is.
2
172
and that's @SIGGRAPH 2026 🎬 one last highlight video from the week 🀩 we showed up with a list of people to meet and we're leaving with an even longer list of ideas and inspiration. thank you to all of you who made that happen. see you next year πŸ‘‹
3
87
day two of @siggraph: a walk through the exhibition floor 🀝 live demos in every direction and builders happy to go deep on how it all works. we left with a long list of tools to try. a few favorites from yesterday πŸ“Έ
1
4
139
the week has started off strong at @SIGGRAPH πŸŽ₯ scenes from the floor so far below if you're here: go get lost in The In-Between, a whole wing of displays and installations. budget more time than you think you need. excited for the rest of the week
1
4
79
thank you to everyone who gave us their monday night πŸ₯‚ one table at @GwenLA during @SIGGRAPH week with studio pipeline leads, vfx supervisors, founders building interactive video, and the researchers they all cite. the kind of conversations that need a dinner table, not a booth. grateful you came. we left feeling full (literally and metaphorically). πŸ₯‚
2
65
we're at @siggraph 2026 in LA this week 🌴 the team is stoked to trade notes on real-time rendering, neural graphics, and world models this week with the community! building something where the pixels have to show up in real time? DMs are open, come say hi πŸ‘‹
1
4
82
every app you've ever used is pixels rendered to a screen. now there are models that can generate those pixels in real time, in response to you. @keeganmccallum3 on where that goes πŸ‘‡
Everything you've ever done on a computer is you interacting with a video. We just never had models that could render those pixels in real time, in response to you. Now we do. You can speak an experience into existence and interact with it as it renders with @urunml.
1
27
we're reminiscing about this year's @aiDotEngineer World's Fair earlier this month πŸŒ‰ πŸ₯° thanks to everyone who filled the room for @keeganmccallum3's talk on real-time generative video, and traded notes with us on the floor. the energy around real-time was hard to miss this year. good to spend a few days with the people building the interactive era with us. already looking forward to the next one πŸ‘‹
1
5
432
talking at @aiDotEngineer in a few hours. "Generative Video at the Speed of Light" @ 2:25pm the idea: real-time generative video isn't a model problem anymore, it's a serving problem. you can't stream 24fps off a different GPU every frame and land inside 300ms. the whole data path from browser to GPU is the surface you optimize. stop by if you're at the conference and say hi πŸ‘‹
1
2
83
Today at @aiDotEngineer World's Fair: @keeganmccallum3 on "Generative Video at the Speed of Light" @ 2:25 PM (room 2010) what breaks when real-time video becomes a serving problem, and why every frame has to land inside 300ms. come through and say hi πŸ‘‹
1
6
865
giving a talk at @aiDotEngineer World's Fair next week in SF: "Generative Video at the Speed of Light" on jul 2 @ 2:25pm. the model side of real-time video is mostly solved. self-forcing, rolling-forcing, 50 diffusion steps down to 3-4. the open problem now is serving it. would love to see you there (and hear your thoughts after)!
1
3
93
we'll be at @aiDotEngineer World's Fair next week in SF (jun 29 – jul 2). our founder @keeganmccallum3 is giving a talk: "Generative Video at the Speed of Light" πŸ—“ thu jul 2 Β· 2:25 PM Β· room 2010 real-time generative video is a serving problem now. come find us and say hi πŸ‘‹
1
2
56
Denver, thank you πŸ™ Last week's happy hour at CVPR 2026 was everything we hoped for. A room full of sharp people and real conversation about where realtime AI media goes next. Already looking forward to the next one. πŸ‘€
2
662
tonight. 6PM. denver. come say hi and tell us your favorite part of the CVPR so far! if you're on the list, you've got the address. if you're not, shoot us a message and we'll sort it. luma.com/hll4zk65
1
83
"Real-time" in marketing materials usually means "we batched it and it returns in 8 seconds." Real-time means under 300ms. That's the threshold UI designers use to decide when to show a loading spinner. There's a big difference. We build for the second one at @urunml.
1
1
93