Partner @a16z investing in infra & AI | Prev product @databricks & @uber, founder at Jemi (YC S20, acq) | Technology optimist ☀️

San Francisco, CA
Jason Cui retweeted
jev isn't production ready. today @modaicdev is launching Mo, the first multimodal decision model that learns from its mistakes. most production tasks involve subjective decision-making, where the last mile of calibration and alignment has always fallen on the developer. Mo closes that gap on its own. it refines its instructions and weights from the signals already in your workflow to learn what a good decision means for you. use Mo for the decisions where alpha matters the most like LLM judging, lead scoring and qualification, content moderation, resume screening, and more. if you're already using another decision model like jev, plug it into Modaic and it becomes self-improving too :) 🐙 modaic.dev docs.modaic.dev
18
16
62
47,182
Just hit one year on the best infra team and they got Starter Bakery pastries to celebrate. Time flies when you’re having fun!!
7
42
1,791
Jason Cui retweeted
@fal's H3 Max generates video faster than you can watch it. With director mode and voice prompting, you can direct a scene as it plays - moving the camera and guiding the action just by talking to the model. The leap isn’t just speed. It’s creative control. That’s what takes AI from impressive demos to a serious technology for Hollywood and professional filmmakers. Inspiring convo with @gorkem and @isidentical on what possibilities are unfolding in generative media.
.@fal's Gorkem Yurtseven and Batuhan Taskaya on making an open source video model 35x faster, and what Hollywood wanted after they built it: Last month, MiniMax released H3, an open source video model. fal rebuilt it - they cut down the steps the model takes to make a video, rewrote the code under each stage, and got the GPUs to 70-80% of their theoretical ceiling instead of their usual 30-40%. No quality loss. Video now generates faster than you can film it. The models have gotten so cheap and fast that end users aren't even asking for improvements in either category anymore. The gap has moved to quality, or how closely the model follows the prompt. Hollywood wasn't a customer a year ago and is now fal's fastest growing segment. Studios love it for the little things - extending a shot, moving the camera, changing the lighting. Those edits land 80-90% of the time. fal is chasing 99.9%. In this conversation with a16z's Jennifer Li: 00:00 Intro 01:45 The first open model worth rebuilding 03:45 The industry ran out of compute in April 06:25 35x faster without losing quality 11:05 5s of video generated in 1.5s 14:25 The weekend fal dropped everything 16:15 3 viral projects nobody planned 18:20 Teaching a video model to remember 20:25 An hour of video that's consistent 23:30 2x the usage of other models in 3 weeks 26:35 The case for giving video away for free 29:00 Blender sketch in, finished shot out 30:40 Chasing 99.9% reliability 34:15 Hollywood, fal's fastest-growing customer 36:50 Why studios wouldn't touch it until now YouTube: piped.video/SDbRJXQrYGY @gorkem @isidentical @fal @JenniferHli
6
10
46
3,773
Arrow 2 beats many big models on design and vector graphics. Super impressive
Introducing Arrow 2 Our latest and most advanced models for generating precise, editable vector graphics. Higher quality. Faster outputs. Available now in App and API.
4
823
We’re partnering with @inferact (the team behind @vllm_project) to make TPUs a first-class citizen for open-source AI inference. Together, we’re delivering: - Upstreamed open-source TPU kernels - Native PyTorch integration via TorchTPU + so much more ↓
@googlecloud and Inferact are announcing today a partnership to make TPU a first-class citizen in @vllm_project. This partnership puts both teams on one engineering roadmap to bring TPU to the broader open model ecosystem, optimizing vLLM as the agentic production serving engine for TPU: • Production serving features and optimized kernels • A native PyTorch path via TorchTPU • Moving towards day-0 support for frontier model releases We're also launching a community program: shared TPU capacity for open-source contributors, plus dedicated review and design help from the core vLLM maintainers at Inferact. Everything this collaboration produces is open source. Read the full announcement: inferact.ai/news/google-tpu-…
8
23
154
32,082
Jason Cui retweeted
Vals AI co-founder and CEO Rayan Krishnan with a16z's Ben Horowitz and Jennifer Li on grading AI, what it costs, and who gets to make the rules: Every big industry eventually grows an independent testing layer. AI has credit ratings to learn from and Enron to avoid. Model capability today is still mostly self-reported. As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. The harder problem is geopolitical. Reagan's "trust but verify" worked during the Cold War because you could fly over and count the missiles. No simple equivalent for AI models exists. In this conversation with Erik Torenberg, they get into how you measure a model's ability to improve itself, why every good benchmark eventually has to be retired, and what happens when token spend begins to rival employee salaries. 00:00 Intro 02:20 Llama 4 on public vs private benchmarks 05:24 Nobody agreed how to test humans either 06:55 What movie ratings teach us about AI 08:55 The Enron problem in benchmarking 11:36 Why a good benchmark has to be retired 13:22 Evals that run for weeks, not seconds 16:20 Where the real workday starts at 4pm 18:08 A firm really is just its evals 20:35 Why Sonnet can cost more than Opus 22:42 One engineer, 6 billion tokens in a day 25:05 Who should set the rules for models 28:55 Public sector enforces, private verifies 33:32 Why sovereign AI is inefficient and happening anyway 35:00 The AI version of trust-but-verify 37:15 Where cyber evals have to go next YouTube: piped.video/watch?v=WO9c9qxD… @RayanKrishnan @ValsAI @bhorowitz @JenniferHli @eriktorenberg
26
17
140
159,950
A very interesting breakdown that points to the future of data agents! Model capability improvements will continue to close gaps, but human guided semantic context for agents is still crucial for real world data messiness.
As better models rapidly eat the stack, what system research problems will remain to enable the next frontier of data agents? We studied the evolution of data agent capabilities the past 2 years, finding general coding agents beat hand-designed data agents by up to 37 points, with 4× fewer turns. So what's actually left for system researchers to solve? More than you might think... 📚 Paper: arxiv.org/abs/2609.03141 🧵👇
3
2
23
1,958
Jason Cui retweeted
OpenAI's Mark Sellke and Mehtaab Sawhney with a16z's Lisha Li, on the state of AI and mathematics: Before OpenAI released GPT‑6 Astra last week, the model was already doing original mathematics. Recorded before the launch, this conversation tells the story of how it got there. It began with GPT‑5. Mehtaab Sawhney pasted in an Erdős problem still listed as open. Five minutes later, the model surfaced a paper that had solved it. The exercise eventually uncovered published solutions to 10 more problems thought to be open. Then Astra went further. Told to "go have fun" with a high-dimensional sphere-packing problem, it improved a bound that had stood since the 1970s. Mehtaab had spent six months on the same problem in graduate school and made "absolutely zero progress." Another Astra result established that non-sofic groups exist with a roughly 15-page proof. A related human breakthrough took 250 pages and machinery from quantum complexity theory. OpenAI’s Mark Sellke and Mehtaab Sawhney join a16z’s Lisha Li on why wrong ideas pollute a human’s context window, why polished papers hide how mathematics is actually made, why a breakthrough can stop one prompt early, and what math rewards once proving stops being the bottleneck. 00:00 Intro 02:44 Cracking an Erdős problem in 5 minutes 06:20 Why a human quits and a model doesn't 08:50 Wrong ideas pollute your context window 11:45 Why math papers are bad training data 16:20 Nobody knows how to stack spheres in high dimensions 18:28 The orange-stacking proof 21:04 The 1970s Russian paper nobody could beat 24:14 The function that won a Fields Medal 27:09 How Astra beat the sphere-stacking record 29:20 Why error correction is sphere packing in disguise 35:15 The breakthrough Astra almost didn't bother with 39:22 Solving harder problems means it has better taste 41:00 One model for taste, one for the grind 44:40 Astra found an infinite group no finite one can imitate 52:00 250 pages of quantum complexity, or 15 of group theory 56:18 Only humans write 200-page proofs 1:00:38 What changes when proving stops being the bottleneck 1:02:44 The problems AI may never solve YouTube: piped.video/watch?v=1JvyLGd2… @mehtaab_sawhney @MarkSellke @OpenAI @lishali88
17
25
147
131,628
With the pace of shipping in consumer AI and personal assistants, it's easy to forget how far we've come in terms of capability and autonomy. When chatGPT first came out it was an incredibly mind-opening moment for the general public - the biggest public exposure to LLM capability at the time. It was impressive in its conversationalness and writing, but also confident in its hallucinations and its knowledge stopped in 2021. Asking something as simple as the score to a recent game quickly outlined the gaps. Early chatGPT was magical, but the trust was not there for many, if not all tasks. Fast forward a few years and AI teammates/assistants like Grok @bot, Instinct, and others are redefining how we interact with AI systems and how autonomous they can be. Iterating on the momentum of OpenClaw and abstracting out and simplifying a lot of the core workflows, AI can do more, effectively. The biggest thing is a huge step function increase in trust. We now trust these systems to code, write, book, message, and act on behalf of us. If this is how far we've come in a few years, the next decade will be truly incredible.
1
1
18
1,802
Inference is one of the most interesting systems problems at scale. Near unlimited demand, wildly different workloads. Constantly shifting targets for models to optimize, and trade offs between cost, latency, throughput.
26
35
578
47,835
Jason Cui retweeted
We're thrilled to lead a $300M investment in Gimlet Labs. AI inference is one of the fastest-growing markets in the history of capitalism, and we are running out of nearly every physical input required to serve it. Inference demand compounds at software speed; power plants, data centers, and semiconductor fabs do not. We believe @gimletlabs has built the solution: the first multi-silicon inference cloud, designed to produce more intelligence from every watt. This system is already delivering up to 10x gains in throughput and interactivity on frontier models within the same power envelope. In a market starved for compute, efficiency is net-new capacity and latency is product differentiation. Wherever the existing stack ends, the Gimlet team starts. Welcome, @zainasgar, Michelle Nguyen, @oazizi, @nserrino, James Bartlett, and the Gimlet Labs team! By @RaghuRaghuram, @sarahdingwang, @shangdaxu, @steph_zhang
1/ Today, we announced Gimlet Labs’ $300M Series B, led by @a16z, joined by @SapphireVC as a major investor bringing our valuation to $3B. We started Gimlet with a simple conviction: inference would become the dominant AI workload, and its infrastructure would need to be rebuilt from the ground up.
21
41
508
93,874
Super insightful convo with @lishali88 and Daniel Litt!
AI is now solving serious open problems in math. I talked to @littmath about what the models have actually mastered so far, what's still hard, and what that gap tells us about the frontier of reasoning capabilities. Measuring progress in math capabilities gives us a way to probe what higher level reasoning might still be missing. It’s also an opportunity to discuss what happens as the capabilities models have already mastered become abundant. Daniel and I get into which recent AI math results are the most impressive and why; whether increasingly difficult conjectures could themselves form a curriculum for better reasoning; whether AI math could mode collapse onto the same styles of thought; what mathematical taste is; and whether beauty is even something worth optimizing for. 👶♾️And, because we both have toddlers, how you teach kids math in a world with increasingly capable AI :) We recorded just before the recent S⁶ result, so maybe an illuminating version of its proof has dethroned the Erdős unit distance problem as the most impressive result so far? Our hour long chat here!
1
5
1,506
Jason Cui retweeted
Our fifth growth fund is now $8.5b. The founder is the asset in late stage, and we back the best in the world. At the same time, the best founders face global complexity earlier in their lives than ever before. So we are building out an enhanced Growth Platform for these exceptional founders. Biggest product cycles of our lives. Never been higher stakes, and never been more excited. @a16z
21
24
255
80,810
Machine Age!
$1.1B for the Machine Age.
2
19
1,300
Super excited for our new Machine Age Fund, which will invest in all of the computer infrastructure on which AI runs - including chips, memory, networking, and storage. And also the full systems AI runs on - from data centers to robots. This is a once in a generation opportunity to invest in improving the bottlenecks around hardware!
2
32
2,375
Jason Cui retweeted
Distributed AI infrastructure works because of the network. GPU clusters rely on multiple network fabrics, each adding complexity as infrastructure scales, tenants change, and new capacity comes online. Traditional network automation wasn’t built for this dynamic environment or the hard tenant isolation it requires. That’s where NAAM – Network Automation, Abstraction, and Multi-Tenancy – comes in. As @a16z’s @martin_casado puts it: “The future of AI is gated on our ability to scale up the network. It is the harder problem.” In this video, Martin and a16z partners @appenz and @RaghuRaghuram join Netris CEO @alex_saroyan and co-founders Tigran Martirosyan and Arsen Arakelyan to discuss what’s changing in AI networking and why a new approach is needed. Netris is already putting NAAM to work across 35+ live AI deployments. Watch the video to learn why a16z believes the network is becoming one of AI infrastructure’s most critical challenges.
1
4
185
Jason Cui retweeted
Honored to cohost another super fun event with the one and only Cursor & SpaceX team!! Sign up below 🙌🙌
we're hosting our first-ever grok bot build night for women with our friends at @a16z 💛 you'll hear from @poteto (!!), @clairevo & @stuffyokodraws with tons of hands-on build time. if you're in sf, join us!
4
2
36
2,144
Jason Cui retweeted
Said it before and saying it again. It always takes a village to make a company, and to make a legend. Very proud of what we’ve built here - no one tries to be the only hero on a company or a deal, everyone works together to help our companies succeed. Go Team Infra!!! See how our team mom @KatieBaynes prepared us for a family photo.
I've never worked on a deal in venture that didn't involve multiple people. Not just in supporting roles, but critical to getting the deal done and making the company successful. I don't think we recognize enough how much venture really is a team game.
4
3
74
13,600
The team in question :)
The best deals get done as a team. And even if things don’t work out the way you intend, you learn as a team. This is such an important and understated part of our team culture that makes it special. And it also just makes everything more fun!
7
1
135
37,852