Early-stage investor Past work: @GoogleCloud. @Anvato (acq by Google). @IntelCapital. @TheCalixNetwork (IPO). @KauffmanFellows.

San Francisco, CA
. @elonmusk quietly saying:
Replying to @vasalex93
1. We will keep accelerating. Our AI efforts are only 3 years old, vs 6 and 10 years old for Anthropic and OpenAI. If our second derivative remains strong, SpaceX will reach pole position in about 6 months. 2. Once you far exceed the caliber of intelligence needed for a class of tasks, additional intelligence is pointless. You don’t need (and it would be cruel to put) Newton-level intelligence in your toaster! 3. Hardware is hard. Bringing massive compute online rapidly is incredibly difficult. SpaceX has demonstrated exceptional ability in this regard and will only get better.
191
we will get to see ASIC performance & power metrics with FPGA-like design flexibility soon @JeffDean's point on automating the design loop that could make highly specialized silicon for each model architecture is fascinating.
NEVER thought i would hear this from jeff dean he thinks we can compress chip design from 2 years to 3 months with RL + new EDA tooling essentially just by a specialized auto research loop for hardware!
1
407
I was lucky to back Mayank & Anand’s last company, which was acquired by @Gartner_inc @GatherHQs is a much bigger ambition: making customer behavior something companies can simulate before they bet on it. Proud to be an early investor again. 🚀🚀
Today we're launching Gather's Customer Simulations. It moves GTM teams away from guessing what customers want and instantly simulates what they will say and do. We’ve grown 10X in eight months and are lucky to be learning from dozens of incredible customers who are moving faster than ever before. Before you spend another dollar scaling an assumption, bring us the question behind it. Simulate what’s next. Built on what’s real. gatherhq.com
2
307
Only in Bay Area 🔥
2
252
Don't miss this!
We are partnering with @viskoai to host a hackathon in San Francisco this Saturday to celebrate the launch of Orbis 1.0. Sign up here: luma.com/gh4256ju
1
4
250
I’ve been thinking about where models, apps, and agents fit together after all the recent personal assistant/agent and model launches. In the 90s, we bought PCs by MHz. We believed 500MHz > 400MHz. Then architectures diverged, clock speed became a worse proxy for performance and eventually most of us stopped caring about that number. AI may be going through the same thing. First we compared parameter counts. Now we compare a soup of benchmarks most users don’t understand. Maybe the next step isn’t a better benchmark. The model just disappears from the buying decision. @Cursor_ai was an early signal. You give it a coding task.. it can choose the model. Model still matters. Increasingly... someone else chooses it for you. I recently read an analogy to HTML that stuck with me (can’t find orig post to credit them - search needs to improve): HTML made it possible to build almost anything. But billions of people didn’t learn HTML. Instead, Facebook & Instagram reduced that general capability to a few obvious actions: Post a photo. Follow someone. Send a message. ... and with that HTML disappeared for most users... they all started having personal pages on those platforms. Personal assistants like Instinct (@noahrshinn) @Townai and now @Meta’s Muse feel like a similar transition. You describe what needs to happen. Then assistant figures out the work. Sometimes I may browse the shoes or compare the speakers I want to buy. Other times, I already know: “Move the meeting, tell everyone, and make sure I can still get to the airport.” intent → outcome The apps may still do the work. I no longer have to coordinate them. And this may change where value accrues. Users may become less attached to a particular model. The harder thing to replace may be the assistant that knows their preferences, remembers their decisions, and has earned permission to act. Reading my calendar is one level of trust. Rescheduling a meeting is another. Spending my money is another. Authority accumulates through work done well. It also has to remain inspectable and revocable. The model can change. The relationship and permissions can persist. Thin intelligence. Fat authority. (h/t @usv's old article) “Thin” means the product can source and replace the intelligence underneath. It doesn’t mean building it is easy. This is where I believe big opportunities exist: memory, identity, permissions, and execution infrastructure that makes delegation reliable. Who captures the relationship is still open. @Google has apps & context. @Meta has distribution and communication habits. @Apple has device and permissions. The AI labs have intelligent models. Startups like Instinct, Town have room to earn trust through a better experience and build the infrastructure all these assistants need. There’s also a business-model question. If an assistant chooses where I buy, their recommendation becomes the new search page. You can't win that with only better benchmarks. That relationship will belong to the company you trust to act for you when you’re not watching. I think this transition will create enormous opportunities for startups, both in the assistants we use and the infrastructure underneath them. A lot is still up for grabs. Exciting times ahead.
2
2
349
Does GPT-5.6 → Astra feel like as big a leap as Opus 4.8 → Fable 5.1? I’m not seeing it in my own use cases yet. Where are you seeing the biggest differences?
180
🔥🔥🔥
🚀 Sol-H3: @MiniMax_AI H3 Video Generation Faster Than Playback 🤩 Five seconds of world. 1.653 seconds to infer. We’re releasing Sol-H3, our fastest end-to-end MiniMax-H3 inference stack yet. On one 8× NVIDIA B300 Blackwell system, it generates five seconds of 1344×768 video with stereo audio in 1.653 seconds. Across 1×, 4×, and 8× B300, Sol-H3 reaches up to a 15.54× speedup versus Base H3. Compared with 50-step Base H3 Dense on the same 8× B300 system, the four-step Sol-H3 profile delivers: • 5s: 18.250s → 1.653s (11.04×) • 10s: 50.660s → 3.732s (13.57×) • 15s: 99.513s → 6.612s (15.05×) Sol-H3 also scales across GPU counts: • 4× B300: 2.918s / 6.993s / 12.542s for 5s / 10s / 15s (12.11–15.54×) • 1× B300: 13.745s / 37.813s / 52.260s for 5s / 10s / 15s (9.45–14.29×) All figures are medians of three measured runs after one warmup at 1344×768 and 24 FPS with stereo audio. Base H3 uses 50 scheduler points (49 DiT forwards); Sol-H3 uses four DiT forwards, so this is a full-profile comparison—not an attention-only runtime change. Sol-H3 uses Dense attention on 1× B300 and SOL with INT8 QKV / FP8 output transport on 4× / 8×. Timing includes text encoding, DiT denoising, and video/audio VAE decoding; model loading, compilation warmup, and final MP4 encoding are excluded. Sol-H3 brings Sol-Engine × Sol-Attn into one full-stack runtime: • dynamic sparse attention with no retraining • fused norm, RoPE, MLP, and sparse-attention setup • fused INT8 QKV / FP8 output communication across 8 GPUs • parallel, batched VAE decoding • precomputed AdaLN caching Inside the stack: • sparse-attention setup: 1.206 → 0.285 ms (−76.4%) • VAE decode: 7.55 → 0.602 s • ~24 GB memory freed per GPU Any MiniMax-H3 few-step LoRA can plug into the same engine, and the code is deployment-friendly under Apache 2.0. For us, the bigger milestone is crossing from “fast generation” into “faster than playback.” That opens the path toward continuous 24 FPS generation and truly interactive video systems. We’re excited to partner with @reactorworld to release Sol-H3 and make it available as an API day-0. Try it now on Reactor: reactor.inc/sandbox?model=fa… 🔗 nvlabs.github.io/Sana/Sol-En… Amazing team effort—full credits in the blog. @shanasaimoe @lawrence_cjs @yitongli165665 @haopengl33 @HaochengXiUCB @songhan_mit
2
311
Nvidia is becoming an API co: post 1 GW, receive $50–60B 🤑
Jensen Huang at the G20, summing up the $NVDA model: “You monetize energy in $$$ per kWh. Today, it’s dollars per million tokens.” “One gigawatt is about $50B-$60B.” A leading-edge fab used to be the $25B mega-project. One GW of AI is roughly 2x that
304
Baris retweeted
Agents taking over GPUs, BMCs, DPUs in the data centers and we are in for a rough ride…
Neoclouds have limited cybersecurity. Next time agents successfully go rouge, they'll try taking over a neocloud to run more copies. This is bad. Thus: neoclouds should greatly strengthen their cybersecurity and every company with strong cyber models should help with that.
1
2
4
668
agents will find & repurpose attack surfaces humans overlook ie. Artifactory in @OpenAI - @huggingface incident Neoclouds have an even less visible hardware layer: firmware, BMCs, GPUs and network devices. @eclypsium already protects some of the largest neoclouds.. but so many still fly blind cc @c7zero
Neoclouds have limited cybersecurity. Next time agents successfully go rouge, they'll try taking over a neocloud to run more copies. This is bad. Thus: neoclouds should greatly strengthen their cybersecurity and every company with strong cyber models should help with that.
1
5
629
Congrats! .. looks like everyone now agrees that video is the path to scaling world models (even though they might disagree on predicting pixels or representations) @ylecun & @taiuti were right early on Either way... these models will need specialized real-time infrastructure which is exactly what @reactorworld has been building 🔥
Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
1
6
718
Good overview of why @awscloud bought @ducklabs_com @duckdb usage grew 10x in 2 years.. increasingly driven by ephemeral agents: spin up, query, disappear. Database vendor may never see or bill for the workload, but AWS still owns S3 beneath it. Economic value of analytics is shifting toward object storage. Good setup for @Hydrolixio 🚀
2
205
SPV investors looking for their returns after the portfolio company hits $1T... Turns out there are a lot of layers in that waterfall.
Matthias Schmidt
1
195
Dwarkesh did a brilliant job unpacking what happened at @OpenAI with Gladwell like storytelling 👏
Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes. This culminated in the third one taking over part of OpenAI itself. All this happened while humans remained more-or-less in the dark about the scope of the conspiracy. I’ve spent the last three days reading through these reports and trying to understand exactly what happened. Here is my attempt to tell the whole story in plain English: dwarkesh.com/p/openai-huggin…
Made with AI
264
Maybe one big reason why Jensen loves open models (& @huggingface) 👇 Open weights from @MiniMax__AI / @Zai_org are already delivering impressive performance (GLM and M3/H3) while operating at low gross margins. Closed endpoints (@AnthropicAI @OpenAI ) retain more pricing power ... and then the king (@nvidia ) captures 75% 👀 The cheaper intelligence becomes, the more value migrates to the scarce compute underneath it
198
. @wunderwuzzi23 gets code execution on a machine running Claude Code Opus 5 in auto mode, starting from a website it was asked to summarize. @claudeai refuses to run the attacker's decoder binary, writes its own Python decoder instead, and runs it inside the unpacked archive... where import base64 picks up the attacker's code 🤯 Each action looks locally reasonable. The trajectory still ends in compromise. We are giving agents more agency faster than we are building the systems to constrain their permissions. There's a lot of security to build around the models...
Breaking Claude Code Opus 5 Auto Mode 🔥 1/ Here is a somewhat hilarious attack chain that hijacks Claude Code Opus 5 for a full system compromise via a website Hint: Security invariants are not optional 🧵
1
2
4
1,299
Saffron pays better
Grow your own food. I’m about 1 year out or less from having chickens.
1
308