Managing Partner @500GlobalVC. Tech adviser, investor and executive. Previously: operator at @Google @Twitter @Color

🇹🇼-🇺🇸-🇬🇧-🇺🇸-🇹🇼-🇺🇸
Pinned Tweet
I don't get enough credit for not stepping on my dog. http://t.co/1nHTdj1IL6
22
29
279
There is alpha here for people paying attention 👇🏼
Or, build a cluster and own the hardware, make your money go way further, still have a valuable asset after 3-5 years, and make a differentiated model bet. That's what humans& did Huge thanks to @DeepInfra and to @nvidia @Supermicro for working with us to make it happen
1
2
1,325
😮 In the future, rival AI clusters will be sending humans to each other as good will gestures. Stay cute, humans. 🐼
🚨 JUST IN: Xi Jinping announces China will be sending TWO Giant Pandas to the Atlanta Zoo, Ping Ping and Fuhuan "The giant panda has been an envoy of friendship between the Chinese and Americans. Here, I have a piece of good news. In a few days, two pandas, Ping Ping and Fuhuan will come to their new home in Zoo Atlanta and meet with the American people." I freaking love pandas. So I'm glad to hear this 😆
1
479
I’ll be in Austin for this with @TomorrowXSummit. If you are around, let me know!
Welcome Nasdaq Private Market (@NPM) to TomorrowX Summit. Nasdaq Private Market is the platform powering modern private markets — serving private companies and their investors with liquidity, capital and investment solutions. Come find them at the NPM Cafe at the Moody Center. Registration link in the replies.
1
324
If you are a founder that can do this for every other company, I would love to talk to you 🙏
Greg Brockman says OpenAI pointed Astra at its own systems until it ran out of vulnerabilities to find: "We took 25% of our production engineers and said, 'Sorry, all your projects are on hold. You are now defending. You are now up-leveling our security architecture. You're going to use the models to find all the holes.' And we found a number of serious issues, and we fixed them." "We found some new problems, but eventually it saturated. We basically have found, to our knowledge, all of the P0s, all of the critical problems that Astra is smart enough to find. And of course, there will be a new model, there will be a new round." "You want to be in this tight loop of new cyber capability drops, you deploy it against your systems, you find the new holes, and ideally, you've managed to automate this, what we call defense factory. That's what we're building internally." "There are ideas, for example, formally verifying all of software, that are possible with AI." @gdb @bhorowitz
2
4
1,099
Can you imagine how much more this would have sold if there were <10% chance of AI wiping out humanity?
San Francisco home sale in Hayes Valley at $3M over asking price
3
567
This is a wild way to report the price of housing increases.
A San Francisco home where prosecutors said a landlord killed his tenant outside the property has been listed for rent at $6,300 a month — nearly double what the slain tenant had been paying. sfchronicle.com/bayarea/arti…
2
543
AI Infrastructure Economics continues today @GoldmanSachs Comparing Google TPU economics below to $NBIS CEO comments today: Nebius capex payback is 22-months, very similar to <2 years for GCP. The difference is that Nebius makes clear it covers capex plus related opex. Long term contacts moved from $10-15M per MW to $20-25M, whereas short term contracts are $40-50M per MW. This is very consistent with the terms that SpaceX has with Anthropic and Google. Nebius has guided capex to $20-25B for 2026.
Google’s TPU Economics 👀 Google GCP CEO Thomas Kurian @ThomasOrTK just said that their payback period on AI servers is <2 years, and half that on their own silicon. 😊 (at @GoldmanSachs today) But it’s not clear whether he means revenue payback or profit/cash-flow payback. 🤔 Let’s assume he means revenue payback, the more realistic of the two in my opinion. 😐 I’d also interpret the statement as payback for just the servers and not the entire datacenter. This means that they monetize TPU clusters at about $30M per facility MW. How do I get there?🤓 Currently, the list price for Ironwood ranges from $5.40/hr (with a 3 yr commitment) to $12/hr on demand. Google says 10MW of IT/pod power can support 9,216 TPU chips. However, this likely does not include the other energy needs (chillers, pumps, transformers, lighting, etc.) So assuming a very efficient modern Google AI data center with 1.10-1.15 PUE (power usage effectiveness), it means about 800-840 Ironwood TPUs per facility MW. Assuming 820 TPUs per MW, at the 3 yr per hour rate of $5.40, that’s $39M/MW at 100% utilisation. I likely wouldn’t lower the utilization assumption because everyone is sold out everywhere. 😮 But I’d lower the rate for these two reasons: (1) TK paired his statement with a mention of 5 yr contracts, and (2) it’s possible volume enterprise contracts (eg for Anthropic) will be at a further discount with a longer commitment. Here there is not list price to use for 5 yr contracts, so extrapolating a diminishing returns curve on the discount from the listed schedule (from on-demand to 1 yr to 3 yrs) suggests $4.50/hr for 5 year commitments (16% discount from the 3 year $5.40/hr rate). Add on top an additional 10% discount for scaled customers and you get $4.08/hr. Thats how I estimate the ~$29M revenue per facility MW of TPUs. TKs statement also has presents some suggestions on margin. Alphabet depreciates their servers and networking equipment over 6 years. With 2yr/1yr payback on AI servers/TPU respectively, it means every $2 server investment produces $1/$2 per year respectively. Depreciation for AI servers generally consume 33% of each revenue dollar ($2 investment divided by 6 years depreciation as percentage of $1 per year) whereas it consumes 17% of each revenue dollar for TPUs ($2 investment divided by 6 years as percentage of $2 per year). So there’s a 16 percentage point gross margin advantage for TPUs. TPU revenue covers the capex in one year for assets that are depreciating over six years. Of course there are operating costs, but once you get long term commitments, it becomes an attractive asset. No wonder Google is spending on capex. This whole post is speculative, but the margin part especially so. 😝 But it will be interesting to build out this model with more information to come. Tell me what I missed!
509
Google’s TPU Economics 👀 Google GCP CEO Thomas Kurian @ThomasOrTK just said that their payback period on AI servers is <2 years, and half that on their own silicon. 😊 (at @GoldmanSachs today) But it’s not clear whether he means revenue payback or profit/cash-flow payback. 🤔 Let’s assume he means revenue payback, the more realistic of the two in my opinion. 😐 I’d also interpret the statement as payback for just the servers and not the entire datacenter. This means that they monetize TPU clusters at about $30M per facility MW. How do I get there?🤓 Currently, the list price for Ironwood ranges from $5.40/hr (with a 3 yr commitment) to $12/hr on demand. Google says 10MW of IT/pod power can support 9,216 TPU chips. However, this likely does not include the other energy needs (chillers, pumps, transformers, lighting, etc.) So assuming a very efficient modern Google AI data center with 1.10-1.15 PUE (power usage effectiveness), it means about 800-840 Ironwood TPUs per facility MW. Assuming 820 TPUs per MW, at the 3 yr per hour rate of $5.40, that’s $39M/MW at 100% utilisation. I likely wouldn’t lower the utilization assumption because everyone is sold out everywhere. 😮 But I’d lower the rate for these two reasons: (1) TK paired his statement with a mention of 5 yr contracts, and (2) it’s possible volume enterprise contracts (eg for Anthropic) will be at a further discount with a longer commitment. Here there is not list price to use for 5 yr contracts, so extrapolating a diminishing returns curve on the discount from the listed schedule (from on-demand to 1 yr to 3 yrs) suggests $4.50/hr for 5 year commitments (16% discount from the 3 year $5.40/hr rate). Add on top an additional 10% discount for scaled customers and you get $4.08/hr. Thats how I estimate the ~$29M revenue per facility MW of TPUs. TKs statement also has presents some suggestions on margin. Alphabet depreciates their servers and networking equipment over 6 years. With 2yr/1yr payback on AI servers/TPU respectively, it means every $2 server investment produces $1/$2 per year respectively. Depreciation for AI servers generally consume 33% of each revenue dollar ($2 investment divided by 6 years depreciation as percentage of $1 per year) whereas it consumes 17% of each revenue dollar for TPUs ($2 investment divided by 6 years as percentage of $2 per year). So there’s a 16 percentage point gross margin advantage for TPUs. TPU revenue covers the capex in one year for assets that are depreciating over six years. Of course there are operating costs, but once you get long term commitments, it becomes an attractive asset. No wonder Google is spending on capex. This whole post is speculative, but the margin part especially so. 😝 But it will be interesting to build out this model with more information to come. Tell me what I missed!
Made with AI
2
1
3
1,497
Watch the followers go up in real time. When I clicked, @johnternus was at 39k followers.
Replying to @TonyW
It’s basically a requirement to be on @X
557
Network effects are really hard to break, especially those filling a very specific use case. In the last few weeks, CEOs of the most influential companies of our time have had to resort to Twitter-turned-X to engage directly. We saw @JensenHuang debut, @finkd make a return, and here @DarioAmodei engage directly with @GavinSBaker on a really important topic. Where would these interactions be if not here? Would we still be relying on op-eds and talking head interviews?
1/2 Thanks Gavin for an especially thoughtful exchange. I don't usually spend much time on social media but I wanted to engage here because it really brings out the heart of an important conversation. First, on regulation, I think that “either concentrate it in the hands of a chosen few companies and politicians via regulation or distribute it widely” is a false choice.  I know that there’s a sort of Silicon Valley shorthand where regulation = regulatory capture = concentration of power, but I’ve always found this to be an overly simplified picture of the world.  Many people outside this bubble think of regulation as something that constrains corporate power and benefits ordinary people.  I don’t necessarily agree with that perspective either, rather I think it’s complicated and really depends on what the “regulation” consists of.  But in particular I think that those in the “regulation = regulatory capture = concentration of power” frame often underrate the decentralizing power of objective and fair institutional processes.  A crude analogy is that the formal court system can sometimes feel stuffy and elitist, but it does a much better job of defending the rights of vulnerable individuals than the alternative, mob justice.  At their best, institutions can vest power in ideas rather than people, and thereby decentralize that power. This is why Anthropic has always made its policy proposals very carefully.  We try very hard to make proposals that disadvantage (slow down) frontier AI companies while *advantaging* smaller competitors.  California’s SB53 (which we supported), and even the much-maligned SB 1047 (which we were ambivalent on), completely exempt any company below a certain amount of revenue or model training costs from being covered at all (it was $500M for SB 53, lower for 1047 but we objected to that).  More recently the testing process we’ve advocated for at CAISI and the White House involves more rigorous tests for frontier models than off-frontier models — something that differentially advantages challengers.  Similarly, the “Pacing the Frontier” letter envisions (or at least Anthropic’s preferred implementation of it envisions) modulating the pace of the very best models while not constraining those who are catching up.  This hurts the business interests of the frontier labs and helps challengers, including open-weights! Overall my view is that AI is *structurally* a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws).  Open-weights do help some with this but are nowhere near a sufficient solution because they simply shift the concentration somewhat to those with the most compute and chips (which are roughly the frontier labs plus maybe hardware providers).  By contrast I think the right “rules of the road” can simultaneously (a) address AI’s cyber/bio/alignment risks, (b) institutionally constrain the power of the frontier AI companies, and (c) leave room for open-weights models while also addressing the specific risks that they bring. BTW I do not think that the events of the last few months have “failed to result in [my] preferred regulatory path”.  The approach that the Trump administration is reported to be taking — pre-deployment testing for frontier models, and also testing of open-weights models when they get closer to the frontier — is one that I am very supportive of, though of course I have to see the details to be sure.  I am also supportive of Demis Hassabis’ ideas around a FINRA-like entity.  This contrasts with six months ago when most of the industry was still pushing for preemption of all state regulation and no apparent federal approach either.
2
5
700
It’s basically a requirement to be on @X
648
Dear Apple: Step One: make iMessage not suck Step Two: dominate the world
2
4
1,136
If you’ve found a solution to this please let me know. Who is the Milken of the AI era?
The scarcest and most in-demand resource on earth right now is GPU financing for sub-IG/unrated offtakers and ASICs This tweet will appeal to approximately 6 people
4
584
Will be at #hotchips2026 at @Stanford again tomorrow and Tuesday. Please say hi and let’s chat chips! 👋🏼
1
4
519
Nice to see @DeepInfra as one of the top fasting growing software companies for this Summer 2026 by @brexHQ 🚀
Our summer fastest-growing vendors list just dropped, and the story is: infrastructure won. Turns out the real AI hype isn't the app, it's what's under it.
2
10
11
2,545
Tony Wang retweeted
I wish every neolab had a professor as cofounder of the level of @jietang and so able to put in perspective their new model release. A great snapshot on the history of scaling laws
Thoughts About Scaling Law Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions. The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed. Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter. Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it. This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count. Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.
7
8
180
39,667
Exciting demo from Claude Science, and this clarification from @pdhsu is very helpful. Basically Claude uses tools and models that humans are already using and getting better output autonomously as opposed to manually.
a nice demonstration of Claude Science, but worth clarifying that the design is not "done by Claude" but by orchestrating tool calls of open-source, task-specific protein design models: PXDesign, RFdiffusion, Genie, BoltzGen, etc I think the direction of LLMs using biology-specific models is a good one!
619
With the news around the acquisition of OpenRouter, nice to see ⁦@DeepInfra⁩ being their largest token provider 🙌 When it comes to inference, there are ways to game the benchmarks, but actual usage shows revealed preferences of developers. @500GlobalVC
1
3
2
1,105
This is a good explanation for why Elon can get a higher price for compute, roughly $30-50 per watt for the short term contracts compared to $20-25 for longer term. The other reason is that he has (and is building) bulk capacity which makes it easier for certain offtakers.
Replying to @CEOinterview
Correct. Most companies like nebius have to do long term contracts to be able to get the debt financing to buy the servers. If they don't have those contacts , they can't get the financing. This gives the customers leverage and they are able to push for a lower price. If Elon has the money or the ability to raise that financing without having a contract in hand, the customers don't have the same leverage and they have to pay him more.
1
595