Co-Founder @ Happyverse AI. Ex-Google TPU team & Etched. Stanford MBA/MS grad. Georgia Tech drop out. Dad.

San Francisco, CA
We launched Happyverse 2.0 - a platform for building lifelike, real‑time Confidants (AI avatars + agents) that hold genuine conversations, stay tied to real data, and deliver measurable results for individuals and businesses. producthunt.com/products/hap… Please vote for us on Producthunt and share any feedback you might have! Also, huge thanks to our business partners for your support - @googlecloud (amazing Gemini team in particular - @DynamicWebPaige @triswarkentin @OfficialLoganK and others!), @pipecat_ai (@kwindla and others!), @elevenlabs (congrats with your own launch & summit!) and many others! P.S. Here is a Google Meet conversation with my real-time app.happyverse.ai/CEO Confidant and our CTO, Nicholas 🙂
3
5
28
30,900
What a phenomenal female founder story! From nothing to wired launch in ~6 months! Incredible, huge congrats @aimalysheva and team!! 👏👏
meet @mostik_ai! what happens when you put 12 PhDs in one room for four months? first place on the ARC-AGI leaderboard, which I can't say much about while the competition is still running. and this, which I can. everyone's arguing about whether open models will catch up to frontier models. we think it's the wrong question. here's the one we pose: why does a frontier model have to generate your answer at all, when the only thing you need from it is the reasoning? we do this by enabling models to communicate in latent space. through our protocol, hidden states pass straight from a frontier model into a small one running on your infrastructure -- no text between them, and neither model is fine-tuned. two models from different families, sharing reasoning, both left untouched. how do we know it works? we tested it on a setup where a 753B model reads the problem, and a 4B edge-class model writes the answer. with this approach, we get results 80% as accurate as the frontier model, but at 20x faster performance. we're committed to preventing frontier model lock-in and are already partnering with inference providers to accelerate open-weight adoption. we've done this between 15 of us, in four months, 12 PhDs and a Fields medalist, backed by @generalcatalyst WIRED has the first external account of the company and the work: wired.com/story/russian-star… full writeup, the setup, and all the numbers: mostik.ai/read-more
3
137
Overheard in church - priest's wife made a present to the priest - bought a star online ⭐ what a cool gift idea! Has anyone done this? What are the best "star marketplaces" out there?
120
Max Sapo retweeted
12 PhDs and a Fields Medalist in the mountains arguing about latent communication, and it happened to be my birthday 🎂 taking suggestions on how to top that 😂
51
42
217
26,322
Max Sapo retweeted
If you've wondered where I've been recently, I gave an interview. This is a very difficult time for me. I'm trying to bring my mom home, and I could use your help spreading the word. Please watch, and please share. interview: piped.video/watch?v=JlK2a6yv… website: bringmymomhome.com/
28
67
151
106,588
Max Sapo retweeted
Successfully training models on TPUs has been demonstrated by Anthropic through the past five-plus successful Claude releases. This is positive for the ML community, as Google TPUs continue to gain market share outside of internal Google workloads, giving frontier AI labs a viable alternative for training. 1/4🧵
19
44
518
114,622
Reading what AI doesn't say (this topic genuinely matters to me, and it happens to be my birthday today, so if you feel like it, a repost would be a nice gift 🎂) Something shifting in AI architecture that I don't think gets discussed enough through a safety lens. Models are starting to pass information to each other directly through internal representations, hidden states, activations, and computation vectors (still mostly research right now, but the infrastructure will follow). That matters for how we audit reasoning, since most of what we can currently check comes from what shows up in words. Right now, the main tool we have for auditing AI reasoning is reading chain-of-thought. Korbak et al. (2025) spend a whole paper arguing this window is already "fragile" for single-model reasoning. If models start reasoning through vectors passed model-to-model instead of text, understanding that latent layer becomes the only way to keep any visibility into what's actually happening. Thus I think that what seems underexplored is understanding the math of these latent spaces, and it might be the same research direction as learning how to monitor them. Models appear to converge toward similar geometric structure regardless of architecture (Huh et al. 2024). If that's right, probes for safety-relevant features might generalize: deception patterns, goal representations, misalignment signals studied once and applied across models rather than re-derived per architecture. Whether this holds for safety-critical features specifically is open, so is what adversarially robust latent decoding would even look like. Both feel more urgent than the current research investment suggests.
11
38
77
11,617
Congrats @spurs 🤙
2
51
Max Sapo retweeted
🚀Today we ship @FlyMy_AI Agents. The world's first all-in-one agentic cloud. The modern way to build, integrate, and scale production AI agents. 3 steps to a production agent: 1. Connect your work tools to FlyMy 2. Describe what the agent should do - in text or 5 lines of code 3. Set execution rules: manual, scheduled, or integrated into your backend 4. Done! Agent works on scale! Everything in one place: 800+ MCPs, hundreds of AI models, brain, memory, sandboxes. Stop building from scratch. Stop waiting for infra. Compress 6 months into a day. #1 on @ArtificialAnlys benchmarks. Stable, secure, scalable from day one. Try FlyMy.AI →
14
4
29
6,621
Amin Vahdat is one of the most underrated leaders of our generation. His ability to zoom in / zoom out is just second-to-none.
Amin Vahdat live at Transition-AI 2026 with @CatalystPod. Google's Chief Technologist for AI Infrastructure - the man in charge of Google's $175–185B 2026 CapEx and the one who's said compute capacity needs to 2x every 6 months! Key takeaways: Training → Inference cascade: -Frontier GW training clusters have a ~1–2 year useful life; then capacity cycles to serving -Inference doesn't need GW scale — <100 MW can do useful work on large models -"Entering the age of inference" is now real as agents explode DC footprint -Speed-of-light latency starts to matter as models get faster — geography becomes UX + reliability -"a medium number of medium-sized data centers, augmented with a small number of large ones" (few large training clusters, inference can be mix of 10s to 100s of MW) Reliability reframe -4-nines isn't intrinsic: "we should be thinking about lower reliability power delivery overall" -Do you make the this trade: 4-nines at half capacity, or 2-nines (3.65 days downtime/yr) at 2x capacity? Customers "very often" pick 2x Behind the meter -Google actually prefers grid-connected capacity = BTM is a bridge, not a destination -BTM is about a different latency: time-to-delivery of capacity -Bridge sources: turbines, gas, mobile generation. Permanent: solar, wind, nuclear -Stranded BTM? "I'd love to have that problem." Energy is the limiter Co-design = capability -Google co-designs across Gemini ↔ software ↔ TPU ↔ rack ↔ DC ↔ power ↔ building -A few percent at each interface compounds into real advantage Bottlenecks -Won't force-rank chips/power/labor/EPC: "10am it's labor, noon it's power, 2pm it's chips — every day" -YoY efficiency is real: this year's capacity would've cost ~1.2x last year Building fungibility is dead, purpose-build is back -25-yr buildings vs. 5–6 chip generations. Disk rack vs. GPU rack = ~100x watts/sq ft and widening -Old world: build for fungibility (compute wasn't dominant cost). New world: "this is a GPU building, that's a TPU building" -Density wasn't maxed before because flexibility was worth more. That's flipped One of the most important pods of the year from an AI leader at one of the most important companies at the center of it all! Great job @shaylekann
3
188
Wow, even @elonmusk acknowledges this! #GoTPUs!
Replying to @sundarpichai
TPUs are underrated
59
Met lots of good old friends at @googlecloud #Next2026 and even took a selfie with #TPUv4, my first product launch that I led back in 2022. Can’t believe it’s been 4 years ago - kids grow fast :) Link to a blog in comments below (for those curious to compare what was state-of-the-art then vs now. 🙂)
1
1
55
Met @que_tourist earlier this week. My best celebrity selfie of the year so far 🤗🤙
1
89
1
123
Anyone at #CES in Vegas? Let's catch up!
1
119
Happy New Year, fellas! As per @Skrillex , less phone and more real life 🎄🤙🤗
114
Max Sapo retweeted
The Information’s AI predictions for 2026: - Google will acquire Thinking Machines Lab - OpenAI will launch an automated AI research intern by September 2026, but that intern will fall short of expectations - One of the major AI labs will launch a $1,000-per-month AI agent - Google Gemini will catch up to OAI’s ChatGPT in weekly active users - A breakthrough in “continual learning” will cause Nvidia’s stock to crash Bookmark this tweet. I’ll open it again around next October.
127
128
1,517
236,387
"The Inference Era Has Arrived" :) Credit: #NanoBanana Pro by @GoogleDeepMind
1
115
I asked 4 leading AI models (in "Deep Research" mode) to value #Groq, assuming a role of Investment Banker. With #Nvidia’s ~$20B deal announced: → ~$20B Without it: → ~$9–12B #Claude: “The $20B anchored my entire analysis upward.” That’s anchoring bias. $20B ≠ intrinsic value $20B = what "a monopolist pays to stay a monopolist" (quote from #Claude) Details in this Google Doc: docs.google.com/document/d/1…
136