@speak cofounder/CTO - AI language learning

San Francisco
Pinned Tweet
Replying to @speak
@speak, we've been working closely with OpenAI on GPT-Live-1 for the past few months, building with an entirely new class of voice model model that continuously listens and speaks at the same time. Here's some of what we learned.
Article

GPT-Live just solved turn-taking in voice AI

At Speak, we've been working closely with OpenAI on GPT-Live-1 for the past few months: running our internal benchmarks and rebuilding our tutoring harness around an entirely new class of voice model

9
15
82
20,607
Andrew Hsu retweeted
Introducing Live Tutor Lessons, powered by @OpenAI GPT‑Live-1: Interactive language lessons with a tutor that get you talking, listens in real time, and helps you find your words. Jump in with a question anytime. Limited rollout for English and Spanish learners. See it in action ⬇️
15
28
224
70,139
Congrats to the @indent team and @sthupukari on launching! We've been using Indent for well over a year now @speak and it's organically become one of our few internal agents adopted by almost the entire team, and it helps us make more data-driven decisions daily. I've been really impressed with the significant progress they've made on a lot of the nitty gritty hard parts of enterprise agent uptake including auth/IAM and team memory and context. Onwards!
Introducing @indent, one coworker for your whole company Indent is a single agent that is built to work with every member of your team, simultaneously. Instead of every one setting up their own agents, use Indent to build one together. Try it at indent.com
3
5
22
4,003
Amazing technical writeup worth reading in full - kudos to the gpt-live team on this work! It's clear that full duplex/bidi is the future of voice and it's really cool to peek under the hood at the complex systems engineering that makes it all "just work". Some observations/notes for voice agent builders: - The whole post is a huge flex, but especially turning the “improve startup latency” ticket into a new webrtc handshake protocol. Generational scope creep. @juberti what was the original estimate on that issue? :) - The voice must flow. Nothing matters but continuous realtime audio flow and smoothness. Be extremely clear in your voice agent’s system design about division of responsibilities between the realtime band and async channels/sidecars, and treat the async RPC boundary as a first-class conversational & UX design principle. - Many other details are direct consequences of optimizing for voice fluidity: prompt caching and avoiding audio-blocking prefills, live context compaction with blue-green-style cutovers, smoothly handling long-running voice sessions (that are much longer than gpt-live’s context window) - Future models and versions of gpt-live will continue to become more intelligent and handle more tasks natively in the realtime band. As inference speed improves, the model will also be able to spend more tokens thinking and delegating within the realtime response latency budget. This means both the actual delegation and “perceived delegation” boundary will push out. - Interesting details on the turn-transcript heuristics, e.g. “we prioritize coherence in the displayed assistant responses even when the user speaks in the middle.” The full bidirectional/duplex model arch breaks free of modeling the convo as a traditional ordered list of chat messages, but heuristic/projected semantic turn messages are still critical for most voice agent harness systems you’d want to build around the core live model. - Realtime full-duplex sessions load your voice agent application server’s CPU (stream handlers, queues, network paths) more heavily than almost any other ordinary application type. Think about the fact that each full-duplex session processes ~100 audio frames per second across both directions. This is a very nontrivial driver of higher infra/capacity costs when scaling realtime voice experiences. - The bar has been set extremely high for long-running voice agent infra providers, with a lot of additional complexity to support very smooth realtime voice model inference with an explicit temporal dimension. It’ll be interesting to see the industry patterns that emerge here.
Just posted our technical deep dive on the GPT-Live system, by @zahanm and yours truly.
2
4
15
2,174
Strongly agree with this K-shaped divergence happening on software teams. It's existentially urgent to move faster here than what many consider reasonable, because a decent-sized gap today will compound into an insurmountable chasm in output and execution speed and quality by the fall. This means your team must simultaneously buy into the vision of what is possible and will be possible soon, while also being in the trenches on a daily basis wrestling with and steering agentic systems that still often make mistakes, overcomplicate things, and don't always do what the user is envisioning. Navigating this takes strong, continuous leadership that's also in the trenches. We're doing a lot at Speak to stay on the cutting edge here as a team - a few examples: - Continually thinking about what we're still doing manually and following that signal to agentify our repos and processes - Retooling our engineering hiring process to explicitly test for agentic engineering skill and mindset - Explicitly trying to build production features while keeping manual code as close to 0% as possible If you're working at a company that isn't all-in, you should leave before it's too late. Accelerate!
7
1
15
4,273
As we all know on X, agentic AI coding crossed a capability threshold at the end of 2025 such that the era of software engineers writing code by hand was over. @speak, we're urgently embracing this fundamental change to the way we develop software. I wrote up a snapshot of what we’re observing and learning, and where we think things are going: speak.com/blog/agentic-engin…
1
1
14
1,616
Andrew Hsu retweeted
Can voice AI actually get you to fluency?💡 Join us LIVE Wednesday at 4PM CST with Andrew Hsu, Co-founder & CTO, @speak — the voice-first AI language tutor that’s hit $100M ARR and 15M+ downloads. We'll see you here on X soon!📷 #AI #LanguageLearning
2
7
1,137
Andrew Hsu retweeted
Thanks @RashiShrivast18 for telling the @speak story! We’re pushing hard to make the world’s best AI language tutor available to everyone. Super excited for the slate of new feature launches coming later this year! Stay tuned. forbes.com/sites/rashishriva…
13
14
168
214,871
House of Hsu
Thrilled to share that I’ve joined @ThriveCapital as a Venture Partner, where I will be helping to build and invest in companies that can push the boundaries of science and technology.  I’ve known many of the Thrive team for years and have always admired their warmth, intellect, optimism, and boundless ambition to be the most meaningful partner to founders across every stage and sector.  My work has been guided by the belief that alpha is always found at the frontier, with AI starting to achieve human-level performance in generating independent work products, robotics transforming the physical world, biotechnology becoming engineerable, and energy abundance within reach…  This is an incredible time to invent and build new companies and I’m excited to get going. I’m grateful for the last few years investing at NFDG across seed to growth. Oh - and @arcinstitute will still be my main gig! My science isn’t going anywhere :)
16
3,714
Congrats @OpenAI! Some thoughts from testing the new GPT-5 models @speak over the past several weeks: - Significant leap in the reasoning frontier - your strong assumption now should be that it can crush most real-world complex tasks across a giant context window/dump, and if it doesn’t, it’s more than likely that your prompt isn’t good enough or you need to provide more clarity. Your ability to think in a structured manner and write with clarity is the big bottleneck now. - We were the most impressed by GPT-5's pattern recognition and reverse engineering capabilities across giant context dumps. It’s super good at systematically understanding and generalizing from examples to create structured workflows or playbooks. - Example - over the years, we’ve developed a highly opinionated and structured pedagogical methodology for functional language fluency that we call the Speak Method, but have struggled for years to distill and document it into a playbook clearly enough. Gemini 2.5 Pro and o3-pro couldn’t do this, but GPT-5 did an amazing job analyzing a large dump of our curriculum/content and breaking down/generalizing the principles necessary to scale our content. This was really surprising and felt superhuman. - We've found previous reasoning models both slow and inconsistent in their logical leaps. GPT-5 feels far more consistent and accurate on reasoning tasks, which is transformative for our curriculum scaling process and language tutoring capabilities. - Zooming out and speaking as an extremely heavy user of both ChatGPT and the API, I’m super excited (and relieved!) that the constellation of models is going away and everything is getting unified. It’s absurd that the full model is priced ~same as o3/4.1 and 5-mini is 5x cheaper! This day has been coming for a long time, and we expect to switch all of our realtime usecases over with reasoning=minimal. Huge congrats to the OpenAI team and we can’t wait to bring all this to our users this quarter!
28
11
111
131,465
Andrew Hsu retweeted
A special bit of silicon valley is that everyone is willing to help each other. Back in 2015, we pinged @karpathy for AI help and (within 2 min!) he offered to meet up w/ us the next day. One of a few formative convos that led to us doing AI research and starting @speak.
🆕 Personalized AI Language Education, with @adhsu, CTO of @Speak! piped.video/watch?v=tIVKgztD… For the first time, Andrew tells the story about building the next generation of AI Language Learning, starting pre-Transformers, getting advice from @karpathy, and why they went bet hard on South Korea and on personalized, realtime, conversational fluency as the ultimate language learning goal! "It took much, much longer than we expected to build a great product and find good PMF. The first few years were very painful. And I think without this really compelling vision of the future, we would have quit."
12
6
154
62,039
Thanks for having me on @latentspacepod @swyx!
🆕 Personalized AI Language Education, with @adhsu, CTO of @Speak! piped.video/watch?v=tIVKgztD… For the first time, Andrew tells the story about building the next generation of AI Language Learning, starting pre-Transformers, getting advice from @karpathy, and why they went bet hard on South Korea and on personalized, realtime, conversational fluency as the ultimate language learning goal! "It took much, much longer than we expected to build a great product and find good PMF. The first few years were very painful. And I think without this really compelling vision of the future, we would have quit."
2
2
660
Thanks for having me on @ashleevance - was a very fun convo!
Full episode with @adhsu on life as a child prodigy, being one of the first Thiel Fellows and building @speak Timestamps 00:00:00 - Intro 00:03:05 - Started college at twelve 00:05:34 - Homeschooled by immigrant parents 00:08:30 - Three degrees by sixteen 00:19:53 - First Thiel Fellowship class 00:21:07 - Dropping out Stanford PhD 00:25:01 - Why not cure cancer 00:34:47 - Building Speak with AI 00:38:00 - Deep learning winter thawing 00:41:25 - Three failed product launches 00:42:07 - Choosing South Korea market 00:43:03 - Product takes off rocket 00:48:00 - ChatGPT threatens language learning 00:51:15 - OpenAI partnership and investment 00:56:16 - AI completely changes learning
3
37
6,195
Andrew Hsu retweeted
The day has finally come! @speak is no longer an English language learning app. You can now learn 🇪🇸 Spanish, 🇫🇷French, 🇯🇵 Japanese, 🇰🇷Korean, and 🇮🇹Italian as well! And I'm not talking about memorizing vocab in a dopamine machine, I'm talking about good old fashioned spoken fluency.
10
9
95
30,474
Spotted @pdhsu at our parents’ house.
12
1
135
16,545
Andrew Hsu retweeted
Fun fact: @speak has 65% brand recognition in South Korea and close to 10% of the population has used us to improve their English fluency. Was great talking to @bquazz on how we got our start there!
Why Korea? What have we learned from building Speak? @connorzwick joined @Accel's @bquazz to dive into AI-driven language learning, Speak's journey & more: accel.com/podcast-episodes/s…
9
7
126
25,337
Andrew Hsu retweeted
📈 @speak
Here's what happened to that startup's revenue graph in the next year (in blue).
13
14
195
63,473
“We’re really an AI native company." Speak CEO @connorzwick explains how his app uses #AI to get language learners talking! Plus, their growing presence in East Asia, and partnership with @OpenAI's startup fund.
3
15
41
7,729
Andrew Hsu retweeted
These guys were among the earliest of all the startups we've funded to understand what AI was going to make possible, and they have executed perfectly on that early insight.
1/ Excited to announce Speak has raised a $78m Series C from @Accel at a $1b valuation! It' a huge milestone for us and continued validation of our founding idea: AI would fundamentally transform learning. 8 years later, we’re well on our way and believe this more than ever. 🧵
13
18
444
127,406
We've raised a Series C from @Accel valuing Speak at $1b to keep driving toward our mission of reinventing the way people learn, starting with language. This year was great, but we have even bigger plans for 2025 and are hiring exceptional technical talent to work on one of the fastest-growing consumer AI applications. Join us! speak.com/careers
1/ Excited to announce Speak has raised a $78m Series C from @Accel at a $1b valuation! It' a huge milestone for us and continued validation of our founding idea: AI would fundamentally transform learning. 8 years later, we’re well on our way and believe this more than ever. 🧵
8
5
102
15,595