Heading to Interspeech 2026 in Sydney this week. Every major speech AI lab in one room. The day worth watching: Tuesday, Indigenous Voices Day. Silencio is Full-Day Sponsor. We bring consented, real-world data for the 6,000 languages the internet has never held.
2
7
24
291
Public service announcement. Record a few clips for @silencioNetwork and you get two things: 1. Voice AI that actually hears your language. 2. A private safeword only your model recognises, for when the robots take over. Yours will listen. Everyone else's will not.
7
8
49
1,182
Every major AI lab will need a voice-data partner or an acquisition within 36 months. The reason is structural. In-house collection cannot reach the tail. Purchased data with the wrong provenance is a legal problem. Voice data companies sit where CV data companies did in 2016.
2
8
36
2,052
Mozilla Common Voice v18: ~31,841 hours across 129 languages. That is the largest open, commercial-cleared multilingual speech dataset in the world. Modern frontier speech models train on multiples of that.
2
9
26
569
Which means everything above the Common Voice ceiling is not open, not free, and, more often than not, not documented. That gap is the real training-data question in 2026. What is inside the training set, where did it come from, and can the provider prove it.
1
5
117
For serious multilingual voice AI, the audit trail matters more than the hour count. Silencio builds on the auditable side. Common Voice is the floor. The industry needs the ceiling raised by 10x, provenance intact. Source: Mozilla Common Voice 18. #SpeechAI #MultilingualAI
1
5
92
Time to ride this wave 🌊
60 organizations from @AnthropicAI , @Google , @OpenAI , @Microsoft , and @nvidia to Masakhane, Lelapa AI, AI4Bharat, and BHASHINI just committed to a shared five-year goal: 3.4 billion people who speak underrepresented languages, able to use AI in their own language and voice. That the frontier labs, national governments, and the low-resource-language research community are now aligned around one number tells you how large this problem actually is. No single institution solves it alone. This is why we are building what we are building. A contributor network of ~2.5 million people across ~180 countries, ~350 languages accessible, per-recording consent that travels with the data. Exactly the supply infrastructure a 3.4 billion goal needs. Every recording in this network is a step toward that number. The work you have done is the reason Silencio can support this effort at the scale it needs. Thank you for being early. gatesfoundation.org/ideas/me…
1
4
36
825
Why '1 million hours per language' is the real 2030 target for serious voice AI. Thread. ↓
4
6
30
506
4) The only architecture that scales to 250M hours is a distributed contributor network with fair economics, real-device capture, per-recording consent, and native-speaker transcription at scale. Studios cannot get there. Scraping cannot get there. Ever.
1
8
61
5) The investment thesis in one line: whoever builds the supply infrastructure between now and 2030 owns the next generation of voice AI. Not the model layer. The data layer underneath it.
9
56
What we need asap - please share: 1. Wolof-Speaking Crypto Community 2. Arabic-Speaking Crypto Communities - all Arabic-speaking countries. @silencioNetwork Silencians please share! We have ongoing projects that will need support and receive payments now!
14
19
63
6,928
Sunday note. Voice data has network effects the industry has not fully priced. Once a contributor network holds ~2.5M contributors on demand across ~180 countries, the marginal cost of a new language approaches zero. Second movers cannot catch up. First mover eats the decade.
3
25
479
Every humanoid robot demo I've seen this year is silent. Impressive walking. Impressive lifting. Nobody in the room ever tries speaking to it. We are about to find out why.
1
4
24
481
Computer vision had ImageNet in 2009. AlexNet followed in 2012. The rest of the decade is history. Voice AI has fragmented benchmarks and no comparable inflection point. Yet. The ImageNet moment for voice will be a 10M-hour multilingual corpus. It's being built now.
1
4
24
457
🔥🔥🔥
In two weeks, we're heading to Interspeech 2026 in Sydney to represent every voice that has ever been part of this network. Silver Sponsor. Full-Day Sponsor of Indigenous Voices Day. Every clip you recorded is why we get to be there. Thank you.
1
1
26
304
On 2 August 2026, EU AI Act Article 50 came into force. Providers of generative AI must mark synthetic audio, image, video, and text with machine-readable provenance signals. First hard-law provenance obligation on training-data-adjacent supply anywhere.
1
6
18
555
For AI labs building on scraped or 'mystery-source' audio, this changes the buyer question. Systems already on the market have until 2 December 2026 to comply. That is 12 weeks. Vendor selection just became a documented chain-of-custody question, not a price question.
1
5
97
The market is splitting into pre-provenance and post-provenance shops. Silencio sits on the post side by design. This is not about us. It is about every AI lab needing an auditable chain of custody by December. Source: EU AI Act Article 50. #SpeechAI #AIRegulation
4
67