Voice AI Research @meta super intelligence labs | Previously PlayAI

San Francisco, CA
Come see our work on "Training Language Models under Resource Constraints for Adversarial Advertisement Detection" today at 8:10 AM IST (19:40 PDT) We explore weak supervision, curriculum learning and multi-lingual training. PDF: aclweb.org/anthology/2021.na… @NAACLHLT #naacl2021
2
1
32
Francis Halzen won today's Physics Nobel for IceCube: a cubic kilometer of Antarctic ice turned into a neutrino telescope. I made a 5-minute explainer on how they catch a ghost particle. 🧊👻 It’s my new favourite way of consuming content.
7
27
762
Hosting 2nd edition of Voice research Club San Francisco w/ @shobhitbanga and @rjrshr Attaching a small teaser of my talk. :) Link to join in comments. Visit us while you are here for COLM or SF tech week.
2
2
18
648
LFG 🚀🚀
We're going against OpenAI. We built something miles ahead of Dots & Space, and it's open source. OpenAI's dot cannot: → Work with your whole team. Only you work with your dot → Learn from everyone on your team → Be the same teammate everyone reaches on Slack, WhatsApp or email → Work behind the agentic apps and automations your team needs for the job: a CRM, a ticket desk, a hiring board, a returns tracker and more A Lemma teammate does all of this. And it's open source and free to start. 🧵
1
7
537
Nishant Nikhil retweeted
In <4 weeks you're gonna be so dead sick of "in short" it's going to make your head spin
47
47
1,227
73,567
Found a @Musecases in the wild!
at a YC Paper Club event about neuromorphic computing i didn’t understand most of it but thanks to @Muse i’ll watch the video podcast it created for me later 🥹
7
532
I got my @Muse to generate a full explainer video with deep research. I think this is an emerging way for thoughtful content consumption. It right now has to work hours to come up with this, and these code generated videos need a place in the mid-training stack! @Musecases
2
1
16
736
Nishant Nikhil retweeted
Introducing Parallax: a small, hackable async RL repo with all the bells and whistles of a modern Async RL library. Mixed policy rollouts (pipelineRL), NCCL M2N learner-to-sampler weight streaming, gradient accumulation, and chunked cross-entropy. Core functionality is ~3,000 lines of Python making it easy to read and ~40k tokens – easily fitting into the context window of even local LLMs. Built for 0 friction to try out new ideas across algorithms, systems, and data. Also, obligatory: Parallax Speedrun -- teach Qwen3-0.6B to multiply two five-digit numbers as quickly as possible. Perfect for weekend hacking and autoresearch. [1/8]
5
16
170
7,280
Here's what I have got my @Muse to do in the last week: - Check in and generate boarding pass - Create a game (link in comments - it is v SF centric though) - Create a tasks file, and a live page to track tasks, and research the internet automatically to what can help me close it (see how it got MBFS details) DM for Muse invite. 📥
5
1
22
1,075
Link to the smol game: muse.ai/s/subway-surfers-sil… Very bullish on this being the next interface for computing. And excited for the next updates 👀👀
1
91
It already created an activity artifact for it. Pro tip: add crons in there and connect it to Google drive.
Replying to @nishnik
Setting up the LLM wiki. See how it created the folders. I’ve already started putting resources there. My current implementation for this is an EC2 server, running Tailscale and an app having an interface for creating these notes - and a live Claude code session.
3
339
This is v cool, Kei’s @Muse - Mochi - is subscribing to all the research news he goes through - a custom RSS feed, arxiv daily … and creates this video podcast everyday! 🚀 Mine added word-timestamps too!
try asking Muse to do video podcasts! will need to iterate on the prompting but this is just me asking to make my daily research paper digest as a video
1
5
476
I have been tinkering with muse recently, and all this is crazy unlock: - Setting up Karpathy’s LLM wiki - Asking it to setup and run random repos - Generating a video podcast - Doing research - Financial news as a reel
1
22
2,236
Setting up the LLM wiki. See how it created the folders. I’ve already started putting resources there. My current implementation for this is an EC2 server, running Tailscale and an app having an interface for creating these notes - and a live Claude code session.
1
8
538
But implementation inside muse is also v fast. (Also you can modify the wiki to show activity, maybe geotag, it’s a new way to create your own note app) 5. Here is a video podcast based on my tickers. Pro tip to add this as a cron job inside @Muse (ask it to)
5
91
Guess what’s powering it! ⚡️
Hot off the presses: all @Muse users can ask their Muse to make an audio podcast episode on any topic.
1
9
625
Thank you for hosting us. Best space!
Voice AI event at the @arrayvc office in Dogpatch SF with 70+ researchers, startups, corporates! @voicearena_ai @smallest_AI @Kristopherfloyd @nvidia @Stanford @parthmodi0105
1
16
686
It’s going to be fun. 🎤 We have two research talks.
Super excited to be hosting the first Voice Research Club in San Francisco today with @rjrshr and @nishnik! If you’re around, come join us. Luma link: luma.com/nbaa4d7r
11
570
🚀
Meta has released Muse Voice Transcribe, taking the #1 spot for Final Transcript accuracy on AA-WER Streaming with 3.1% WER at 0.16s after end of speech Muse Voice Transcribe is the first streaming Speech to Text model developed by Meta Superintelligence Labs. Meta states that the model was trained on more than 70 languages, with 25 extensively verified, and supports audio inputs exceeding one hour without required post-processing. It processes audio in 80ms chunks and is available through the Meta Model API, Meta AI for Mac and Muse Code. Key takeaways ➤ Final Transcript: Muse Voice Transcribe achieves 3.1% WER at 0.16s after end of speech. It is more accurate and faster than Cartesia Ink-2 (semantic endpoints) at 3.4% and 0.43s, and more accurate but slower than Cartesia Ink-2 (external endpoints) at 4.0% and 0.07s. It is also more accurate, though slightly slower, than ElevenLabs Scribe v2 Realtime at 3.6% and 0.14s. ➤ First Partial Transcript: The model achieves 3.6% WER at 0.13s, just ahead of ElevenLabs Scribe v2 Realtime on accuracy and latency. It is more accurate and faster than Cartesia Ink-2 (semantic endpoints) at 4.9% and 0.17s, and more accurate but slower than Cartesia Ink-2 (external endpoints) at 4.0% and 0.07s. ➤ Price: Muse Voice Transcribe costs $0.18 per hour, or $3 per 1,000 minutes. This is below Cartesia Ink-2 at $4 and less than half the $6.50 charged for ElevenLabs Scribe v2 Realtime and Deepgram Flux. See more details below ⬇️
12
472
Nishant Nikhil retweeted
Created a very simple benchmark to measure the actual knowledge cutoff of models. It's surprising (or maybe not?) that OpenAI and Anthropic are the only big labs keeping their models fresh. Everyone else - including the chinese labs - is at least 12 months behind
26
93
1,035
254,741