AI agents for concierge customer experiences

San Francisco
Based in United States
Conversational voice agents need to know when a different person is speaking, but not every background voice should count. We combined speaker embeddings with a post-trained audio-language model to determine when a speaker change matters.
Article

Detecting Relevant Speaker Changes with Audio-Native Models

By @alon_rag and @EveraertDante The problem: knowing when another user joins the conversation In customer support, voice agents need to know not only what was said, but who is participating. If

12
6
62
24,139
Modern TTS models can sound great — and still fail badly on pacing, pauses, and prosody. We adapted DPO + GRPO to flow-matching models to tackle the tail end of TTS behavior:
Article

Teaching Flow-Matching Text-to-Speech Models with RL

Written by Rohan Siva (@_rsiva) and Cyrus Asgari (@cyrusasg) Good demos, unreliable distributions Modern text-to-speech (TTS) systems can sound remarkably natural. But average quality hides the

4
7
42
23,961