Added a new pre-seed YC company, @usemoss to Plushcap: plushcap.com/moss. H/t to @diabhey for telling me about them. They’re working on pushing down the latency threshold on retrieval which is a blocker for conversational agents, especially Voice AI agents 🧵

Apr 20, 2026 · 6:26 PM UTC

4
2
7
984
As with any good early stage DevTool, most of their content is in their docs docs.moss.dev/docs, with a few recent technical blog posts they wrote on their approach to solving issues in this problem space
1
222
Moss competes in the real-time data infrastructure space (a heterogenous grouping, some more focused on scaling protocols versus latency requirements): plushcap.com/competitive/spa… As well as Voice Agents: plushcap.com/competitive/spa…
1
1
33
Would be interesting to see content from these guys on what problems are truly solvable with their edge-based approach, versus what is still intractable, for example 3rd party data services like Clearbit, Crunchbase, ZoomInfo, etc can’t be served this way without pre-caching.
2
365
Lots of room for improvement with voice agents, most of them still suck (pressing “0” on repeat while shouting “talk to operator” is still the default). Check Moss’ homepage out here: moss.dev/
325
Sort replies: Relevant Recent Liked
I've been building with Moss on a Voice AI agent and the speed difference is noticeable in conversation. The gap between "technically fast" and "feels fast enough to not break the flow of speech" is smaller than people think.
I am building a course on shipping production Voice AI agents for a major online education platform. Before I teach it, I want to live on it. So I've been building a voice agent for diabhey.com. Here is my current stack and learnings from v1: • @livekit for realtime transport and the agent framework. Handles sessions, room events, and metrics out of the box. • Silero VAD with min_silence_duration set to 250ms. The plugin default is 550ms. VAD tuning is the single biggest lever on how a voice agent actually feels. 550ms felt sluggish in conversation, 250ms felt natural, but go much lower and you'll cut users off mid-thought. • @DeepgramAI for STT. • @cerebras running Llama 3.1 8B for the LLM. Picked it for raw token throughput. In voice, tokens per second matters more than model size. You're racing a user's attention span, not a benchmark. • @cartesia for TTS. • @usemoss for retrieval. It's an in-process semantic search engine in Rust/WebAssembly, so lookups stay in the agent process with no network hop. If you're shipping voice agents right now, what's moved latency the most for you? Drop it below. I'm collecting real patterns for the course.
1
1
57
Pure value content
2
23
lets gooo!
1
23