The Evolution of Voice as UI:
1. Hardcoded Intents: TTS and STT have been largely solved for years, but early assistants like Siri and Alexa relied on rigid, prescripted intent mapping.
2. Raw Transcription: Voice input just captured verbatim speech filler words, false starts, awkward pauses, and all.
3. LLM Cleaners: Tools like Whisper Flow arrived, using LLMs as a harness to turn messy human speech into structured, coherent text prompts.
4. Multimodal Local Agents: Systems like Codex integrated voice with vision and direct on-device action.
5. Remote Agent Control: With Meta Muse and OpenAI's latest agentic workflows, the compute is entirely remote. Hardware like Meta RayBans or your phone becomes pure audio I/O, an ambient steering wheel for an agent running a full computer in the cloud.