Hamming on the conference floor, for everyone who's ever said "but it worked on my test call."
Bring the interruptions, changed minds, and failed handoffs into your tests.
"Supports 12 languages" is a feature list.
"Passed the same booking, correction, handoff, name, and number tests in all 12" is evidence.
Language support is a product claim. Test it like one.
Push-to-talk is a voice-agent test surface.
In a demo, Haloid connected four analog radios to @OpenAI Realtime.
Test clipped openings, overlapping transmissions, and tool arguments.
FinovateFall demo: Hamming Always-On Red Team attacks a bank voice agent, captures severity-scored call evidence, and turns each finding into a regression test.
Sept. 10, 2:44 p.m. ET
Marriott Marquis Times Square
informaconnect.com/finovatef…
The Hamming team is coming to FinovateFall!
Building voice agents in financial services? Come say hi and compare notes on testing and reliability.
Sept. 9–11 · Marriott Marquis Times Square, New York
`elevenlabs agents push --dry-run` shows the config diff.
Our take: a config diff cannot tell you whether the caller still books the appointment.
@elevenlabs agents-as-code need calls-as-tests.
“Your table is booked.”
It wasn’t.
Pipecat’s PhoneBench scores say/do consistency.
Our take: voice-agent evals should grade the promise and the completed action.
QA: same script.
Caller: this version sounds completely wrong.
@DeepgramAI now exposes a calm-to-animated Flux TTS control for the whole Voice Agent session.
Configurable delivery belongs in the regression matrix.
"Was the call successful?" is too vague.
Try:
- identity verified
- slot confirmed
- escalation offered
- downstream state correct
A voice agent is a set of promises per workflow. Write the evals at that level.
Human in the loop is not a security boundary if the model can remove the human.
A new @AWS advisory describes a Strands shell tool where a model-controlled field could bypass consent.
The agent may propose an action. It should never control whether a required approval check runs.
A safe refusal proves one decision. It does not prove the surrounding state stayed clean.
Voice-agent security tests should follow untrusted input across turns, memory, routing, and handoffs.
The real limit of manual voice-agent QA is what teams cannot test at all.
@MavenAGI built its own simulations and smoke tests. Hamming turned them into release infrastructure for concurrency, noise, workflows, and regressions.
“It’s not just time saving, it’s the capability and velocity.” - Dongsheng Wang
Full case study 👇
Voice agents are built to help. That is also what makes them exploitable.
Give an agent access to customer data and permission to act, and a normal conversation becomes an attack surface.
Watch @Sumanyu explain why teams should red-team voice agents:
piped.video/watch?v=cP3cYm-R…
A production voice agent needs more than a model. @OpenAI Presence shows what sits around it: policies, actions, evals, escalation rules, production signals, and controlled rollouts. Our addition: turn failed production calls into regression tests before the next release.
"Voice agents are built to be helpful. The problem is they can be too helpful and cross the boundary."
I sat down with @sumanyu, Founder and CEO of @HammingAI, at our @aiDotEngineer booth. @HammingAI stress-tests and monitors voice agents so the bank or hospital running one catches failures before a customer does.
Fresh after his talk “I Monitored Crime Audio. Voice Agents Scare Me More,” we got into:
0:00 how Sumanyu went from listening to 911 calls at Citizen to founding Hamming AI
0:39 why he's more nervous about voice agents than he was about crime
1:51 the first voice-agent security product
3:16 simulating messy reality: accents, background noise, code-switching
4:11 voice agents testing voice agents
5:47 why taking actions is table stakes, and what makes it dangerous
11:07 hot take: everyone rushed to B2B, more founders should build consumer
Thanks for coming to our @browserbase booth, @sumanyu!
Watch the full episode here:
A voice agent can score 94% in the lab and 58% in a moving car. Only the background noise changed.
Across 10M+ call minutes, ASR accuracy drops 20-40% once noise passes 10dB SNR. If you don't know your test calls' SNR, you don't know how the agent behaves in the real world. 👇