Flock Cameras Can Read Our Lips Under Favorable Conditions
-----
The Modern Rogue tested whether Flock-style street cameras and current AI tools can reconstruct speech from silent video.
Hosts Brian Brushwood and Justin Young recorded conversations at several distances indoors and outdoors, removed the audio, and ran the footage through publicly available lip-reading software.
Specialized lip-reading tools recovered some phrases when the face was large, well lit, and facing the camera.
--- Key findings ---
Newer pan-tilt-zoom cameras can use strong optical zoom to capture a usable face, a prerequisite for lip reading.
Dedicated lip-reading projects, including Open-Alterego, performed better on close and mid-range indoor footage.
Extra processing helped: cropping to a single speaker, running multiple analysis passes, and using a language model to fill gaps from partial matches.
Context was a major factor. Isolated mouth movements were hard to read; a few confirmed words let the system infer more of the conversation.
The hosts said better cameras, more training data on a specific speaker, and more computing power would likely improve performance.
Their conclusion: current public tools cannot reliably read every conversation from a street camera, but they can extract usable speech from good silent footage under favorable conditions.
Source:
piped.video/watch?v=Ks5Dzbt6β¦