Sunday roundup: six posts from a week in voice AIOriginalSix posts in one August week: EU watermarking rules, emotion prompts, open source voice agents, a raccoon heist game, and the industry roundup.Voice AI
TIL: Generating Audio Waveform Images With ffmpegOriginalffmpeg generates a waveform PNG from any audio file. Visualise silence gaps, volume levels, and speech patterns without a DAW.TIL
Voice Emotion Control Moves From SSML to PromptsOriginalVoice emotion control is moving to natural language prompts. Kakao's Kanana-o scores 94.50 on the Korean InstructTTSEval benchmark.Voice AI
EU AI Act Voice Watermarking: What TTS Builders Must KnowOriginalEU AI Act voice watermarking rules took effect August 2, 2026. What TTS providers and voice developers must do to stay compliant.Voice AI
WebSocket vs REST TTS APIs for Voice AgentsOriginalChoosing between WebSocket and REST for your streaming TTS API changes your voice agent's latency floor. Here is how to pick the right protocol.Text-to-Speech
Fish Audio's $52M Seed: Open Weights Got Them HereOriginalFish Audio raised a $52M seed on July 28, 2026 with $21M ARR and 8M users. Its new S2.1 Pro model is closed, API-only. What that shift means for TTS.Text-to-Speech
What a Week at SpeechifyOriginalOne week at Speechify: Simba went multilingual, streaming TTS gained timestamps, voice agents learned to switch language mid-call, headers got renamed.Voice AI
MAI-Voice-Flash, Opus 5, and a Red LineOriginalMicrosoft MAI-Voice-2-Flash enters TTS at $15/M. Anthropic ships Opus 5 at half Fable's cost. OpenAI faces red-line questions after Hugging Face.Voice AI
TTS Quality Has No Single Number YetOriginalTTS quality has no single metric. MOS scores cluster, Elo only tests short English clips, and TTS WER measures a different thing than STT.Text-to-Speech
TIL: Test TTS Voices With CurlOriginalOne curl command to compare TTS voice output side by side before you write any integration code. Handy for prototyping or auditing voice quality.TIL
GPT-Live Failed Its First Viral Test in 12 SecondsOriginalOpenAI's GPT-Live launched July 8 and the internet found its weak spot in hours. TikToker Husk broke it with a spelling test, and the full-duplex interruptions areVoice AI
Grok Voice Gets 21 New Voices. The Price Is the PointOriginalxAI added 21 multilingual voices, voice cloning from one minute of audio, and a no-code agent builder to Grok Voice at $0.05 per minute of audio.Voice AI
An Open Source TTS Model That Edits Words After RecordingOriginalViiTorVoice-NAR is an open-source TTS model that can replace individual words inside finished audio without regenerating the surrounding content.Voice AI
TIL: Playing Audio From the Terminal With SoX PlayOriginalplay (from SoX) plays any audio file through your speakers with one command. No media player needed when you are already in the terminal.TIL
Pricing wars are good for developersOriginalThree TTS providers changed pricing in the same week, all in one direction: down. Cheaper voice is good for developers, and the trend is not slowing.Opinion
What surprised me about TTS API design after years of STTOriginalAfter years of speech-to-text, TTS API design broke my assumptions: output parameters everywhere, SSML that fights you, and speech marks nobody mentions.API Design
5 voice AI stories that shaped the start of JulyOriginalThe five voice AI stories that shaped early July: OpenAI's voice reasoning model, ElevenLabs at $22B, Deepgram going multilingual, and Gemini using a computer.Voice AI
Qwen-Audio-3.0-TTS Flash Comes for Real-Time VoiceOriginalAlibaba's Qwen-Audio-3.0-TTS Flash targets the real-time TTS market on price and latency. What it means for voice developers and the API field.AI
TIL: Clone a voice in one API call with SpeechifyOriginalThe Speechify API creates a cloned voice from a 10-second audio sample with a single POST, returning a voice ID that works on any speech endpoint you already use.TIL
Voice cloning goes open source, voice agents go enterpriseOriginalNetEase open-sourced voice cloning from 3 seconds of audio. ElevenLabs partnered with IBM and added SynthID. Coval raised $28M. UK laws are unfit.Voice AI
My Voice Was Cloned Before Lunch on Day OneOriginalDay one at Speechify and my voice was cloned before I finished onboarding. Hearing yourself through a TTS engine is a rite of passage I was not ready for.Voice AI
Voice AI's Real Competition Shifted From Models to PlatformsOriginalTTS model quality converged by June 2026. The real competitive moat shifted to developer experience, platform integration, and compliance tooling.Voice AI
TTS Latency: How Time to First Audio Actually WorksOriginalA deep dive into Time to First Audio, the metric that defines voice agent responsiveness, and what happens in the latency pipeline from text to speech.Voice AI
TIL: Read speech marks from the Speechify API responseOriginalEvery Speechify TTS response includes word-level timing data alongside the audio. Here is how to read speech marks and what you can do with them.TIL
Six stories shaping voice AI in mid-June 2026OriginalMicrosoft MAI-Voice-2, Google live translation, DeepL bought Mixhalo, and open-weight TTS models kept shrinking. Six stories from a busy month in voice AI.Voice AI
Voice data residency decides where your agent runsOriginalVoice data residency decides where speech APIs process audio. Deepgram Australia went live June 17, 2026, shaping how teams pick STT and TTS providers.Voice AI
Streaming TTS: Rethinking the Voice Audio PipelineOriginalStreaming TTS changes what voice applications expect from audio APIs. Time-to-first-byte drops, complexity moves into buffer management and chunk boundaries, andVoice AI
Sunday roundup: four posts on audio debugging and work cultureOriginalFour posts from the week of May 11 in one Sunday roundup: afinfo, speaker diarization, Slack culture on remote teams, and ffmpeg for voice AI.Developer Experience
Sunday roundup: debugging habits and multilingual speechOriginalOne post on voice AI debugging habits, plus Flux going multilingual, AssemblyAI's Voice Agent API and Twilio Conversation Relay, from early May 2026.Voice AI