Pipecat changelog: what's new each month
Every stable Pipecat release summarised by month: the highlights, new features, improvements, fixes and anything you need to act on. 3 months covered; the current month updates daily.
September 2026
4 releases: 1.9.0 → 1.12.0
September brought expanded speech and text service capabilities across multiple providers, improved latency visibility and eval testing, and support for newer SDK versions. The month focused on fine-grained control over transcription, synthesis quality, and evaluation tools.
- SpeechmaticsSTTService now targets Speechmatics Agent STT (/v2/agent) and requires the speechmatics-agent-stt SDK instead of speechmatics-voice.
- Pipecat now supports OpenAI 3 SDK; openai dependency widened to >=1.74.0,<4, which may resolve to version 3.
- Anthropic dependency widened to >=0.49.0,<2 to support anthropic 1 SDK.
Highlights
- Latency breakdown now shows the timeline of where delays occur in user-to-bot exchanges.
- Evaluation scenarios can now play audio files as user input instead of synthesizing text.
- Multiple speech services gained new configuration options: Azure STT silence timeout, DeepgramFlux profanity filtering, OpenAI transcription keywords, and voice effect controls.
- SpeechmaticsSTTService now reconnects with exponential backoff and buffering on connection failures.
- Eval framework now supports classifier-based judging alongside LLM judges, with confidence scores.
New
- LatencyBreakdown.contributions timeline shows the interval breakdown of measured latency.
- AssemblyAISyncSTTService for segmented speech-to-text using AssemblyAI's Sync API.
- BaseAudioResampler.flush() and reset() methods for manual stream boundary marking.
- Eval function call expectations with per-call judge prompts to verify arguments.
- TTSService.pronunciationtransformipa() to define custom pronunciation for specialized terms.
- MOQTransport client-mode reconnect with backoff when relay session drops.
- MOQRunnerArguments.relayurl for dialing a relay by full URL.
- Interruptible frame flag on every frame to control interruption handling.
Improved
- SmallestTTSService now uses continuation API to keep context across text fragments in the same LLM turn.
- SpeechmaticsSTTService reconnects with exponential backoff and audio buffering on errors.
- TTS sentence aggregation now follows Settings.language instead of defaulting to English.
- AzureSTTService segmentation silence timeout is now configurable (100–5000 ms).
- Azure TTS voice parameters support SSML attributes for tuning HD voices.
- Gradium STT can now use server-side end-pointing instead of pipeline VAD.
- JudgeVerdict includes a confidence score (0 to 1) calibrated by judge type.
- Eval judge supports allowcontinue flag to disable continue responses for final-answer-only scenarios.
Fixed
- SpeechmaticsSTTService now properly handles rejected credentials and rejected sessions without reconnect attempts.
- LiveKitParams.audiooutqueuesizems defaults to 1000 ms to preserve existing behavior.
- TTS language updates now apply correctly with TTS settings.
August 2026
3 releases: 1.7.0 → 1.8.1
August brought expanded metrics and usage tracking across services, improved integration with external tools and APIs, and enhanced support for specialized speech recognition features. The month also introduced the Pipecat Context Hub to the CLI for easier agent development workflows.
Highlights
- Added token and usage metrics reporting for LLM, STT, and TTS services to help track resource consumption.
- Integrated live web search capability through KeenableWebSearch service for voice agents.
- Added MCP (Model Context Protocol) client enhancements including automatic tool registration and injection of fixed arguments.
- Introduced Pipecat Context Hub to the CLI with automatic setup for coding agents in Cursor, VS Code, and Zed.
New
- Token usage metrics for AWS Nova Sonic LLM service with speech and text token tracking.
- STT usage metrics across all STT services reporting audio seconds submitted.
- Local avatar image support for LemonSlice transport.
- Inbound SIP DTMF support on LiveKit transport for PSTN calls.
- Google Speech-to-Text v2 adaptation with phrase set biasing for domain-specific recognition.
- Gemini adapter image URL support via external URLs.
- KeenableWebSearch service for live web search and page reading in voice agents.
- Pipecat Context Hub CLI integration with automatic coding agent registration.
Improved
- OpenAI TTS and Whisper-based STT services now accept custom httpx.AsyncClient parameter.
- ElevenLabs real-time STT now supports background audio filtering before transcription.
- Deepgram Flux STT can convert spoken numbers to numeral form in transcripts.
- Google LLM and Vertex LLM services expose Gemini content safety filters directly in settings.
- MCPClient.tools() simplifies tool integration with automatic connection and cleanup.
- pipecat init now includes Context Hub setup for coding agents.
- Context Hub freshness notices alert users when index is stale or built for different Pipecat minor version.
- Directory-based pipecat eval run now recognizes both .yml and .yaml scenario files.
Fixed
- pipecat init no longer repeatedly offers to build a Context Hub index that already exists.
- Stale-index warning now properly appears after Context Hub refreshes the index.
July 2026
2 releases: 1.5.0 → 1.6.0
July brought significant expansions to Pipecat's transport and AI service options, including a new Media over QUIC transport for low-latency connections and support for multiple new speech and language model providers. The month also added important metrics, reasoning capabilities, and flow control features.
- Gemini TTS GenAI backend does not support prompt/style instructions or multispeaker output; these settings are ignored with a warning.
Highlights
- New MOQTransport enables bidirectional audio and RTVI messaging over QUIC with lower latency than WebRTC, with built-in server mode for local development.
- Together AI, NVIDIA, Gemini, Crusoe, and OpenAI reasoning support expand the range of speech and language model backends available.
- Time To First Audio metrics now measure actual audio output latency alongside existing first-byte metrics.
- Pipecat Flows gained NORESPONSE to let functions complete without triggering state transitions or LLM responses.
- Heartbeat timeout events help detect and handle stalled pipeline connections.
New
- TogetherSTTService and TogetherTTSService for real-time speech processing via Together AI WebSocket APIs.
- MOQTransport for Media over QUIC with bidirectional audio and RTVI support, including built-in server mode and TLS configuration.
- Gemini TTS can now use the Google Generative AI backend (google-genai) in addition to Google Cloud.
- Per-sentence synthesis mode and zero-shot audio prompt support in NVIDIA TTS.
- OpenAI reasoning support with configurable effort and summary settings.
- NORESPONSE return value in Pipecat Flows to complete functions without state transitions.
- Heartbeat timeout event handler on PipelineWorker.
- CrusoeoLLMService for Crusoe Cloud's Managed Inference API.
Improved
- Time To First Audio (TTFA) metrics added to TTS services, measuring latency to the first audible sample.
- Gemini TTS backend selection is now automatic based on authentication method, with manual override option.
- Development runner gained MoQ configuration flags for server setup and TLS in local development.
- Eval scenario expectations can now check for absence of events within a time budget.
- HttpOptions parameter added to forward options to Gemini's GenAI client.
Summaries are written automatically from the official release notes (full changelog ↗); check the original notes before relying on a detail. Pipecat: pricing, features and alternatives · All changelogs
