kmail.at
← learning

technologies · difficulty ◆◆

Pipecat — build real-time voice & multimodal agents

Scaffold a talking agent, wire it to any STT/LLM/TTS, and ship a real project.

2026-08-23 · 10 min read

$ uv tool install "pipecat-ai[cli]"

What it is

Pipecat is an open-source Python framework (BSD-2 license) for building real-time and multimodal AI agents. It is not a model — it is an orchestration framework that connects speech-to-text, an LLM, and text-to-speech into a single agent. You assemble voice assistants, multi-agent systems with handoffs, and multimodal interfaces that see, hear, and speak. It works with any AI provider and any hosting environment; you pay only for the services you choose.

The client/server shape

A typical Pipecat application has a client and a server. The client connects users via a browser, mobile app, or phone. The server runs a Pipecat pipeline that processes audio, runs LLMs, and generates speech in real time. Your hosting — Pipecat Cloud or self-hosted — manages deployment and scales instances to handle concurrent sessions.

The pipeline idea

Every Pipecat agent is a Pipeline of modular processors. A transport connects (WebRTC or WebSocket), a speech-to-text processor transcribes, an LLM thinks, and a text-to-speech processor voices the answer. You control the data flow between them. Because every pipeline is an agent, you can nest them — hand off to specialists, fan out in parallel, or run sidecars.

Why it matters

Real-time voice is awkward to build from scratch: WebRTC plumbing, speech detection, low-latency audio, and provider integration. Pipecat packages that. It orchestrates 100+ AI services with ultra-low latency, and ships Pipecat Flows (structured conversations with state) and Pipecat Cloud (managed hosting) on top of the core runtime.

Example

$ pipecat init quickstart
# Install the Pipecat CLI
uv tool install "pipecat-ai[cli]"
# Scaffold a quickstart project (also writes AGENTS.md + CLAUDE.md)
pipecat init quickstart
# Change into the new project
cd pipecat-quickstart
# The server lives in a subdirectory — go there
cd server
# Copy the environment template, then fill in your keys
cp .env.example .env
# Sync dependencies and run
uv sync
uv run bot.py
# 🚀 WebRTC server starting at http://localhost:7860/client

The CLI scaffolds a server/ subdirectory — the .env.example lives there, so run cp .env.example .env from inside server/.

$ bot.py — wire up the pipeline
# Create the AI services (STT → LLM → TTS)
stt = DeepgramSTTService(api_key=os.getenv("DEEPGRAM_API_KEY"))
tts = CartesiaTTSService(
    api_key=os.getenv("CARTESIA_API_KEY"),
    settings=CartesiaTTSService.Settings(
        voice=os.getenv("CARTESIA_VOICE_ID",
                         "71a7ad14-091c-4e8e-a314-022ece01c121"),
    ),
)
llm = OpenAIResponsesLLMService(
    api_key=os.getenv("OPENAI_API_KEY"),
    settings=OpenAIResponsesLLMService.Settings(
        model=os.getenv("OPENAI_MODEL", "gpt-4.1"),
        system_instruction="You are a helpful assistant in a voice conversation."
    ),
)

# A conversation context, split into user + assistant halves
context = LLMContext()
user_agg, assistant_agg = LLMContextAggregatorPair(
    context,
    user_params=LLMUserAggregatorParams(vad_analyzer=SileroVADAnalyzer()),
)

# Compose the pipeline — this is the whole agent
pipeline = Pipeline([
    transport.input(),      # Receive audio from the browser
    stt,                    # Speech-to-text (Deepgram)
    user_agg,               # Add the user message to context
    llm,                    # Language model (OpenAI)
    tts,                    # Text-to-speech (Cartesia)
    transport.output(),     # Send audio back to the browser
    assistant_agg,          # Add the bot response to context
])

Real code from the official quickstart. The pipeline IS the agent.

$ Tuning voice activity (VAD)
from pipecat.audio.vad.silero import SileroVADAnalyzer
from pipecat.audio.vad.vad_analyzer import VADParams

vad = SileroVADAnalyzer(
    params=VADParams(
        confidence=0.7,   # minimum confidence for voice detection
        start_secs=0.2,   # wait before confirming speech start
        stop_secs=0.2,    # wait before confirming speech stop
        min_volume=0.6,   # minimum volume threshold
    )
)
user_agg, assistant_agg = LLMContextAggregatorPair(
    context, user_params=LLMUserAggregatorParams(vad_analyzer=vad)
)

VAD decides when the user started and stopped talking — tune it for your mic and noise floor.

Common flags

pipecat init <name>
Scaffold a runnable project in one command
uv tool install "pipecat-ai[cli]"
Install the framework + CLI
uv add "pipecat-ai[daily,deepgram,openai]"
Add transports and service extras
uv sync / uv run bot.py
Install deps and run the bot
cp .env.example .env
Copy the environment template before filling keys

← all learning