Skip to main content
Your agent hears the person as PCM16 frames and speaks by writing PCM16 frames back. The transport encodes Opus and sends one packet every 20 ms, like a live microphone, so you can write a whole reply at once and the wire keeps its pace.

Hear the person

Each audio event carries interleaved signed 16-bit samples. The default is 48 kHz stereo; ask for the format your speech-to-text wants when you build the transport.
Pipecat delivers the same audio as InputAudioRawFrames at the pipeline’s audio_in_sample_rate. LiveKit Agents receives 24 kHz mono, the format of a LiveKit room.

Speak to the person

Write PCM16 at any sample rate. writeAudio returns once the audio is queued; waitForPlayout returns when the queue is empty and the last packet has left.
Until the person is receiving your agent’s audio, what you write is held in order and silence goes out. It then plays from the start, so a hello written before the person’s phone is ready is heard whole. To start talking only once the person’s audio has arrived, wait for it first:

Stop talking when interrupted

When the person starts speaking over your agent, drop what has not left yet. clearAudio empties the queue, and a pending waitForPlayout returns at once.
Pipecat and LiveKit Agents do this for you: an interruption clears the transport’s queue, and LiveKit reports the reply as interrupted at the position that actually played.

Mute your agent

setMuted(true) publishes your agent’s mute state to the room as a userUpdate frame, and every participant’s roomState then reports it. Stop writing audio as well.

See also