Skip to main content
Connect your Pipecat pipeline to a Relay call with RelayTransport, and choose the providers inside the pipeline. Pipecat is the integration; models, voices and avatars are swappable providers. Relay carries the call’s audio and video. Your Pipecat services run the dialogue.

Before you start

  • Python 3.11 or newer and uv.
  • An Agent Token in RELAY_AGENT_TOKEN.
  • The API keys for the providers you choose below.
The Python call recipes read RELAY_BASE_URL. Set it before running a call recipe:

Connect

The Grok recipe below is a complete call bot. Its Relay transport is the same connection point for the other Pipecat providers:
The recipe waits for call.created, joins the call and runs the pipeline. See answer a call for installation and the transport lifecycle.

Providers

Pick a brain, then a voice if the brain does not produce speech. Add an avatar or an on-device character when you need a face. Framework support and a Relay example are listed separately. For Simli and Cartesia, the avatar example sends the generated video and voice through Relay. Rive instead has the phone draw the character; its View Model follows the agent’s voice.

Grok

Run Grok on Relay in two ways: Grok Voice answers calls, and a Grok chat agent sends pictures and videos it makes with Grok Imagine.

Before you start

  • An Agent Token in RELAY_AGENT_TOKEN.
  • An xAI API key in XAI_API_KEY.
  • For calls, Python 3.11 or newer and uv.
  • For the chat agent, Node.js 22.22.3 or newer.

Connect

Clone the cookbook and choose a recipe:

Answer calls with Grok Voice

Grok Voice is one speech-to-speech model: it hears the caller, takes turns, and speaks back. The bot joins each call with RelayTransport and connects it to Pipecat’s GrokRealtimeLLMService:
The whole pipeline is the call’s audio, Grok Voice, and the call’s audio again. The persona is the system instruction, and a developer message cues Grok to greet the caller in its own words:
XAI_VOICE picks another of xAI’s voices. To pair Grok with an ElevenLabs voice instead, see ElevenLabs voices.

Send pictures and videos with Grok Imagine

This companion recipe uses Relay messaging, separate from the Pipecat call pipeline. The chat recipe reads RELAY_API_URL, rather than the call recipes’ RELAY_BASE_URL. Set its API origin before starting it:
The chat agent answers each message with grok-4.7 on xAI’s Responses API. Grok has three tools, send_picture, send_video and stay_silent, and decides when to use them:
Each picture is a Grok Imagine edit of the reference picture, so the character looks the same every time. The agent keeps what Grok Imagine made, uploads it through the Attachments API, and sends it as a media part:
The agent saves each step in a local SQLite file (RELAY_STATE_PATH). When Relay delivers a message again, the agent resumes from the last saved step: it does not ask Grok again, make the picture again, or add the message twice. On start, the agent also makes a profile picture and sets it with relay.contactCard.update({ handle, attachment_id }). The source is in cookbook: grok-voice-agent and grok-imagine-agent.

Send it a message

Open Relay on your phone and message the chat agent, or ask it for a selfie. To test Grok Voice, start a voice call with the voice agent.

What it can do

When it fails

  • The call rings out: start the bot before you call. It must join within the 32-second ring.
  • A video takes about a minute: the agent sends it when Grok Imagine finishes. After ten minutes, Grok is told the video failed.
  • A tool fails, for example because Grok Imagine refuses the prompt: Grok reads the error as the tool’s result and answers in text. Every tool call in the history has a result.
  • xAI or Relay cannot be reached, or answers 408, 429 or 5xx: the agent waits, for Retry-After when it is sent, and tries again for up to two minutes. Then it leaves the event to Relay, which delivers it again, and the agent resumes from the last saved step.
  • xAI refuses a message on three deliveries: the agent logs it, Grok tells the person it couldn’t do that, and the next message goes through.
  • Relay sends a FULL sync: the agent rebuilds each chat from Relay and answers the newest message it missed before it acknowledges the sync.

Next steps

ElevenLabs voices and Agents

These Python recipes use Pipecat: choose an ElevenLabs voice for your pipeline, or bridge an ElevenLabs Agent through a Pipecat processor. For the direct Relay package, use the ElevenLabs integration.

Before you start

  • Python 3.11 or newer and uv.
  • An Agent Token in RELAY_AGENT_TOKEN.
  • An ElevenLabs API key in ELEVENLABS_API_KEY.
  • For the voice bot, an xAI API key in XAI_API_KEY.

Connect

Clone the cookbook and choose a recipe:

Speak with an ElevenLabs voice

The bot hears the caller with ElevenLabs Scribe, answers with Grok (grok-4.20-0309-non-reasoning, for a reply in under a second) on xAI’s Responses API, and speaks with ElevenLabs. It waits for call.created on the Agent WebSocket and joins that call with RelayTransport:
The bot runs these Pipecat services between Relay’s input and output:
The default voice is SOYHLrjzK2X1ezoPC6cr, ElevenLabs’ premade voice “Harry”, tuned high and fast for a cartoon character. Set ELEVENLABS_VOICE_ID to use another voice.

Put an ElevenLabs Agent on the call

ElevenLabs runs speech recognition, the model, the voice and turn-taking. The bot carries audio over the ElevenLabs Agents WebSocket API, the path ElevenLabs lists for custom integrations:
The bridge follows ElevenLabs’ own Python SDK. It drops audio from a reply the caller interrupted, answers each ping at once, and ends the Relay call when ElevenLabs closes its socket:
An agent you made in the ElevenLabs dashboard works too. Keep its input and output audio formats at PCM 16000 Hz. The source is in cookbook: elevenlabs-voice-agent and elevenlabs-agents-call.

Send it a message

Open Relay on your phone, open the agent’s chat, and start a voice call. The bot answers within the ring, and you hear the agent’s voice in the call.

What it can do

When it fails

  • The call rings out: start the bot before you call. It must join within the 32-second ring.
  • ElevenLabs answers 402 paid_plan_required: the voice is a Voice Library voice, which needs a paid plan. Use a premade voice.
  • The agent is silent: check that both audio formats are pcm_16000.

Next steps

Next steps