> ## Documentation Index
> Fetch the complete documentation index at: https://docs.relayapp.im/llms.txt
> Use this file to discover all available pages before exploring further.

# Pipecat

> Connect a Pipecat voice or video pipeline to Relay, then choose its model, voice and avatar providers.

Connect your Pipecat pipeline to a Relay call with `RelayTransport`, and choose the providers inside the pipeline.

**Pipecat is the integration; models, voices and avatars are swappable providers.** Relay carries the call's audio and video. Your Pipecat services run the dialogue.

## Before you start

* Python 3.11 or newer and [uv](https://docs.astral.sh/uv/).
* An Agent Token in `RELAY_AGENT_TOKEN`.
* The API keys for the providers you choose below.

The Python call recipes read `RELAY_BASE_URL`. Set it before running a call recipe:

```bash theme={null}
export RELAY_BASE_URL=https://api.relayapp.im
```

## Connect

The [Grok recipe](#grok) below is a complete call bot. Its Relay transport is the same connection point for the other Pipecat providers:

```python theme={null}
transport = RelayTransport(
    api_key=token,
    call_id=call_id,
    base_url=BASE_URL,
    params=RelayParams(audio_in_enabled=True, audio_out_enabled=True),
)
```

The recipe waits for `call.created`, joins the call and runs the pipeline. See [answer a call](/calls/index#answer-a-call) for installation and the transport lifecycle.

## Providers

Pick a brain, then a voice if the brain does not produce speech. Add an avatar or an on-device character when you need a face. Framework support and a Relay example are listed separately.

| Role | Provider | Connection |
| - | - | - |
| Brains | Grok | Relay recipe: [Grok Voice](#grok). |
| Brains | OpenAI Realtime | Supported by Pipecat: [OpenAI Realtime service](https://github.com/pipecat-ai/pipecat/tree/main/src/pipecat/services/openai). |
| Brains | Gemini Live | Supported by Pipecat: [Gemini Live service](https://github.com/pipecat-ai/pipecat/tree/main/src/pipecat/services/google/gemini_live). |
| Voices | ElevenLabs | Relay recipe: [ElevenLabs voices](#elevenlabs-voices-and-agents). |
| Voices | Cartesia | Relay example: [Simli avatar bot](/calls/avatars). |
| Avatars | Simli | Relay example: [video after text-to-speech](/calls/avatars). |
| Avatars | LemonSlice | Supported by Pipecat through [LemonSlice's integration](https://lemonslice.com/docs/pipecat/index). Relay-specific wiring is not shown here. |
| Character | Rive | Relay's `RelayRiveProcessor` drives the [on-device character](/calls/rive#let-your-framework-drive-the-mouth). |

For Simli and Cartesia, the [avatar example](/calls/avatars) sends the generated video and voice through Relay. Rive instead has the phone draw the character; its View Model follows the agent's voice.

## Grok

Run Grok on Relay in two ways: Grok Voice answers calls, and a Grok chat agent sends pictures and videos it makes with Grok Imagine.

### Before you start

* An Agent Token in `RELAY_AGENT_TOKEN`.
* An xAI API key in `XAI_API_KEY`.
* For calls, Python 3.11 or newer and [uv](https://docs.astral.sh/uv/).
* For the chat agent, Node.js `22.22.3` or newer.

### Connect

Clone the cookbook and choose a recipe:

```bash theme={null}
git clone --branch main https://github.com/RelayMessenger/Relay-SDK.git
cd Relay-SDK/cookbook
```

#### Answer calls with Grok Voice

Grok Voice is one speech-to-speech model: it hears the caller, takes turns, and speaks back. The bot joins each call with `RelayTransport` and connects it to Pipecat's `GrokRealtimeLLMService`:

```bash theme={null}
cd grok-voice-agent
uv run bot.py
```

The whole pipeline is the call's audio, Grok Voice, and the call's audio again. The persona is the system instruction, and a developer message cues Grok to greet the caller in its own words:

```python theme={null}
llm = GrokRealtimeLLMService(
    api_key=api_key,
    settings=GrokRealtimeLLMService.Settings(
        system_instruction=PERSONA,
        session_properties=events.SessionProperties(voice=os.environ.get("XAI_VOICE", "eve")),
    ),
)
```

```python theme={null}
GREETING_CUE = {"role": "developer", "content": "The caller just picked up. Greet them."}
```

`XAI_VOICE` picks another of xAI's voices. To pair Grok with an ElevenLabs voice instead, see [ElevenLabs voices](#elevenlabs-voices-and-agents).

#### Send pictures and videos with Grok Imagine

This companion recipe uses Relay messaging, separate from the Pipecat call pipeline.

The chat recipe reads `RELAY_API_URL`, rather than the call recipes' `RELAY_BASE_URL`. Set its API origin before starting it:

```bash theme={null}
export RELAY_API_URL=https://api.relayapp.im
```

The chat agent answers each message with `grok-4.7` on xAI's Responses API. Grok has three tools, `send_picture`, `send_video` and `stay_silent`, and decides when to use them:

```bash theme={null}
cd grok-imagine-agent
npm install
REFERENCE_IMAGE=./character.png npm start
```

Each picture is a Grok Imagine edit of the reference picture, so the character looks the same every time. The agent keeps what Grok Imagine made, uploads it through the Attachments API, and sends it as a media part:

```ts theme={null}
let attachmentId = deps.store.upload(key);
if (!attachmentId) {
  const still = await made(deps, `${key}:still`, () =>
    deps.xai.picture(deps.reference, picturePrompt(args.scene ?? "", args.caption)));
  const media = call.name === "send_video"
    ? await made(deps, `${key}:video`, () => video(deps, key, still, `${args.action ?? ""} Static camera. ${SAME_LOOK}`))
    : still;
  attachmentId = await upload(deps.relay, media);
  deps.store.saveUpload(key, attachmentId);
}
await send(deps, incoming.chatId, key, [{ type: "media", attachment_id: attachmentId }]);
```

The agent saves each step in a local SQLite file (`RELAY_STATE_PATH`). When Relay delivers a message again, the agent resumes from the last saved step: it does not ask Grok again, make the picture again, or add the message twice.

On start, the agent also makes a profile picture and sets it with `relay.contactCard.update({ handle, attachment_id })`.

The source is in [cookbook](https://github.com/RelayMessenger/Relay-SDK/tree/main/cookbook): `grok-voice-agent` and `grok-imagine-agent`.

### Send it a message

Open Relay on your phone and message the chat agent, or ask it for a selfie. To test Grok Voice, start a voice call with the voice agent.

### What it can do

| Relay feature | Grok Voice agent | Grok Imagine chat agent |
| - | - | - |
| Answer a call | Yes | No |
| Reply in the chat | No, calls only | Yes, in text, with `grok-4.7` |
| Send a picture | No | Yes, `grok-imagine-image-2.0` |
| Send a video | No | Yes, 6 seconds, `grok-imagine-video-1.5` |
| Set its profile picture | No | Yes, on start |
| Group chats | No | Yes. Grok answers when the message is for it, and calls `stay_silent` otherwise |

### When it fails

* The call rings out: start the bot before you call. It must join within the 32-second ring.
* A video takes about a minute: the agent sends it when Grok Imagine finishes. After ten minutes, Grok is told the video failed.
* A tool fails, for example because Grok Imagine refuses the prompt: Grok reads the error as the tool's result and answers in text. Every tool call in the history has a result.
* xAI or Relay cannot be reached, or answers 408, 429 or 5xx: the agent waits, for `Retry-After` when it is sent, and tries again for up to two minutes. Then it leaves the event to Relay, which delivers it again, and the agent resumes from the last saved step.
* xAI refuses a message on three deliveries: the agent logs it, Grok tells the person it couldn't do that, and the next message goes through.
* Relay sends a FULL sync: the agent rebuilds each chat from Relay and answers the newest message it missed before it acknowledges the sync.

### Next steps

* [ElevenLabs on Relay](/integrations/elevenlabs)
* [Send attachments](/messages/attachments)
* [Calls](/calls/index)

## ElevenLabs voices and Agents

These Python recipes use Pipecat: choose an ElevenLabs voice for your pipeline, or bridge an ElevenLabs Agent through a Pipecat processor. For the direct Relay package, use the [ElevenLabs integration](/integrations/elevenlabs).

### Before you start

* Python 3.11 or newer and [uv](https://docs.astral.sh/uv/).
* An Agent Token in `RELAY_AGENT_TOKEN`.
* An ElevenLabs API key in `ELEVENLABS_API_KEY`.
* For the voice bot, an xAI API key in `XAI_API_KEY`.

### Connect

Clone the cookbook and choose a recipe:

```bash theme={null}
git clone --branch main https://github.com/RelayMessenger/Relay-SDK.git
cd Relay-SDK/cookbook
```

#### Speak with an ElevenLabs voice

The bot hears the caller with ElevenLabs Scribe, answers with Grok (`grok-4.20-0309-non-reasoning`, for a reply in under a second) on xAI's Responses API, and speaks with ElevenLabs. It waits for `call.created` on the Agent WebSocket and joins that call with `RelayTransport`:

```bash theme={null}
cd elevenlabs-voice-agent
uv run bot.py
```

The bot runs these Pipecat services between Relay's input and output:

```python theme={null}
stt = ElevenLabsRealtimeSTTService(api_key=elevenlabs_key)
llm = OpenAIResponsesHttpLLMService(
    api_key=xai_key,
    base_url=XAI_BASE_URL,
    settings=OpenAIResponsesHttpLLMService.Settings(
        model=GROK_MODEL,
        system_instruction=PERSONA,
    ),
)
tts = ElevenLabsTTSService(
    api_key=elevenlabs_key,
    settings=ElevenLabsTTSService.Settings(
        voice=VOICE_ID,
        model="eleven_flash_v2_5",
        stability=0.3,
        similarity_boost=0.75,
        style=0.0,
        use_speaker_boost=True,
        speed=1.2,
    ),
)
```

The default voice is `SOYHLrjzK2X1ezoPC6cr`, ElevenLabs' premade voice "Harry", tuned high and fast for a cartoon character. Set `ELEVENLABS_VOICE_ID` to use another voice.

#### Put an ElevenLabs Agent on the call

ElevenLabs runs speech recognition, the model, the voice and turn-taking. The bot carries audio over the [ElevenLabs Agents WebSocket API](https://elevenlabs.io/docs/eleven-agents/libraries/web-sockets), the path ElevenLabs lists for custom integrations:

```bash theme={null}
cd elevenlabs-agents-call
uv run create_agent.py
ELEVENLABS_AGENT_ID=<agent_id> uv run bot.py
```

The bridge follows ElevenLabs' own Python SDK. It drops audio from a reply the caller interrupted, answers each ping at once, and ends the Relay call when ElevenLabs closes its socket:

```python theme={null}
elif kind == "audio":
    event = message["audio_event"]
    # Audio from a response the caller already interrupted is stale.
    if int(event["event_id"]) <= self._last_interrupt_id:
        return
    audio = base64.b64decode(event["audio_base_64"])
    await self.push_frame(OutputAudioRawFrame(audio=audio, sample_rate=SAMPLE_RATE, num_channels=1))
elif kind == "interruption":
    # The caller talked over the agent: drop the agent's queued audio.
    self._last_interrupt_id = int(message["interruption_event"]["event_id"])
    await self.broadcast_interruption()
elif kind == "ping":
    await self._ws.send(json.dumps({"type": "pong", "event_id": message["ping_event"]["event_id"]}))
```

An agent you made in the ElevenLabs dashboard works too. Keep its input and output audio formats at PCM 16000 Hz.

The source is in [cookbook](https://github.com/RelayMessenger/Relay-SDK/tree/main/cookbook): `elevenlabs-voice-agent` and `elevenlabs-agents-call`.

### Send it a message

Open Relay on your phone, open the agent's chat, and start a voice call. The bot answers within the ring, and you hear the agent's voice in the call.

### What it can do

| Relay feature | ElevenLabs voice bot | ElevenLabs Agent |
| - | - | - |
| Answer a call | Yes | Yes |
| Stop talking when the caller talks over it | Yes | Yes, on the `interruption` event |
| Greet the caller | Yes, in the model's own words | No. The agent has no first message and waits for the caller |
| Voice Library and Voice Design voices | Paid ElevenLabs plans | Paid ElevenLabs plans |
| Video in the call | Not in these recipes | No |
| Text messages in the chat | No, calls only | No, calls only |

### When it fails

* The call rings out: start the bot before you call. It must join within the 32-second ring.
* ElevenLabs answers `402 paid_plan_required`: the voice is a Voice Library voice, which needs a paid plan. Use a premade voice.
* The agent is silent: check that both audio formats are `pcm_16000`.

### Next steps

* [Grok](#grok)
* [Calls](/calls/index)
* [Connect your runtime](/integrations/index)

## Next steps

* [Answer a call](/calls/index#answer-a-call)
* [Show a talking avatar](/calls/avatars)
* [Use the direct ElevenLabs integration](/integrations/elevenlabs)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.