> ## Documentation Index
> Fetch the complete documentation index at: https://plivo.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Advanced Features for Audio Streaming

> Build voice agents that call your systems, follow a scripted flow, remember returning callers and use your existing tools — on Plivo Audio Streaming.

<Info title="TL;DR">
  * **What:** What a voice agent on Plivo Audio Streaming can do beyond talking.
  * **Watch out:** A tool call that runs longer than a second is silence to the caller unless you cover it.
</Info>

A voice agent that only talks is a recording with extra steps. The agents customers actually deploy look things up, follow a script, remember who is calling and act on what they hear.

Everything below runs inside your agent, above Plivo's WebSocket. The examples use [Pipecat](/docs/voice-agents/audio-streaming/integration-guides/pipecat/overview), which supports Plivo out of the box.

<Note>
  Written against **pipecat-ai 1.11.0**. Pipecat moves quickly and import paths change between releases — check the version you have installed before copying a snippet.
</Note>

***

## Call Your Own Systems Mid-Conversation

The agent decides, during the call, that it needs real data — an order status, an available slot, a customer record — and calls your code to get it.

```python theme={null}
from pipecat.adapters.schemas.direct_function import tool_options
from pipecat.services.llm_service import FunctionCallParams
from pipecat.processors.aggregators.llm_context import LLMContext

@tool_options(cancel_on_interruption=False, timeout_secs=30)
async def get_order_status(params: FunctionCallParams, order_id: str):
    """Look up the status of a customer's order.

    Args:
        order_id: The order number the customer gave you.
    """
    status = await my_backend.lookup(order_id)
    await params.result_callback({"status": status})

context = LLMContext(tools=[get_order_status])
```

The function's parameters and description come from its signature and docstring, so there is no separate schema to maintain.

### Cover the Silence While It Runs

There is no spinner on a phone call. A lookup that takes three seconds is three seconds of dead air, which is when callers say "hello? are you there?" and start hanging up.

```python theme={null}
from pipecat.frames.frames import TTSSpeakFrame

@llm.event_handler("on_function_calls_started")
async def on_function_calls_started(service, function_calls):
    await tts.queue_frame(TTSSpeakFrame("Let me check on that.", append_to_context=False))
```

<Warning>
  `cancel_on_interruption` defaults to `True`. With the default, the caller saying "hello?" during your lookup **cancels the lookup** — so a call that was already too quiet also loses its answer. Set it to `False` for anything the conversation depends on.
</Warning>

***

## Define the Call Flow as Data

A single system prompt holding a whole conversation together stops working the moment the call has real steps. Pipecat Flows models the call as named nodes, each with its own instructions and its own tools, and transitions between them.

The flow is a YAML file the agent loads at runtime.

```yaml theme={null}
initial_node: initial

nodes:
  initial:
    role_message: >-
      You are a friendly insurance agent on the phone. Your responses will be
      converted to audio, so avoid special characters and say amounts in words.
    task_messages:
      - role: developer
        content: Greet the customer and ask how old they are.
    functions:
      - name: collect_age
        transition_to: marital_status

  marital_status:
    task_messages:
      - role: developer
        content: Ask whether they're single or married.
    functions:
      - name: collect_marital_status
        transition_to: quote_results
```

Tools that capture data or do work live in a handlers file. A function that only moves the conversation along needs no code at all — it is a `transition_only` entry in the YAML.

Because the flow is data, one deployed agent can run whichever flow a call needs, and changing what the agent says or where a step leads is a YAML edit rather than a deploy. Write the flow in Python instead when a tool needs JSON Schema constraints an ordinary signature cannot express, or when a node's shape depends on the conversation.

***

## Remember a Returning Caller

The agent's conversation history is a list you can read and write, so a caller who phones back can be greeted with what happened last time rather than starting from nothing.

Save it when the call ends:

```python theme={null}
messages = context.get_messages()
with open(f"conversations/{caller_number}.json", "w") as f:
    json.dump(messages, f)
```

Restore it when the same number calls again:

```python theme={null}
with open(f"conversations/{caller_number}.json") as f:
    context.set_messages(json.load(f))
```

You need the caller's number to key that file, and the stream's `start` event does not carry it — it carries `callId`, `streamId`, `accountId`, `tracks`, `mediaFormat` and `extra_headers`, and nothing else. The caller's number reaches you as `From` on the call webhook. Pass it into the stream yourself:

```xml theme={null}
<Response>
  <Stream bidirectional="true" extraHeaders="caller={{From}}">
    wss://your-server.example.com/stream
  </Stream>
</Response>
```

It then arrives on the `start` event as `extra_headers`, and the agent can open with context instead of "how can I help you?".

<Note>
  Reloading a whole prior conversation grows the context every call. For a caller who rings often, summarize the old history before restoring it rather than replaying it verbatim.
</Note>

***

## Give the Agent Your Existing Tools

If your systems already speak MCP, the agent can use those servers directly rather than having every integration rewritten as a voice-specific function.

```python theme={null}
from mcp.client.session_group import StreamableHttpParameters
from pipecat.services.mcp_service import MCPClient

mcp = MCPClient(
    server_params=StreamableHttpParameters(url=os.getenv("MCP_SERVER_URL")),
    tools_filter=["search_orders", "create_ticket"],
)
tools = await mcp.register_tools(llm)
```

`tools_filter` matters on a phone call. An MCP server exposing forty tools gives the model forty things to consider before every reply, and that shows up as latency the caller hears. Register the handful the agent actually needs.

***

## Add Background Audio

Customers describe an agent with complete silence behind it as sounding artificial. Real people are never in silence — there is always a room around them. Faint ambience under the speech makes the call sound like a person at a desk rather than a recording.

Pipecat mixes it in the output transport, so it plays for the whole call, including the gaps when the agent is not speaking.

```python theme={null}
from pipecat.audio.mixers.soundfile_mixer import SoundfileMixer
from pipecat.pipeline.worker import PipelineParams

params = FastAPIWebsocketParams(
    audio_in_enabled=True,
    audio_out_enabled=True,
    audio_out_mixer=SoundfileMixer(
        sound_files={"office": "office-ambience-8000-mono.wav"},
        default_sound="office",
        volume=0.6,
    ),
)

pipeline_params = PipelineParams(
    audio_in_sample_rate=8000,
    audio_out_sample_rate=8000,
)
```

The `PipelineParams` line is not optional here. The mixer loads the file only if its sample rate equals the pipeline's output rate, and Pipecat defaults to 24 kHz out — so an 8 kHz ambience file is silently skipped unless you set the rate to match.

Change it during the call without rebuilding the pipeline — turn the volume down while the agent reads out a reference number, or switch the ambience off entirely.

```python theme={null}
from pipecat.frames.frames import MixerEnableFrame, MixerUpdateSettingsFrame

await task.queue_frame(MixerUpdateSettingsFrame({"volume": 0.2}))
await task.queue_frame(MixerEnableFrame(False))
```

<Warning>
  The sound file's sample rate must match the pipeline's output rate. If they differ, the file is skipped with a warning and no ambience plays for the entire call — the agent still works, so this is easy to miss. At the 8 kHz telephony default, use an 8 kHz mono file.
</Warning>

Keep the volume low. On a narrowband phone line, ambience that competes with the speech makes the agent harder to understand rather than more natural.

***

## Tune Barge-In

By default the agent stops the moment it hears the caller. On a phone line that is too eager: a cough, an "mm-hmm", or someone talking in the background all cut the agent off mid-sentence. A browser would have removed most of that with local echo cancellation. A phone line does not.

Require a minimum number of real words before treating speech as an interruption.

```python theme={null}
from pipecat.audio.vad.silero import SileroVADAnalyzer
from pipecat.turns.user_start import MinWordsUserTurnStartStrategy
from pipecat.turns.user_turn_strategies import UserTurnStrategies

user_aggregator, assistant_aggregator = LLMContextAggregatorPair(
    context,
    user_params=LLMUserAggregatorParams(
        user_turn_strategies=UserTurnStrategies(
            start=[MinWordsUserTurnStartStrategy(min_words=3)],
        ),
        vad_analyzer=SileroVADAnalyzer(),
    ),
)
```

When an interruption does register, Pipecat sends Plivo a `clearAudio` event, which flushes audio already queued for playback so the caller hears the agent stop rather than talk over them.

***

## Tuning for a Phone Line

These are settings rather than features, but they are the ones that decide whether the agent sounds right on a call.

| Setting                                                                 | Why it matters on a phone call                                                                                                                            |
| ----------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `PipelineParams(audio_in_sample_rate=8000, audio_out_sample_rate=8000)` | Plivo streams 8 kHz. Pipecat defaults to 16 kHz in and 24 kHz out, so it resamples both ways unless told otherwise.                                       |
| `user_idle_timeout=5.0`                                                 | Callers go quiet. Set a timeout so the agent prompts instead of both sides waiting.                                                                       |
| `VoiceFormatter()` on the TTS service                                   | Reads `$42.50` as "forty-two dollars and fifty cents" instead of "dollar sign four two point five zero".                                                  |
| `STTUpdateSettingsFrame` / `TTSUpdateSettingsFrame`                     | Switch language or voice mid-call. `Language` covers the Indian region variants — `HI_IN`, `TA_IN`, `TE_IN`, `KN_IN`, `ML_IN`, `MR_IN`, `BN_IN`, `GU_IN`. |

<Note>
  `SileroVADAnalyzer` supports 8 kHz natively, one of the few audio components built for telephony rates rather than adapted to them. It accepts only 8000 or 16000.
</Note>

***

## Related

<CardGroup cols={2}>
  <Card title="Build with Pipecat" icon="wrench" href="/docs/voice-agents/audio-streaming/integration-guides/pipecat/overview">
    Connect Pipecat to Plivo Audio Streaming
  </Card>

  <Card title="Audio Streaming Reference" icon="book" href="/docs/voice-agents/audio-streaming/concepts/audio-streaming-reference">
    Every WebSocket event, with schemas and types
  </Card>

  <Card title="Transfer to Human Agent" icon="user" href="/docs/voice/use-cases/transfer-to-human-agent">
    Hand a live call to a person
  </Card>

  <Card title="Best Practices" icon="shield-check" href="/docs/voice-agents/audio-streaming/concepts/best-practices">
    Voicemail detection and connection handling
  </Card>
</CardGroup>
