AI voice agents reach production limits on concurrency, number control, incident ownership, and compliance scope once volume grows beyond a pilot. Platforms that rent telephony add hops that affect latency, support boundaries, and the subcontractor chain required for regulated workloads.
Voice AI agents succeed in pilots on rented infrastructure because low call volume masks the added transit and divided responsibility. At scale the same architecture surfaces constraints on calls per second, caller ID reputation, and escalation paths that only the carrier relationship can resolve. Teams sizing production therefore evaluate whether the telephony layer sits inside or outside the platform boundary. For a pilot the distinction rarely decides the outcome. For production it often does.
The Two Architectures
Voice AI platforms follow one of two models. Some run the agent pipeline on telephony they rent from a separate carrier or provider. Others own the carrier network and execute the agent logic close to the PSTN handoff.
Rented telephony shortens the path to a first working agent. The platform team avoids carrier relationships, number inventory management, and direct PSTN routing. Vapi, Retell AI, Bland AI and ElevenLabs do not operate as telephony providers: customers bring the number, carrier or SIP trunk, or use one the platform resells. LiveKit and Pipecat are full-code frameworks where outbound calling and most numbers come through a third-party SIP provider. This approach suits prototypes and low-volume inbound services where operational surface remains small and speed to a demo matters more than control of the call path.
Owned telephony keeps the full audio path and control plane inside one organization. Plivo and Telnyx own network infrastructure and can place the agent pipeline next to the carrier edge. The platform can enforce its own pacing limits, manage number reputation directly, and hold the carrier relationship during faults. How audio reaches the PSTN is a separate mechanics question covered in SIP trunking for Voice AI Agent Platforms; this article focuses on what that ownership implies for buyers. The trade-off appears in initial setup time and the need for network expertise, which is why many teams start on rented telephony and revisit the choice only after the pilot succeeds.
What the Extra Hop Costs
Every extra network hop between the caller and the agent adds delay. ITU-T Recommendation G.114 on one-way transmission time states that a one-way delay of 400 ms should not be exceeded for general network planning, and that highly interactive tasks such as many voice calls can be affected by much lower delays. The recommendation also notes that when one-way delay stays below 150 ms, most applications are largely unaffected. Each extra hop consumes part of that budget before speech-to-text, model inference, and text-to-speech even run. Natural turn-taking degrades when cumulative transit leaves little headroom for the agent pipeline itself.
During an incident the support boundary multiplies. A fault in routing or audio quality may require coordination across the agent platform, the rented carrier, and any intermediate transit provider. Resolution time lengthens because no single team owns the entire path. The party that holds the carrier relationship is usually the one that can open a trunk-level ticket, inspect SIP signaling end to end, or force a route change. Teams that rent telephony inherit a multi-party escalation ladder even when the agent logic is healthy.
Regulated workloads add a compliance layer. Every model provider and transcription service in the chain, and any telephony provider that records or stores call audio, becomes a subcontractor that must fit inside a business associate agreement. NIST SP 800-66 Revision 2, published in February 2024 with HHS Office for Civil Rights collaboration, guides covered entities and business associates on assessing and managing risks to electronic protected health information under the HIPAA Security Rule. The platform that rents telephony cannot shorten this subcontractor list on its own. Buyers should map every hop before signing a BAA, and review the platform’s published security posture against that map rather than assuming the agent vendor covers the full path.
Concurrency and Scale Behavior
Concurrency limits appear first when a pilot moves into production. Rented infrastructure often enforces calls-per-second pacing at the carrier level rather than inside the agent runtime. A campaign that exceeds the rented ceiling encounters throttling that the platform team cannot adjust directly. Outbound dialers feel this first: a burst that looked fine at 10 concurrent sessions can queue, reject, or stretch connect times once the rate climbs into production ranges.
Where the limit is enforced matters as much as the number on the slide. If pacing lives in a third-party trunk, the agent team files a ticket and waits. If pacing lives inside an owned network, the same team can raise a ceiling, shift traffic across regions, or isolate a noisy campaign without crossing a vendor boundary. IETF RFC 3261 defines the Session Initiation Protocol used to establish these sessions; capacity and admission control still sit with whoever operates the network that accepts the INVITE.
A low-volume pilot reveals little about this behavior. The same agent logic that handles a handful of concurrent calls may encounter queueing or dropped attempts once the rate reaches hundreds. Production sizing therefore requires load tests against the actual enforcement point rather than extrapolated pilot metrics. Ask every vendor where concurrent-session and calls-per-second limits are applied, who can change them, and what happens to in-flight calls when a ceiling is hit. Those answers separate a demo-ready stack from a production-ready one.
Numbering, Caller ID and Deliverability
Number provisioning speed and control depend on direct carrier relationships. Platforms that rent numbers inherit the provider’s inventory and approval workflows, which can delay rotation or market expansion when a campaign needs fresh local presence. Teams that own the numbering path can search, assign, and rotate inventory without waiting on a second vendor’s queue. Direct control of phone numbers is often the first operational gap teams notice after a successful pilot.
Caller ID reputation forms over time through consistent, low-complaint traffic. When the platform does not originate calls on its own network, reputation signals sit with the rented carrier. Spam labeling or blocked display then requires coordination outside the agent team. The FCC’s guidance on unwanted robocalls frames illegal and spoofed calling as a top consumer-protection priority, which is why downstream carriers and analytics engines aggressively flag suspicious patterns. Number rotation helps only when you can actually obtain and retire numbers quickly.
STIR/SHAKEN attestation levels also trace to the originating carrier. The FCC’s caller ID authentication program requires voice providers to implement the STIR/SHAKEN framework to combat spoofed robocalls. IETF RFC 8224 specifies how authenticated identity is carried in SIP. An STI-GA announcement published by ATIS explains that a call receives full A-level attestation only if the signing service provider can identify the end user and attest to their right to use the number. Higher attestation improves trust signals in many networks, and the decision sits with the originating provider that signs the call: A-level attestation requires that provider to verify the caller’s right to use the number. Branded caller-name display is offered by some carriers and platforms, and availability differs by country. Providers that rely on third parties for technical signing still must make attestation-level decisions themselves under the FCC’s 2025 Call Authentication Trust Anchor rules.
Uptime and What the Number Actually Means
A published 99.99% figure covers only the scope the provider defines. It may exclude the rented carrier segment, model inference latency, or transcription service availability. Teams must examine the status page and published incident history rather than accept the headline number at face value. Ask what components sit inside the SLA, how multi-region failover is tested, and whether audio-quality degradations (not only hard outages) appear in public incident reports.
Redundant carrier connections and failover routing become visible only when the platform owns the network layer. A rented path can still be highly available in practice, but the buyer cannot verify dual-carrier diversity or force a route change without the underlying provider’s cooperation. Status-page transparency on routing events and audio-quality metrics further distinguishes platforms that control the full path from those that report only the application tier.
Voice service providers in the United States must also file in the FCC’s Robocall Mitigation Database, certifying STIR/SHAKEN implementation or detailing alternative mitigation practices. That filing is a compliance artifact, not an uptime guarantee, but it is one more signal that the originating network is a first-class operational concern. Plivo reports 99.99% platform uptime while processing more than 1 billion conversations monthly on its owned network. Treat any vendor’s figure the same way: read the scope notes, then judge the architecture that sits underneath the percentage.
When It Does Not Matter
Low-volume inbound services, internal tools, and single-country prototypes remain viable on rented telephony. The added latency and divided support boundaries stay within acceptable bounds when daily call counts remain small and compliance scope is narrow. Speed to a first working agent is a real advantage in those settings, and pretending otherwise weakens the evaluation. Many teams should stay on rented infrastructure through the pilot and early ramp.
The decision point arrives when any of the following appear: outbound campaigns that test concurrency ceilings, number needs across countries, regulated data flows that require a short subcontractor list, or production incidents that demand direct carrier escalation. At that stage the architecture choice moves from convenience to operational requirement. The build versus buy framework for voice AI agents covers how teams weigh ownership across the full stack; telephony ownership is one axis inside that larger decision, not a mandate for every deployment.
Use a simple threshold rather than a slogan. If you can name a concrete failure mode (throttled outbound, spam-labeled caller ID, multi-party incident delays, or a BAA subcontractor list you cannot defend) that rented telephony cannot fix inside your timeline, revisit the architecture. If you cannot name one, rented telephony is still doing its job.
Frequently Asked Questions
Where do AI-first platforms hit their limits when the network is not theirs?
Rented telephony adds transit hops that increase latency, multiplies support boundaries during incidents, and lengthens the subcontractor chain required for a business associate agreement.
How reliable should a production voice platform actually be?
Expect published figures to cover only the provider’s defined scope. Check whether the provider runs more than one carrier route, how failover is triggered, and whether its status page reports routing and audio incidents.
Which latency and audio-quality problems really matter?
One-way latency targets from ITU-T G.114 leave limited headroom once an extra hop enters the path. Turn-taking in natural conversation degrades when cumulative transit exceeds guidance.
Can an agent run on the numbers we already have?
Platforms that own their carrier network allow direct number import and reputation management. Rented models route through the provider’s inventory and approval process.
Which parts of the stack does production calling genuinely need?
Production requires control over concurrency pacing, STIR/SHAKEN attestation, number reputation, and incident escalation. Owned carrier infrastructure supplies these controls inside a single boundary.
When should a team move from rented to owned telephony for voice AI agents?
The shift becomes relevant once outbound volume tests calls-per-second limits, numbering across countries appears, or regulated workloads require a shorter subcontractor list.
Conclusion
Review any candidate platform against 5 concrete items: concurrency ceiling and where it is enforced, number provisioning path and reputation controls, STIR/SHAKEN attestation level, incident escalation ownership, and the full subcontractor list required for a business associate agreement. Plivo’s owned carrier network and Plivo AI Agents platform place these controls inside one operational boundary. Teams that reach production volume can evaluate the fit directly through Plivo SIP trunking documentation and the broader voice AI platform evaluation guide. Bring the checklist to the next vendor call and score the architecture, not the demo. Teams ready to test the difference can start at signup, or talk to the Plivo team about concurrency ceilings and number provisioning for a specific production footprint.