Direct answer: Two separate buyer questions are often mixed together: (1) will the conversation feel paused? and (2) must we migrate our phone system? Latency can be reduced and managed; it is not “eliminated,” and figures only mean something with measurement conditions. Many clinics can keep their published number and existing business VoIP by routing or forwarding into an AI receptionist-but that is not a universal “no telecom change ever” promise.
This page is an evidence-led primer for concerns about AI phone pauses, voice AI latency and deployment with existing phone systems. Deeper latency engineering: reducing voice AI latency. Keep vs migrate architecture: AI vs traditional clinic phone systems. UK landline retirement: PSTN switch-off guide.
Evidence policy: This article does not publish Clero end-to-end “sub-1 second” conversational latency as a product SLA. No approved public measurement pack (method, percentile, sample, conditions) is attached here. Homepage animations are not measurement reports.
Question 1 - Conversational latency (“the pause”)
What callers actually notice
The perceived pause is the gap after the caller stops speaking before agent audio returns. Humans expect a short turn-taking gap; a long void feels like a drop, a misunderstanding or a broken line. Callers repeat themselves; the agent then talks over them.
That gap is a sum, not one knob:
- Telephony / network path
- Speech endpointing (deciding they finished)
- Transcription
- Model inference
- Tool / API calls (availability, patient match, booking)
- Speech synthesis and playback
Tool-heavy turns (diary checks) often dominate FAQ turns. Streaming stacks feel snappier than batch “wait for full transcript → full reply → full audio” pipelines-but physics and APIs still take time. Detail: latency explainer.
Performance depends on conditions
Expect variance with:
- Handset vs softphone, Wi-Fi vs wired
- Forwarding chains and geographic path
- Concurrent load
- Agent configuration and prompt size
- Whether a PMS/tool call runs mid-turn
- Noise, accents and barge-in behaviour
Qualified wording (approved for this page): well-tuned deployments can feel conversationally responsive on simple turns; do not claim “no pause,” “zero silence,” or “sub-second always” without a stated measurement method, percentile, sample size, conditions and approved source. We do not publish such a pack here-so we do not publish the figure.
Demo checklist (buyers)
Run these on your number path (or a path that mirrors production routing):
- Simple FAQ turn - note time from your silence to first agent audio (stopwatch is fine for a feel check).
- Interruption / barge-in - speak over the agent; does it stop cleanly?
- Noisy line - background clinic noise or speakerphone.
- Tool call - ask for availability or a booking that must hit a live system; listen for fillers vs false confirmation.
- Transfer - request a human; confirm destination rings and unanswered behaviour.
- Failed integration - simulate or request a failed write; caller must not hear a fake success.
- Peak-ish load - if possible, two near-simultaneous test calls on the overflow path.
Record pass/fail with date, telephony path and whether tools ran. Re-test after provider or agent changes.
Question 2 - Phone migration vs keep-and-route
Layers (keep these distinct)
| Layer | Job |
|---|---|
| Telephony | Delivers the call (carrier, VoIP, PBX, forwarding) |
| AI receptionist | Conversational handling and approved actions once the call arrives |
| Practice systems | Diary/PMS write-back when a connector exists |
Clero is an AI reception layer, not a replacement phone carrier by default. Partner telephony (for example Dental VoIP Connect / The VoIP Shop) is optional for clinics that want a consolidated stack - not a mandatory lock-in for every deployment.
What “keep existing phones” can mean (realistic)
Often possible:
- Keep the patient-facing number
- Keep the current business VoIP provider when it can overflow, forward or otherwise route selected calls to the voice agent
- Use AI for overflow, out-of-hours or all hours per your rules
May still require telecom work:
- Still on analogue / PSTN (see PSTN switch-off)
- Provider or hardware cannot implement reliable routing, recording or multi-site rules you need
- You choose to port onto a single cloud telephony platform before or with AI
- You want partner-only telephony features beyond AI answering
Always required for AI to answer: a delivery path that reaches the voice agent (commonly forward/overflow to an AI-reachable number, or an equivalent dial-in design). That is routing into the AI stack, which is not the same as “rip and replace every handset and contract”-but it is also not “zero network change in every case.”
Production patterns commonly use forwarding or VoIP rules from clinic lines into the voice platform. Caller-ID masking and CDR recovery can matter for analytics-verify on your PBX. Do not assume “native in-network with no hop” unless that exact architecture is demonstrated for your site.
Absolute claims removed
| Old absolute | Current position |
|---|---|
| “No phone migration” / “never swap provider” | Softened: many keep provider + number via routing; some need telecom changes |
| “No pause” / always sub-second | Softened: latency managed; no universal figure published here |
| “Native integrate - no forwarding impact” | Softened: forwarding/routing is a common path; test audio and CLI |
| Outbound “completely insulated from spam flags” | Softened: routing via verified business numbers can help reputation vs generic diallers; carriers still apply their own filters - no absolute guarantee |
How the two questions interact
A clean VoIP path can improve audio and reduce avoidable network delay; it does not by itself remove tool-call time. Conversely, keeping an old analogue line may force a telephony project and leave you with a worse base path for any voice AI. Sequence honestly: make the phone path workable, then judge conversational feel on that path-including diary tools.
Claims and validation summary
| Claim type | Validation |
|---|---|
| Perceived pause is multi-component | Supported (engineering consensus); see latency guide |
| Clero sub-1s conversational latency as universal fact | Not published here - no approved measurement pack attached |
| Conversation-init / tool webhook timing targets in code | Internal engineering budgets ≠ full mouth-to-ear SLA; not restated as patient-facing figures |
| Keep existing VoIP + number via route/forward | Often true when provider supports it; documented in phone-system / PSTN guides |
| Never need any telecom change | False as absolute - PSTN, locked routing, porting choices |
| Partner VoIP mandatory | False - optional partner stack |
| Spam-risk “eradicated” | Not claimed - qualified only |
Frequently asked questions
What causes the AI pause?
The full stack delay from end-of-speech to agent audio-including tools - not one mythic hop.
Sub-second always?
Not a claim on this page. Measure under stated conditions, including tool turns.
Must we change provider?
Not always. Many route/forward from existing VoIP; some need migration (especially analogue/PSTN or inflexible routing).
Is forwarding enough?
Often used; test quality and caller-ID. See phone-system comparison.
Where next?
Latency guide · PSTN switch-off.
Judge pause feel and phone architecture separately, then test both on a path that mirrors production. If you want a supervised checklist run on your lines, use the CTA below.