Telcos don’t have an AI problem. They have a voice estate visibility problem.

By Satish Barot, Co-founder and CTO, Klearcom

AI is increasing telecom interdependence

I’ve spent around twenty years building telephony products, and the last few watching what happens when AI gets added to them. The models are usually fine, but the voice estates around them are not. By voice estate I mean everything a call touches: the number dialed, the carriers carrying it, the menu that answers, the systems behind it, in every country you serve. Most of it you don’t actually own.

NVIDIA’s February 2026 survey of roughly one thousand telecom respondents found that sixty percent of organizations are using or evaluating generative AI, up from forty nine percent in the 2024 edition. The call that once met a gateway and a queue now meets speech recognition, a language model, a synthetic voice, and also identity checks.

Every one of those can be healthy while the call goes wrong. So the useful test is whether you can prove, from outside your own network, that a real customer’s call still ends the way it should.

Why AI amplifies operational complexity

Old fashioned faults announce themselves: a line drops and something logs it. AI is quieter, and it fails in ways worth naming. It gets worse quietly: a model loses accuracy on one kind of request and nothing reports an error. It’s also confidently wrong. WildASR tested seven speech recognition systems on recordings degraded the way phone audio is. When a caller is cut off mid word, by a network delay or by the system deciding they had finished, the models fill in words nobody said. Further down the line that looks like a clean answer. The same benchmark makes a broader point: robustness measured in one language can substantially mispredict behavior in another, so a release validated in one market tells you little about the next.

Speed fails at the edges, not the average. A system that usually answers in under a second, but takes two and a half seconds on one call in ten, feels broken to those callers and healthy on every voice testing team chart. Nor is the result repeatable: EVA-Bench tested twelve systems and found a wide gap between passing once and passing every time, which undoes testing against a single expected answer.

Then the ones nobody looks at: the fallback menu that takes over when the AI gives up is usually the least maintained thing you own, and it runs exactly when things go wrong. None of this turns a light red on a contact center voice dashboard.

The other exposure is a number that anybody can dial. OWASP ranks prompt injection first among risks to applications built on language models, and a voice channel is an open microphone into the instruction path itself: a caller talks the model out of what it was built and told to do. Mitigations exist and they are worth naming. Treat caller speech as untrusted input and keep the transcript out of the context that carries instructions, constrain what the model can emit to a defined set of intents and slots rather than free text, and require confirmation for anything that moves money, changes credentials, or reads account data back.

The customer journey as the real test environment

Take a composite example, assembled from patterns that recur across multi country voice estates rather than drawn from one operator. An operator runs support numbers in fourteen countries behind one AI system. On Friday the team delivers an updated model and a new synthetic voice. Both pass in testing, no issues. By Monday, the share of calls handled without an agent is up by six points, complaints are up in three countries where the release went live first, and every dashboard is green. Nobody knows why.

The new voice made the menus two seconds longer, and nobody adjusted the window in which the system listens for a caller to stop talking. Callers spoke too early, the system heard fragments, filled the gaps, and guessed just confidently enough to keep a person out of the call.

The headline number improved because failures had stopped escalating. Every part did its job, so nothing raised an alarm. What broke was the outcome, and no single part owns the outcome. People given those same recordings transcribed them easily.

Those failures are not hypothetical. Test calls run by my own team over the past two years have found a global biopharmaceutical company whose menu prompts ran into each other faster than callers could answer, a financial services provider whose speech recognition intermittently failed on the word “agent”, so the route to a human worked on some calls and failed on others, and a healthcare data provider whose misrouted number reached an AI assistant that was never meant to answer it. In every case the systems carrying the call reported nothing wrong.

From system monitoring to outcome validation

Monitoring looks inward: are my systems up? Validation looks outward. It places a real call in country and checks the result against what should have happened: is a customer in this country, on this network, right now, getting the right answer?

Monitoring only sees what you own, and these failures hide inside the part causing them. So the two work together: validation tells you something is wrong and where, monitoring tells you why. Three changes follow. Report three numbers where you currently report one: calls the AI resolved correctly, calls it escalated correctly, and calls it failed. Report the slower calls alongside the average ones, and report by country and carrier instead of one global figure.

Figure 1. A single call crosses the voice estate an operator owns and systems it does not. Monitoring covers only the shaded stage. Validation traverses the whole path from the caller’s side.

Continuous testing as an assurance layer

The industry already has vocabulary for this. TM Forum’s autonomous network levels give operators a shared scale for how much of the “operate, assure and optimize” loop runs without people, and the same NVIDIA survey places 88% of organizations at Levels 1 to 3 on a scale of 0 to 5. Moving up that scale rests on evidence about delivered outcomes.

If the journey is what breaks, the journey is what you test. That means dialing the real public number, from a real device, on a real carrier in the country the customer is calling from, rather than looping a call back inside your own data center. It means testing from the outside in, and it means doing it on a schedule, across every country and carrier your customers use, fixed and mobile.

Figure 2. Testing the journey from the outside in. A scheduled call dialed from a real device in the country the customer calls from, scored on the delivered outcome.

Check outcomes, not connections: audio clear both ways, the right menu, key presses registered, the caller reaching the right team with their details intact. Score it across many calls with a library of realistic phrases, accepting a pass rate rather than an exact match, so decline shows as a trend. Treat a model update like a network change: try it on a few numbers first, then pull it back if the pass rate drops.

Several approaches are in this category. Synthetic transaction monitoring from operator owned probes, carrier side test call generation and crowdsourced testing on real handsets all produce outside in evidence, and they trade off differently on country coverage, cost and how closely the test resembles a real customer’s call. The common requirement is that the evidence originates outside the systems being assured, on a call placed in the country where the customer is dialing from.

Practical implications

Keep human escalation as a safety valve, not a number to drive down: cutting it without checking correctness makes the metric look better while the service gets worse. Then give the outcome an owner. Gartner’s June 2025 forecast that over 40% of agentic AI projects will be canceled by the end of 2027 attributes those cancellations to escalating costs, unclear business value and inadequate risk controls. Model capability doesn’t appear anywhere on that list. This is organizational before it’s technical.

These failures are seams between systems that break without producing an error, and monitoring your own equipment cannot see them: the gap sits between what your systems report and what your customer gets. Networks have been tested end to end for decades rather than trusted part by part. Voice estates deserve the same.

References

  1. NVIDIA (2026) State of AI in Telecommunications: 2026 Trends. Fourth annual survey report, published 19 February 2026, based on responses from 1,038 telecom professionals worldwide. Summary of findings: Survey Reveals AI Advances in Telecom.
  2. Tay, G., Ma, W., Lee, J., Tang, Y., Lee, D., Yin, W., Shen, D., Meng, S., Zhu, Y., Li, M. and Smola, A. (2026) Back to Basics: Revisiting ASR in the Age of Voice Agents. Boson AI. arXiv preprint arXiv:2603.25727, 26 March 2026. Introduces the WildASR benchmark. Dataset and code: bosonai/WildASR.
  3. Bogavelli, T., Gauthier Melançon, G., Stankiewicz, K., Bamgbose, O., Riols, F., Nguyen, H.H., Mehndiratta, R., Brin, L.D., Marinier, J., Subramani, H., Madamala, A., Nemala, S.K. and Sunkara, S. (2026) EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents. arXiv preprint arXiv:2605.13841, 13 May 2026, revised 27 May 2026. DOI: 10.48550/arXiv.2605.13841.
  4. OWASP GenAI Security Project (2025) OWASP Top 10 for LLM Applications 2025. OWASP Foundation. Prompt injection is listed first, as LLM01.
  5. Gartner (2025) Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027. Press release, Sydney, 25 June 2025.
  6. TM Forum (2025) Autonomous Networks Framework v2.0.0 (IG1218F). Introductory Guide, published 20 March 2025, TM Forum Approved 9 May 2025. Levels of autonomy 0 to 5. Evaluation criteria for assigning a level are in Autonomous Network Levels Evaluation Methodology (IG1252).

About the author:

Satish Barot is co-founder and CTO of Klearcom. He has spent around twenty years building telephony products and leads the engineering team behind Klearcom’s voice testing platform.

Leave a Reply

Your email address will not be published.

You may use these HTML tags and attributes: <a href="" title=""> <abbr title=""> <acronym title=""> <b> <blockquote cite=""> <cite> <code> <del datetime=""> <em> <i> <q cite=""> <s> <strike> <strong>

*