Bain: AI to greatly increase network operator expenses; network re-engineering needed!

Introduction by Bain:

“Over the next three to five years, the operating cost structure used by telecom operators is likely to undergo one of the most significant shifts in decades. As AI agents become embedded across customer care, network operations, software engineering, and enterprise functions, tokens will account for a growing share of operating expenditures.”

Key Points:
  • Unless telcos proactively manage their costs, scaling up AI will simply add expenses to an already-heavy legacy base.
  • An agentic operating model is emerging: a 70-to-30 ratio of legacy to AI costs, with AI automating or augmenting work across processes.
  • Cost traps lurk: cheaper AI models, bigger bills; bolting AI onto legacy processes; and demos mistaken for transformation.
  • Telco leaders can take five key actions today to avoid the traps.

Executive Summary:

Telecom operators are extending AI agents across network operations as they pursue higher levels of autonomy, but the shift could introduce a significant new operating-cost burden unless legacy processes, tooling and organizational structures are retired alongside the automation, according to Bain & Company.

Bain’s warning comes as operators accelerate plans for autonomous networks. TM Forum reported in June that 81% of 80 surveyed operators are targeting Level 4 autonomous networks or higher by 2030, and 20% expect to reach that threshold by 2027.

Under TM Forum’s Autonomous Networks framework, Level 4 moves beyond rule-based or preconfigured automation toward closed-loop, intent-driven decision-making within defined network domains. Current Level 4 work includes deployment of closed-loop operations in production networks, agent-based operating architectures, and metrics intended to quantify the operational and business value of autonomous-network use cases.

Artificial Intelligence and Budgets:
    • Lagging Returns: According to Bain’s Automation and AI Pathfinder Survey, nearly 40% of companies saw AI cost savings land below 10%.
    • Growing Budgets: Despite missing initial savings targets, 90% of these companies are still increasing their AI budgets.
    • Autonomous Agents: Only 7% of companies currently run fully autonomous AI agents in production. Data access remains the top barrier to progress. 

Zero-Click Search and Marketing:
  • AI Summaries: Bain’s research on Zero-Click Search shows that 80% of consumers rely on AI-written results for at least 40% of their searches. 
  • Fewer Clicks: About 60% of searches now end without the user clicking through to another website. This shift reduces organic web traffic by 15% to 25%.

Bain estimates that AI agents and associated token consumption could represent 20% to 30% of a telecom operator’s operating-cost base within the next three to five years, leaving conventional operating costs at 70% to 80%.

As AI spending rises, telcos risk increasing total costs without generating proportional gains in productivity or growth.

Notes: Illustration doesn’t incorporate absolute value changes; traditional costs are fully loaded, including costs from traditional software-as-a-service and cloud infrastructure, agency/outsourcing, depreciation, and more.  Source: Bain estimates

Sources: Wells Fargo (October 2025); Barclays (November 2025); company websites; news and industry reports; Bain analysis.

…………………………………………………………………………………………………………………………………………………………………………………..

However, cost is not the main issue.  The principal risk is that network operators can create a parallel operating model when they layer agentic AI onto established network-operations processes without eliminating the people, software, outsourced functions and infrastructure those processes were designed to support.

In that scenario, AI compute, model inference and agent orchestration become incremental expenses, while legacy network operations centers, monitoring platforms, user-facing software licenses, managed-service arrangements and manual operational handoffs remain largely intact.

Network Re-engineering Required:

Avoiding this outcome requires process re-engineering rather than task-level automation. Bain recommends that operators redesign end-to-end workflows around the work agents assume, including:

  • Reducing manual monitoring, ticket triage and operational handoffs.

  • Reassessing workforce requirements as agents absorb repeatable diagnostic and remediation tasks.

  • Reviewing software, tooling and managed-service contracts when agents begin performing functions previously executed through conventional applications or outsourced processes.

  • Removing redundant operational steps rather than simply automating them.

  • Measuring the cost of an operational outcome, rather than the unit cost of model inference or token consumption.

This distinction is especially important in network operations, where an agent may continuously ingest alarms, correlate events, retrieve telemetry, diagnose faults, select a remediation action, trigger network tools, validate results and escalate exceptions to human operators. Each stage can add cost.

Inference is therefore only one component of the AI operating model. A production-grade agentic workflow may also require orchestration, tool and API calls, runtime evaluation, observability, data storage, policy enforcement, security controls, human-in-the-loop escalation and supporting compute infrastructure.

Bain argues that operators should evaluate the complete cost of the resolved operational event. For a service-affecting incident, that includes the total cost of detection, triage, diagnosis, remediation, validation and any remaining human intervention—not simply the marginal cost of the model invocation.

Closed-Loop Operations and Opex Reduction:

Bain cited Vivo in Brazil as an example of an operator redesigning a complete network workflow around AI-enabled automation rather than applying automation to discrete tasks. As part of Telefónica’s Autonomous Network Journey program, Vivo implemented a self-healing mechanism for its virtualized standalone 5G core.

The implementation monitors network-function performance, detects anomalies, identifies root causes and applies corrective actions automatically. It then validates whether the action restored normal operation and can progress to an additional remediation level when required.

Telefónica said the system correlates events across logical and physical infrastructure and completes the detect-to-resolve sequence without human intervention. For the targeted incidents, the company reported a 30-minute reduction in mean time to resolution.

The deployment also reduces repetitive work and manual intervention, while improving the use of computational resources. Telefónica has not disclosed a monetary estimate of the operating-cost savings associated with the implementation, however.

The significance of the Vivo deployment is architectural as much as operational. It integrates detection, correlation, diagnosis, remediation and verification into a closed-loop workflow. That is materially different from deploying an AI assistant within an otherwise unchanged operating model.

Other operator deployments illustrate the potential conventional opex benefits associated with higher autonomy. In a TM Forum case study, China Mobile reported that intelligent agents helped its network operations center achieve Level 4 autonomy under its self-assessment using TM Forum’s Autonomous Networks Levels framework.

China Mobile reported:

  • More than 30% reduction in backend operations-and-maintenance manpower.

  • More than 5% savings in frontline installation-and-maintenance manpower.

  • An average 30% reduction in mean time to repair for network faults and customer complaints.

An earlier TM Forum autonomous-network case study involving China Mobile reported O&M efficiency improvements of 10% to 20%, service-provisioning time reductions of 30% to 50%, and energy-consumption reductions of 3% to 5% across participating internet data centers and base stations.

The reported figures are operator-reported results published through TM Forum case studies. They demonstrate the possible efficiency gains from closed-loop automation, but they do not eliminate the need to account for AI-specific costs such as inference, orchestration, tool execution, supporting infrastructure and operational governance.

Measure Autonomous Networks by Outcomes:

TM Forum is also developing mechanisms to measure the value generated by autonomous-network deployments. In July, it approved version 2.0 of its Autonomous Networks High-Value Scenarios Effectiveness Indicators guide, intended to help operators quantify the impact of Level 4 autonomous-network scenarios.

This outcome-based approach aligns with Bain’s recommendation. For network operations, operators should move beyond narrow AI measures such as token counts, model cost per query or inference latency. Those measures remain operationally useful, but they do not establish whether an AI deployment improves the economics of network operation.

The most relevant measures include:

Operational measure Why it matters for agentic operations
Cost per resolved incident Captures inference, orchestration, tooling, infrastructure and human escalation across the complete workflow
Mean time to detect Measures whether autonomous monitoring improves fault recognition and event correlation
Mean time to diagnose Shows whether agents reduce time spent isolating root causes across multi-domain infrastructure
Mean time to repair or resolve Measures the operational result that most directly affects service assurance and customer experience
First-time resolution rate Indicates whether autonomous remediation is effective without repeated intervention or escalation
Human-touch rate Shows the proportion of incidents still requiring operations-center intervention
Change failure rate Tests whether autonomous actions improve operations without introducing additional service risk
Cost per successful service activation Applies the same discipline to provisioning and service-fulfillment workflows
Energy per completed workflow Helps assess whether agentic operations create material infrastructure or compute overhead

An operator may accept higher AI spending per workflow if it meaningfully reduces outage duration, truck rolls, customer-impacting incidents, workforce requirements or service-activation delays. Conversely, a deployment that lowers model-inference costs but leaves manual handoffs, duplicate monitoring tools and legacy support structures unchanged may offer limited net operating benefit.

AI Consumption at Telecom Scale:

AT&T has illustrated the potential scale of enterprise AI consumption, although its reported figures span AI workloads across the business and are not limited to autonomous network operations. The operator said in July that it processes an average of 45 billion tokens per day.

AT&T uses an AI gateway to route tasks among models based on cost, latency and expected output quality. The company said the platform can switch models during multi-turn interactions and has reduced costs for certain AI workloads by as much as 90%, producing multimillion-dollar savings.

The operating principle is relevant to telecom network automation: only a minority of tasks require the most capable—and most expensive—models. Routine classification, alarm enrichment, knowledge retrieval, configuration validation and other bounded operational tasks may be suitable for smaller models, purpose-built models or conventional deterministic automation. More capable reasoning models can be reserved for ambiguous, multi-domain or exception-heavy cases.

Bain similarly recommends matching model capability to task complexity rather than applying a single model class across every AI workload. Operators should establish dedicated compute budgets, instrument workflow-level economics and treat inference capacity as an operational resource that requires active governance.

Governing Agentic Network Operations:

Agent behavior itself can become a material source of cost and operational risk. Poorly designed agents may repeatedly transmit large context windows, loop without completing a task, invoke overlapping diagnostic tools or conduct duplicative checks that add token and infrastructure consumption without improving the result.

Bain recommends guardrails that limit both expenditure and runtime. In a telecom network-operations environment, those controls could include:

  • Maximum token, compute and tool-call budgets per incident or workflow.

  • Time limits before an agent must escalate an unresolved task to a human operator.

  • Context-management rules that prevent unnecessary repetition of telemetry, alarms and historical ticket data.

  • Controls to consolidate overlapping diagnostic checks and duplicate agent activity.

  • Policy constraints governing which network changes an agent may propose, execute or validate autonomously.

  • Continuous monitoring for model drift, abnormal agent behavior, spending anomalies and degraded operational outcomes.

  • Explicit business and financial ownership for each production agent and workflow.

The central issue is that autonomous networks will not necessarily lower opex simply because they reduce manual work. Operators must also remove the legacy cost structures that agentic systems replace. Otherwise, AI agents risk becoming an additional layer of expense on top of existing network operations rather than the foundation for a more efficient operating model.

……………………………………………………………………………………………………………………………………….

 

Leave a Reply

Your email address will not be published.

You may use these HTML tags and attributes: <a href="" title=""> <abbr title=""> <acronym title=""> <b> <blockquote cite=""> <cite> <code> <del datetime=""> <em> <i> <q cite=""> <s> <strike> <strong>

*