Agentic AI in Mission-Critical Operations

Where the cost of an incident compounds by the minute, spotting it isn’t the job.

Cumulative impact rises slowly for the first ten minutes then climbs steeply, reaching 100 units at twenty minutes. Resolving at two minutes accrues one unit.
Ten times faster resolution averts a hundred times the accumulated impact. The steepest part of the curve is the part never reached. Source: Appledore Research.

Research by Appledore Research. Read the full report: Agentic AI in Mission-Critical Operations: Setting a New Standard

Mission-critical operations aren’t a harder version of IT operations. They’re a separate category, and what defines them is that the cost of an unresolved incident compounds rather than accumulating, until it reaches a point you can’t come back from.

That one property changes what an operations system has to do. In this Appledore Research paper, Consulting Analyst Robert Curran works through what follows from it: five requirements that a system has to meet at the same time, each turned into a threshold you can test a vendor against, and a set of questions that will tell you fairly quickly whether a platform does what its datasheet says.

The proving ground is large-scale telecom network operations, where the problem has been studied longest. The standard itself is industry-neutral.

Five requirements

They follow from treating real-time resolution as the overriding priority. The difficulty isn’t any one of them. It’s that a mission-critical context demands all five at once.

Real-time detection and rapid resolutionSeconds to minutes, ideally ahead of customer impact — not the hours a cross-team investigation takes.
Multi-domain scaleA serious incident doesn’t stay inside one domain, so a system that sees one domain can’t see the event.
Correlation across domainsDrawing scattered partial signals into a single incident, rather than leaving them as alarms in separate stacks.
Behavioral semanticsKnowing what a relationship means: whether a dependency is full or partial, and how far an impact will actually propagate.
Recommendation and validationChecking the proposed action against the model, the policy and the current state, inside the window the incident allows.

Quantified, not asserted

The operations market has described capability in qualitative terms for years, which is why identical claims get made for systems that are nothing alike. Every requirement in the paper is turned into a number.

CriterionConventional AIOpsMission-critical
Scale of ingestionSampled subsets from a single domainPetabytes per day, distilled to terabytes
DetectionMinutes to hours; batch or sampledStreaming, sub-second to seconds
Signal generationConsumes monitoring tools’ pre-made signalsGenerates its own from raw fault, performance and change
Semantic correlationSingle-domain events over a property graphBehavioral semantics: cause and effect across domains
ResolutionHours to days; human-led investigationMinutes, ideally ~20 minutes ahead of customer impact
Fix validationNone, or a manual checkValidated against operational knowledge before trusted
LearningStatic rules or periodic retrainingContinuous, from operational outcomes

The window is the requirement

Those rows are a closed loop, and every stage of it runs against the clock. A platform that assembles its view on a five-minute batch has lost the incident before it recognised anything.

Time budgetStage & description
Sub-second to seconds1. Ingest and detect — Raw fault, performance and change data arrives and a real problem is picked out of the stream.
Seconds2. Correlate — Related events across every affected domain are drawn into one incident.
Tens of seconds3. Diagnose — Candidate root causes are ranked and the true blast radius established.
Seconds4. Recommend and validate — A remediation is proposed and checked against the model, the policy and the current state.
Seconds to minutes5. Execute and confirm — The action is applied and the system verifies the fault has actually cleared.

Knowledge that arrives after the decision point is indistinguishable from no knowledge at all.

Six questions for any system trusted to act

The last page of the paper is the part most readers use first. These separate a capable system from a re-badged monitoring platform, and they’re worth asking of every vendor you’re looking at.

  1. Does it reason over raw streaming data from across domains, or consume pre-made signals from other tools at whatever cadence those tools deliver?
  2. Is the knowledge graph grounded in a formal ontology that answers consequence questions, or a property graph that models connectivity without meaning?
  3. Can every autonomous action be traced to the relationships and policies that justified it, in an account that would survive a technical audit?
  4. Does it quantify its own confidence and feed that into autonomy thresholds, so low-confidence situations escalate rather than proceed?
  5. Does the knowledge model update itself continuously, or through consultant-led curation that goes stale?
  6. Are there production deployments at scale, with quantifiable, validated outcomes?

Score your own operation against the six questions

Read the paper

Fourteen pages. The full standard, the thresholds behind it, the telecom deployments it was drawn from, and how to introduce a system into an estate that can’t be switched off.

Download the PDF

Robert Curran is a Consulting Analyst at Appledore Research, covering automation and operations across telecom and enterprise IT. This paper was sponsored by Vitria Technology. The analysis and conclusions are Appledore’s.

Cumulative impact of an unresolved incident rising slowly then steeply over twenty minutes.

FutureNet World 2026 – The Self Evolving Knowledge Plane the Missing Link to Autonomous Operations

learn more
Vitria logo
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognizing you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.