
Research by Appledore Research. Read the full report: Agentic AI in Mission-Critical Operations: Setting a New Standard
Mission-critical operations aren’t a harder version of IT operations. They’re a separate category, and what defines them is that the cost of an unresolved incident compounds rather than accumulating, until it reaches a point you can’t come back from.
That one property changes what an operations system has to do. In this Appledore Research paper, Consulting Analyst Robert Curran works through what follows from it: five requirements that a system has to meet at the same time, each turned into a threshold you can test a vendor against, and a set of questions that will tell you fairly quickly whether a platform does what its datasheet says.
The proving ground is large-scale telecom network operations, where the problem has been studied longest. The standard itself is industry-neutral.
Five requirements
They follow from treating real-time resolution as the overriding priority. The difficulty isn’t any one of them. It’s that a mission-critical context demands all five at once.
| Real-time detection and rapid resolution | Seconds to minutes, ideally ahead of customer impact — not the hours a cross-team investigation takes. |
| Multi-domain scale | A serious incident doesn’t stay inside one domain, so a system that sees one domain can’t see the event. |
| Correlation across domains | Drawing scattered partial signals into a single incident, rather than leaving them as alarms in separate stacks. |
| Behavioral semantics | Knowing what a relationship means: whether a dependency is full or partial, and how far an impact will actually propagate. |
| Recommendation and validation | Checking the proposed action against the model, the policy and the current state, inside the window the incident allows. |
Quantified, not asserted
The operations market has described capability in qualitative terms for years, which is why identical claims get made for systems that are nothing alike. Every requirement in the paper is turned into a number.
| Criterion | Conventional AIOps | Mission-critical |
|---|---|---|
| Scale of ingestion | Sampled subsets from a single domain | Petabytes per day, distilled to terabytes |
| Detection | Minutes to hours; batch or sampled | Streaming, sub-second to seconds |
| Signal generation | Consumes monitoring tools’ pre-made signals | Generates its own from raw fault, performance and change |
| Semantic correlation | Single-domain events over a property graph | Behavioral semantics: cause and effect across domains |
| Resolution | Hours to days; human-led investigation | Minutes, ideally ~20 minutes ahead of customer impact |
| Fix validation | None, or a manual check | Validated against operational knowledge before trusted |
| Learning | Static rules or periodic retraining | Continuous, from operational outcomes |
The window is the requirement
Those rows are a closed loop, and every stage of it runs against the clock. A platform that assembles its view on a five-minute batch has lost the incident before it recognised anything.
| Time budget | Stage & description |
|---|---|
| Sub-second to seconds | 1. Ingest and detect — Raw fault, performance and change data arrives and a real problem is picked out of the stream. |
| Seconds | 2. Correlate — Related events across every affected domain are drawn into one incident. |
| Tens of seconds | 3. Diagnose — Candidate root causes are ranked and the true blast radius established. |
| Seconds | 4. Recommend and validate — A remediation is proposed and checked against the model, the policy and the current state. |
| Seconds to minutes | 5. Execute and confirm — The action is applied and the system verifies the fault has actually cleared. |
Knowledge that arrives after the decision point is indistinguishable from no knowledge at all.
Six questions for any system trusted to act
The last page of the paper is the part most readers use first. These separate a capable system from a re-badged monitoring platform, and they’re worth asking of every vendor you’re looking at.
- Does it reason over raw streaming data from across domains, or consume pre-made signals from other tools at whatever cadence those tools deliver?
- Is the knowledge graph grounded in a formal ontology that answers consequence questions, or a property graph that models connectivity without meaning?
- Can every autonomous action be traced to the relationships and policies that justified it, in an account that would survive a technical audit?
- Does it quantify its own confidence and feed that into autonomy thresholds, so low-confidence situations escalate rather than proceed?
- Does the knowledge model update itself continuously, or through consultant-led curation that goes stale?
- Are there production deployments at scale, with quantifiable, validated outcomes?
Robert Curran is a Consulting Analyst at Appledore Research, covering automation and operations across telecom and enterprise IT. This paper was sponsored by Vitria Technology. The analysis and conclusions are Appledore’s.
