Six questions for any system trusted to act

Where the cost of an incident compounds, the gap between a capable operations system and a re-badged monitoring platform doesn’t show up on a datasheet. These six questions bring it out.

From Agentic AI in Mission-Critical Operations, Appledore Research, September 2026. Worth asking of every vendor you’re looking at, including whoever sent you this.

Unanswered
1 Does the system reason over raw streaming fault, performance and change data from across domains, or does it consume pre-made signals from other monitoring tools at whatever cadence those tools deliver?
2 Is the knowledge graph grounded in a formal ontology that answers consequence questions, or is it a property graph that models connectivity without meaning?
3 Can every autonomous action be traced to the specific relationships and policies that justified it, in an account that would survive a technical audit?
4 Does the system quantify its own confidence and feed it into autonomy thresholds, so that low-confidence situations escalate rather than proceed?
5 Does the knowledge model update itself continuously from operational data, or through consultant-led curation that goes stale?
6 Are there production deployments at scale, with quantifiable, validated outcomes?

How to read the score

Each question scores 0, 1 or 2. Twelve is the maximum.

10–12 — Credible for a mission-critical context

This clears the bar the paper sets, so the work now is checking rather than assessing. Ask for a reference deployment at comparable scale and event rate, and find out whether the outcomes they quote were measured or modelled.

6–9 — Capable in parts

Strong in places, weak in others. The question worth asking is whether each gap is a roadmap item or an architectural limit. Ingestion and correlation can usually be improved. Reasoning over stored data rather than data in motion usually can’t, because the pipeline already made that decision.

0–5 — Built for a more forgiving problem

This is a system built for an environment that can absorb a wrong answer, where a bad correlation costs a ticket and someone is still watching. That’s a perfectly good product, just not this one. Adding a clever layer on top won’t close the gap.

What a weak answer tells you

Scoring 0 or 1 on any question points at something specific. These are the notes the scorecard returns.

  • 1. It’s working off other tools’ conclusions, at their cadence. Its answers can’t be fresher than the slowest thing upstream of it.
  • 2. A connectivity model tells you what sits downstream. It can’t tell you what happens if you act, which is the minimum you need before acting.
  • 3. An explanation written after the decision isn’t the reason for the decision. It won’t survive a post-incident review.
  • 4. With no confidence gate, it’s just as willing to act on its weakest inference as its strongest.
  • 5. A hand-maintained model drifts away from the estate between refreshes, and you won’t see the drift until it produces a wrong answer.
  • 6. How it behaves at demo volumes tells you very little about how it behaves at estate volumes.

The score is a way into a conversation rather than a verdict. The useful part is usually whichever question a vendor can’t answer quickly. Nothing you enter here is stored or sent anywhere.

Aiops readiness scorecard

FutureNet World 2026 – The Self Evolving Knowledge Plane the Missing Link to Autonomous Operations

learn more
Vitria logo
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognizing you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.