Six questions for any system trusted to act
Where the cost of an incident compounds, the gap between a capable operations system and a re-badged monitoring platform doesn’t show up on a datasheet. These six questions bring it out.
From Agentic AI in Mission-Critical Operations, Appledore Research, September 2026. Worth asking of every vendor you’re looking at, including whoever sent you this.
How to read the score
Each question scores 0, 1 or 2. Twelve is the maximum.
10–12 — Credible for a mission-critical context
This clears the bar the paper sets, so the work now is checking rather than assessing. Ask for a reference deployment at comparable scale and event rate, and find out whether the outcomes they quote were measured or modelled.
6–9 — Capable in parts
Strong in places, weak in others. The question worth asking is whether each gap is a roadmap item or an architectural limit. Ingestion and correlation can usually be improved. Reasoning over stored data rather than data in motion usually can’t, because the pipeline already made that decision.
0–5 — Built for a more forgiving problem
This is a system built for an environment that can absorb a wrong answer, where a bad correlation costs a ticket and someone is still watching. That’s a perfectly good product, just not this one. Adding a clever layer on top won’t close the gap.
What a weak answer tells you
Scoring 0 or 1 on any question points at something specific. These are the notes the scorecard returns.
- 1. It’s working off other tools’ conclusions, at their cadence. Its answers can’t be fresher than the slowest thing upstream of it.
- 2. A connectivity model tells you what sits downstream. It can’t tell you what happens if you act, which is the minimum you need before acting.
- 3. An explanation written after the decision isn’t the reason for the decision. It won’t survive a post-incident review.
- 4. With no confidence gate, it’s just as willing to act on its weakest inference as its strongest.
- 5. A hand-maintained model drifts away from the estate between refreshes, and you won’t see the drift until it produces a wrong answer.
- 6. How it behaves at demo volumes tells you very little about how it behaves at estate volumes.
The score is a way into a conversation rather than a verdict. The useful part is usually whichever question a vendor can’t answer quickly. Nothing you enter here is stored or sent anywhere.
