Cases from the MBA and the Vanguard · MBA module — how far a 24/7 NOC should trust an AI Co-Pilot in a live incident
Subtitle: How far should a 24/7 telecom NOC trust an AI Co-Pilot during a live incident?
02:17. The night shift in a telecom Network Operations Center sees three BTS/cell-unavailable alarms appear in the RAN monitoring tool. Three sites are not unusual enough, by themselves, to explain the cause. They could be unrelated local failures. They could be power. They could be transport. Or they could be the first visible part of a wider incident.
The operator opens the indoor site infrastructure tool. There is no common battery, temperature or power pattern across the three locations. One hypothesis becomes less likely, but nothing is solved. A minute later the transport monitoring tool shows an alarm on a path shared by several of the affected sites. The IP/performance view shows a traffic drop in the same time window. Then two more BTS alarms appear.
The incident has changed shape in four minutes. What first looked like several small RAN events may now be one wider event. The impact is growing, the priority may need to change, and the NOC has to engage the right technical team before too much time is lost.
This is the part of NOC work that is difficult to explain from a procedure. The company already has monitoring systems, incident rules and a Service Management ticketing platform. The ticketing system is not the problem: once the situation is understood, the NOC can record the incident, priority, ownership, timestamps, updates and closure. The harder task comes before that - building a reliable picture while the incident is still moving.
There is no single screen that does that job. Server and database alarms live in one tool. Indoor site power and environmental alarms live in another. Outdoor BTS and cell alarms live in the RAN platform. Transport and optical paths are monitored elsewhere. IP traffic and link performance have their own views. Billing, datacenter/DR, packet-switched, circuit-switched, VAS and fixed-service domains each bring additional alarms, abbreviations and dependencies.
On a normal day, specialization is useful. During a complex incident, it creates a different problem: the operator becomes the integration layer. An experienced person moves between tools, remembers which systems depend on which others, checks whether a planned work is active, searches the approved use case or procedure, speaks to technical teams, and tries to keep one mental timeline of what is known, what is only suspected, and what has already been ruled out.
The BTS scenario is relatively visible because physical sites can often be connected through topology. Higher-priority incidents can be more ambiguous. A customer-facing service may clearly be degraded while alarms appear in IP, transport, PS Core, CS Core, IT/server/database infrastructure or billing at roughly the same time. The NOC sees the impact before it sees the cause. Engaging every team at once creates noise; engaging the wrong team first loses time.
The NOC manager has started thinking about a different operating model. Instead of asking operators to manually connect every signal, what if an internal AI Co-Pilot could read approved, read-only operational feeds from the existing monitoring tools, normalize the events and show one evolving incident picture? It could flag that several BTS alarms share a transport dependency, show that the power picture is clean, retrieve the relevant procedure and similar validated incidents, and prepare a concise explanation of what is known and what still needs to be checked.
The idea sounds attractive precisely because it does not require replacing the existing tools. Each monitoring platform would remain authoritative for its own domain. Service Management would remain the incident record. A correlation layer could handle timestamps, dependencies and repeated patterns; deterministic rules could protect impact/urgency logic; internal knowledge retrieval could bring in procedures and history; a language model could summarize evidence and draft an update in language a human can quickly review.
But the closer the idea gets to a live NOC, the less simple it becomes. A false correlation can look convincing. A language model can explain weak evidence fluently. Stale data can produce a technically neat but operationally dangerous summary. Operators may gradually trust the recommendation instead of checking the source. If the AI suggests the wrong domain during a major outage, who owns the delay? If the team begins relying on it every night, does expertise improve because knowledge is easier to access - or weaken because people stop building the same mental model themselves?
Security creates another boundary. A useful Co-Pilot needs enough context to understand topology, status, procedures and incident history, but it should not receive unrestricted production access or credentials. Even read-only integration across many operational systems expands the surface that must be governed and audited.
The NOC manager can imagine three paths. The NOC could keep the current model and improve procedures, training and cross-tool dashboards without introducing generative AI. It could run a narrow read-only pilot on one RAN/transport use case, where the AI only correlates and explains. Or it could aim for a broader cross-domain assistant that also recommends incident classification, suggests the next team to involve and drafts NEW/UPDATE/SOLVED communication for human approval.
Each path solves a different problem and creates a different risk. The narrow pilot is easier to control but may prove too little. The broader assistant could create real value during high-priority incidents, but this is exactly where a confident mistake would matter most. Keeping everything manual preserves clear accountability but leaves the operator as the human bridge between an increasing number of specialized systems.
The decision is not whether AI can read alarms. The decision is where, inside a live incident, AI should be allowed to influence judgment.
Before asking the technical, security and governance teams to support a pilot, the NOC manager has to put one proposal on the table. How much of the incident-understanding process should the NOC hand to an AI Co-Pilot - and what must remain unmistakably human?
Open the hour. Put the 02:17-02:22 BTS sequence on the table and ask: “At 02:22, what would you already trust an AI system to do - and what would you refuse to delegate?” Do not start with technology architecture; start with the decision pressure.
Ask first. Ask someone who has led a 24/7, safety-critical or high-availability operation. Then ask someone from a non-technical leadership role to challenge whether the proposed boundary is understandable and governable.
The fact that could change the room. The most important evidence would be replay performance on validated historical incidents: not only how often the Co-Pilot finds the right correlation, but how often it produces a confident false correlation or misses a critical signal in high-priority cases. That error profile would materially change the acceptable scope.
The current NOC is already networked in the technical sense, but its knowledge is fragmented across specialized systems. The NEO turn would be to move from a human acting as the integration layer to a human acting as the orchestrator: monitoring platforms, procedures, historical knowledge and AI remain distinct but are connected around one incident picture.
The exponential element is not “more automation”. It is the speed at which context can be assembled when AI is allowed to synthesize many signals at once. The orchestration challenge is governance: preserving deterministic rules, authoritative data sources and human accountability while using AI to reduce the time from alarm to clarity. The desired shift is not from human to machine; it is from fragmented tools to a deliberately orchestrated human-AI operating system.
“The goal is not to let AI run the NOC. The question is whether it can help the NOC reach clarity before complexity wins.”
A live case: every round can be improved, and the author's feedback is the next one.
The same two questions, answered twice: first without the mentor's corpus, then from it — his decision doctrine, Vanguard Leadership, the HAI5 framework, the chapter on Heuer and the Task-to-Agent protocol.
Without the mentor's corpus
1. Yes, pilot now, but let the first phase touch no live decision. Begin with replay, not with the night shift. Take a year of validated RAN and transport incidents from Service Management, replay their alarms to the Co-Pilot in the order they arrived, and compare its picture, minute by minute, with what each incident turned out to be. Microsoft's RCACopilot, built for nearly this job, gathers diagnostic data, predicts the root-cause category and writes an explanatory narrative. Judged on a year of Microsoft's own incidents, it was right at best about three times in four. Borrow nobody's accuracy figure. An independent check at Michigan Medicine, published in 2021, found that a sepsis-warning model used in hundreds of American hospitals did substantially worse than its developer had reported: across 38,455 hospitalisations it missed two-thirds of the sepsis cases while raising alerts on 18% of all of them. So the narrowest pilot is:
Measure one number: minutes from the first alarm to the right team. Ask yourself: do I know that number for our own shift today?
2. Draw the line between knowing and deciding. In 2000, Parasuraman, Sheridan and Wickens split automation into four stages: acquiring information, analysing it, selecting a decision and carrying it out. Their advice fits a NOC closely. Acquisition and analysis may be automated to a high level if the operator keeps access to the raw data, "highlighting, but not filtering", and knows how unreliable the automation is. Where the stakes are high, the machine should go no further than suggesting, not executing, a preferred option. So:
Two findings shape the design. In a 2021 study, Gagan Bansal and colleagues found that explanations made people more likely to accept an AI's recommendation whether it was right or wrong. So every sentence the Co-Pilot writes links to its source alarm and timestamp; fluency earns no trust. And Lisanne Bainbridge warned in 1983 that automated systems watched by former manual operators are riding on skills that later generations cannot be expected to have. Keep regular drills without the Co-Pilot. Ask yourself: if the Co-Pilot names the wrong domain at 02:22, whose name is on the decision?
Sources: Y. Chen et al., "Automatic Root Cause Analysis via Large Language Models for Cloud Incidents" (RCACopilot), EuroSys 2024; A. Wong et al., "External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients", JAMA Internal Medicine, 2021; OWASP Top 10 for LLM Applications 2025, LLM06: Excessive Agency; R. Parasuraman, T. B. Sheridan and C. D. Wickens, "A Model for Types and Levels of Human Interaction with Automation", IEEE Transactions on Systems, Man, and Cybernetics, Part A, 2000; G. Bansal et al., "Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance", CHI 2021; L. Bainbridge, "Ironies of Automation", Automatica, 1983.
From the mentor's corpus
1. Yes, and the mentor's decision doctrine explains why the manual path is not the safe one. Its first rule treats human failure as a design axiom: a system that depends on one permanently trustworthy guardian is badly designed. At 02:17 the operator is that guardian, the only place where RAN, power, transport and IP meet. The second rule forbids the opposite error: AI is not a trusted root either, and its power must be limited by scope, time, evidence, independent review and revocation. So do not choose from the menu as offered. The doctrine's rule of frame before choice says an A/B/C menu is tested before it is accepted, and if all its options are bad, the sequence changes. Here the three paths are one sequence, and the sixth rule gives the order. Roughly 80% confidence can be enough for a reversible decision, but critical infrastructure still advances through isolated tests, AI adversaries, bounded live users and progressively wider deployment, for as long as the evidence requires:
Ask yourself: which of these steps could I start next month without asking anyone for production access?
2. The mentor's first volume draws this line in its pattern for analysis: AI identifies patterns and correlations; the human distinguishes correlation from causation, determines the implications for action and makes the decision. A shared transport path is a correlation. That it caused five BTS outages is a judgment, and so is which team to wake. His HAI5 framework, built on the same Parasuraman taxonomy with governance added against automation bias, turns the line into a declaration made task by task:
Two ideas shape the screen. The mentor's decision doctrine requires an epistemic status for every material claim, and three of its categories fit a NOC exactly: case fact, inference and unknown. Every line of the picture says which it is, with its source and time. His chapter on Heuer describes the terrain of 02:22: ambiguous signals, pressure toward premature closure, and the temptation to reward fluent confidence over honest uncertainty. The case's own fear, weak evidence explained fluently, is that temptation. So rival explanations stay on the screen: evidence for one hypothesis may fit the others too. Ask yourself: at 02:22, could my operator tell which parts of the picture are facts and which are the machine's guesses?
On the NEO Turn. The case's turn, from the operator as integration layer to the operator as orchestrator, is the mentor's own reading of OODA in the NEO era. His Task-to-Agent protocol puts it in one line: the machine observes and orients at scale, while the human keeps authority over deciding and acting. His dissertation, Vanguard Intelligence, adds the governance: the AI orients without constant supervision but within a clear commander's intent and kill indicators, and whatever it flags or recommends carries a justification a human can evaluate. One priority stays above speed. Orientation, the construction of mental models, is the decisive phase, and every step of the protocol is meant to strengthen it rather than merely accelerate execution. So the Co-Pilot's last test is whether an operator, after six months with it, explains an incident better without it.
Sources: the mentor's decision doctrine of 22 August 2026 (§1–3, §5, §6); Vanguard Leadership, vol. 1, Domain 2: Analysis Augmentation; the HAI5 framework, as set out in Vanguard Leadership, vol. 2 and the Vanguard Task-to-Agent Mapping Protocol; Vanguard Leadership, vol. 2, the chapter on Heuer's ancestry; the Vanguard Task-to-Agent Mapping Protocol, theoretical foundations (the OODA architecture); the mentor's dissertation, Vanguard Intelligence, chapter 3 (the role of AI in parallel OODA).
The case asks how far to trust an AI inside a live operation, where a confident mistake costs the most. The mentor faced the same question in a harder place: Vanguard AI, his project for AI that coordinates care for vulnerable patients across hospitals in Croatia and Slovenia. Its answer is a path, and a NOC can walk it in months rather than years.
The path. The project's critical path runs from retrospective validation to shadow mode to an active pilot. In shadow mode the system runs inside real clinical workflows without driving any clinical action, and gathers evidence on workflow fit and on its errors under live conditions. Its risk register adds two rules: retrospective validation before prospective, and a GO/NO-GO decision at a fixed milestone. For the NOC:
The champion. The same register names a second risk: low clinician adoption and resistance to AI. It answers it with a clinical champion and co-design with clinicians from the first month. The NOC's version is its most experienced night-shift operator, the person who is the integration layer today. Make them the Co-Pilot's co-designer from week one and a signatory of the GO/NO-GO. That answers the case's deepest worry. The Co-Pilot is built to carry that operator's mental model to others, not to replace it. And the person who holds the model decides whether the machine has earned the night shift.
The record. The project insists on verifiable provenance for the data it acts on. In a NOC, that means three things:
So the answer to the case's closing question, how much of incident understanding to hand to AI, is this: as much as the evidence has earned, in this order, and never the decision itself. Ask yourself: who is my champion, and what number would make them say no-go?
Sources: the mentor's Vanguard AI project, Horizon Europe proposal (the critical path from retrospective validation through shadow mode to an active pilot; the risk register: retrospective before prospective validation, a GO/NO-GO milestone, a clinical champion and co-design with clinicians; verifiable data provenance).
The pilot: yes, of course, as soon as possible. The NEO stack is ideal for this case, and we will gladly run a free pilot.
A question for the table, a disagreement, what you would have done. The case lead reads every comment; the ones the table takes up enter the chapter as questions from the room, with your name.