“Incident response was built to end with a name. An agent incident ends with a question you can’t answer: who was accountable?”
Every incident response lifecycle converges toward the same destination: a cause and a responsible party. Detect, contain, eradicate, recover, and—at the end—attribute. Name the adversary, name the flaw, name the owner of the fix, and close the ticket. It works because there has always been someone or something on the other end to name. That destination survives malware. It does not survive an autonomous agent acting on authority you delegated, through a channel it was built to trust, at a speed no human reviewed. In that world, incident response reaches the point where it should name an accountable party—and finds a delegation instead of a person. You will close the incident. You will not be able to close the question the board is actually asking.
Before You Read Further — Know Which Question Goes Unanswered
Incident response has always closed on one question. AI changed which one you can’t answer. Your runbook was written for the first column. This debrief is about the third.
Attribution
Who attacked us?
Classic IR closes here. Identify the adversary, cut their access, write the report. The incident ends when the attacker has a name.
Root Cause
What failed?
Modern IR traces the chain — the misconfig, the exploited flaw, the control that didn’t fire. The incident ends when the technical cause is understood.
Accountability
Who was accountable?
The agent acted on authority you delegated. No adversary, maybe no failure — and no human who owns the decision. There is no log of accountability, and no one to hand the question to.
The Operation
It is a Thursday afternoon. A national bank runs an autonomous fraud-response agent with standing authority to freeze accounts and reverse transactions the moment it detects a risk pattern. It is valuable precisely because it acts without waiting for a human — when the signal crosses a threshold, the agent moves. That speed is the product. That standing authority is also the exposure.
Just after two o’clock, the agent begins freezing accounts at scale — tens of thousands in under an hour — after a risk score spikes across a segment of the customer base. An upstream data feed the agent trusts had shifted that morning. Whether the feed was poisoned by someone or simply drifted, no one can yet say. It does not slow the agent down. The agent did exactly what it was authorized to do: it saw the pattern, it crossed the threshold, it acted. Customers are locked out mid-transaction. The regulator’s phone is ringing before the SOC has opened a ticket.
Incident response runs by the book. The agent is paused, the freezes reversed, accounts restored by evening. Textbook execution: detected, contained, recovered. And here is what unsettles the responders — there is nothing to eradicate. No malware detonated. No boundary was crossed. Every action the agent took was authorized, credentialed, and logged. The runbook closes clean, on schedule, with a green dashboard.
Then the reckoning. The board and the regulator ask the only question that matters to them: who authorized this, and who is accountable? And the bank discovers it cannot answer. The threshold that triggered the freeze was set by a team that no longer owns the agent. The upstream feed belonged to another group entirely. No human was in the loop at the moment of the decision. The logs record what the agent did to the fraction of a second — and nothing about who was accountable for letting it. The incident is closed. The accountability has no owner to hand it to. And nobody can say for certain whether it was ever an attack at all.
Three Perspectives
The Trusted Leader
“We ran the playbook perfectly. Detected, contained, recovered by dinner. And when the board asked who was accountable, I had nothing.”
I have run incident response for twenty years. Every one of them ended the same way — we found who or what caused it, and someone owned the fix. This one broke the pattern. There was nothing to eradicate; nothing broke in. The agent did what we authorized it to do. But “we authorized it” is not an answer a board accepts, and it is not an answer a regulator accepts. I could tell them exactly what happened, to the millisecond. I could not tell them who was accountable — because we never decided that before we handed the agent the authority. We provisioned the power and skipped the ownership.
The Defender
“I could reconstruct every action the agent took. I couldn’t reconstruct who was supposed to own the decision to let it.”
My tooling is built to answer what happened, and it did — I had a flawless timeline within the hour. What I did not have was the other half: who granted this agent that authority, at what threshold, with what scope, and who was in the loop when it acted. Those aren’t in my logs, because nobody logs authority the way we log activity. And here is what genuinely unsettled me — I could not even tell you whether this was an attack. The upstream feed might have been poisoned; it might have simply drifted. My detection stack cannot separate the two, and the accountability gap is identical either way.
The AI-Native Diamond Model is built for exactly this. Traditional intrusion analysis assumes an adversary sitting at one vertex of the diamond. Point the model at an agent incident and that vertex may be empty — the live question is no longer “who attacked” but “whose authority did the agent act on, and who owned it.” The model doesn’t just describe the attack. It locates the accountability — which is why every debrief in this series runs on it, and why this one is where it’s built.
The Attacker
“I don’t have to get caught. I don’t even have to exist. Your agent acts, and there’s no one to blame but the authority you gave it.”
Here is the elegant part — I don’t need to be found, because you can’t even prove I was here. I nudge one upstream signal your agent already trusts, and your own standing authority does the rest. Maybe it was me. Maybe your data just drifted. You will spend the entire post-mortem unable to prove which, and it will not matter, because either way there is no human holding the decision. You built an agent that can act at scale and never decided who answers for it when it does. I didn’t breach you. I just chose the moment your own authority would fire — and let you argue with yourselves about whether I existed.
Technical Assessment
The Threat Architecture
AI incident response inverts the endpoint of the discipline. Traditional IR is a convergent process — it narrows toward a cause and a responsible party: an adversary, a vulnerability, a misconfiguration, an owner. The entire lifecycle is engineered to terminate in attribution and accountability. An agent incident breaks the convergence. The agent acts on delegated authority through a trusted channel; the action is authorized and logged; and when the process reaches the point where it should name an accountable party, it finds a delegation where a person should be.
The defining property is that authority and accountability were separated at design time. You can delegate authority to an agent in an afternoon — a threshold, a scope, a set of permitted actions. Accountability does not transfer with it. It stays with whoever should have owned the decision, except no one was ever assigned, because the agent was going to make the decision. So the incident produces a perfect record of what happened and no record of who was accountable — and, uniquely, it often cannot even establish whether an adversary was present at all. Attack and accident collapse into the same evidentiary void, and the accountability gap is identical in both. You cannot attribute your way out of an incident with no attacker, and you cannot restore accountability from a backup.
The Diamond Model, Rebuilt for Agent Incidents
The threads between the vertices carry the analysis, not the vertices themselves. Trace instruction → channel → workflow → authority → impact and the incident reconstructs as a single chain — one that’s shareable as threat intelligence even when the adversary vertex stays empty. In an agent incident the model cannot always tell you attack from accident. What it can always tell you is where the authority ran, and therefore where the accountability should have lived. That is what makes it the backbone the rest of the series is built on.
Accountability Gap Analysis
| Capability | Establishes What Happened | Establishes Who Was Accountable |
|---|---|---|
| Detect / contain / eradicate / recover runbook | Yes | No — a clean recovery names no owner |
| SIEM / activity & audit logging | Yes | No — logs what the agent did, never who authorized it |
| Adversary attribution / threat intel | Partial | No — there may be no adversary to name |
| Authority & delegation registry (who granted what to which agent) | No | Yes — purpose-built; the missing layer in most orgs |
| Decision-owner attestation / human-in-the-loop record | No | Yes — the only artifact that assigns a human; almost nobody keeps it |
The Accountability Multiplier
The gap compounds with autonomy and scale. A single human decision has a single owner. Delegate that decision to an agent and the owner does not multiply — it vanishes, while the number of decisions explodes. Every consequential action the agent takes on delegated authority is an action with no accountable human at the moment it is made. Scale the agent across a workflow, then across a fleet, and you scale unowned decisions, not owned ones. Traditional IR ends by assigning accountability to a party. Agent IR ends by discovering there was never a party to assign it to — and the more you automate, the wider that void grows. Authority scales in an afternoon. Accountability was never wired to follow it.
CISO Debrief
“Nobody attacked their way out of accountability. The agent acted on authority they delegated and never owned — and a clean incident report named no one.”
AI incident response is not a faster version of the old discipline. It is the old discipline arriving at a void it was never built to encounter. The authority was granted. The action was logged. The runbook closed on schedule. What was missing was the one layer the board and the regulator actually needed: a record of who was accountable for the agent’s authority — who granted it, at what bound, and who owned the decision when the agent acted on it. You can restore a system from backup. You cannot assign accountability, after the fact, to a decision no human ever made. That has to be wired to the authority before the agent ever uses it — because after the incident, there is nothing left to wire it to.
IR Directives
Wire accountability to authority before you delegate it. Every grant of standing authority to an agent needs a named, accountable human — recorded at grant time, not reconstructed after an incident. Unowned authority is the finding waiting to happen.
Log authority, not just activity. Your SIEM records what the agent did. Add the layer it doesn’t have: who granted the agent this authority, at what threshold and scope, and who was in the loop at the decision point. Activity logs alone name no one.
Redefine “incident closed.” A restored system is not a closed incident. Define closure as an assigned accountable owner and an answered “who authorized this” — not a green dashboard and a reversed action.
Build IR that doesn’t depend on proving an attacker. Assume attack and accident are indistinguishable. Whether the trigger was poisoned or drifted, the accountability question is identical — your process must reach the same owner either way, without waiting to name an adversary.
Keep a live delegation registry. Maintain a current map of which agents hold which authority, granted by whom, owned by whom, revocable how. If an incident is the thing that forces you to draw it, you have already lost the post-mortem.
Put a human owner at every consequential decision boundary. Not necessarily in the loop on every action — but accountable for the bound. Someone must own the threshold the agent acts within, and answer for it when the agent acts.
Close the Governance Gap
Authority Provenance — First-Class Log Source. Treat delegated authority the way you treat authentication. Tie every consequential agent action to the authority that permitted it and the human who owns that authority — not just the API call that executed it.
Accountability Architecture — Owner at Grant Time. Assign the accountable human when authority is delegated, not when the incident review demands one. Accountability that has to be reconstructed after the fact usually cannot be.
Delegation Governance — Scoped, Owned, Revocable. Govern every grant of agent authority as a managed relationship with a named owner, an explicit bound, and a revocation path. Authority with no owner is an incident with no answer.
Five Questions for Your Next Executive Meeting
1. If an agent took a consequential action today, could we name the human accountable for the authority it used — without reconstructing it after the fact? If we can’t, that is the finding.
2. What is our definition of “incident closed” — restored systems, or assigned accountability? Do we have a procedure for the second?
3. How many standing grants of authority do our agents hold, who owns each, and which are revocable?
4. Could we tell a regulator who was accountable for a decision our agent made autonomously — and would “the agent did it” survive the question?
5. Does our IR plan depend on proving an attacker — and what does it do when we cannot say whether there was one?
Technical Reference
Threat Category: Incident Response & Accountability in Autonomous Agent Systems
Techniques: Delegated-Authority Abuse · Instruction Injection · Trusted-Channel Manipulation · Attribution Denial · Accountability Gap · Human-Out-of-the-Loop Action
MITRE ATT&CK: T1078 — Valid Accounts · T1199 — Trusted Relationship · T1565 — Data Manipulation · T1656 — Impersonation · T1485 — Data Destruction
OWASP LLM Top 10 (2025): LLM06 — Excessive Agency · LLM01 — Prompt Injection · LLM04 — Data & Model Poisoning
Detection & Governance Controls: Authority Provenance · Delegation Registry · Decision-Owner Attestation · Human-Accountability Mapping · Attack/Accident-Agnostic IR
Framework: AI-Native Diamond Model — the establishing reference for IR question reframing across the series
“When AI Attacks” is a practitioner-grade security intelligence series written for CISOs, security leaders, and defenders navigating the AI threat landscape.
The scenarios described in this series are grounded in documented, publicly reported threat intelligence patterns and forward-looking analysis of multi-agent system risk. They describe an emerging threat pattern for defensive planning; they do not depict a specific named incident and do not reflect confidential information from any employer.