Merit AC™
Guide · systems engineering

Incident response for AI agents: what a postmortem actually needs to capture

A database outage postmortem asks what broke and why. An agent incident postmortem has to ask that plus a second set of questions a deterministic system never raises: what did the agent decide, based on what it believed was true, and would a different, equally plausible decision have looked identical right up until it didn't.

1. Why the standard template falls short

A conventional incident postmortem is built around a timeline of system states: service X returned errors starting at time Y, caused by deploy Z, resolved by rollback. That structure assumes the system's behavior at each point is fully determined by its code and its inputs, which makes "why did it do that" answerable by reading the code. An agent's behavior at each point is also shaped by a model's output given a context window, which means the same code, the same tools, and very similar inputs can produce different decisions on different runs -- so "why did it do that" requires reconstructing what the agent actually believed and considered at the moment of the decision, not just what code path executed.

This isn't a reason to abandon the standard template -- the timeline, the impact assessment, the root-cause-to-remediation structure all still apply. It's a reason to add a second track alongside it, specific to the reasoning layer, that a deterministic-systems postmortem format has no slot for.

2. The five questions worth adding

What did the agent believe was true, and was it? Reconstruct the actual context the agent was working from at the decision point -- not what was true, but what the agent's context window contained that it treated as true. A large share of agent incidents trace back to the agent correctly reasoning from an incorrect premise (stale data, a misread tool result, an injected instruction it trusted) rather than to a reasoning failure at all. Conflating "the agent reasoned badly" with "the agent reasoned correctly from bad information" leads to the wrong fix -- the first calls for a prompt or model change, the second calls for fixing what fed the context.

What options did it have, and why this one? Where the trace captured a decision record (see agent observability), review what alternatives existed and why the agent's reasoning favored the one it took. Where no decision record exists, this is often unreconstructable after the fact, which is itself a finding worth writing down -- "we cannot determine why the agent chose this path" is a legitimate, if uncomfortable, postmortem conclusion, and it's a direct argument for better decision logging rather than a gap to paper over with a guess.

Would a re-run produce the same result? Because model outputs carry some irreducible variability, a single bad outcome doesn't by itself tell you whether the system has a systematic problem or hit a low-probability tail case. Where it's safe to do so, re-running the same input (ideally several times) against the same harness configuration gives a real signal: a result that reproduces consistently points to a systematic cause worth fixing structurally; a result that doesn't reproduce points toward the harness needing a guardrail that catches this specific tail case regardless of cause, since you can't patch out a behavior that only shows up probabilistically.

What would have caught this before it shipped? Specifically: was there an eval case that should have covered this scenario and didn't, or did the scenario genuinely fall outside anything reasonable to have anticipated? The answer determines whether the fix is "add this case to the eval suite" (see eval-driven development) or something harder to plan for, and conflating the two leads either to an eval suite that never grows from real incidents, or to a team quietly accepting as "unforeseeable" something a slightly more thorough eval would have caught.

Did a control that should have stopped this exist, and if so, why didn't it? If an approval gate, an authority boundary, or an observability alert was supposed to catch this class of problem and didn't, that's a harness defect distinct from whatever the model did -- and arguably the more important finding, since a harness gap that let one bad decision through will let the next one through too, regardless of whether this specific model behavior ever recurs.

3. Capture the trace before it rotates out

Implementation note: the evidence an agent incident postmortem depends on -- the full context window at the decision point, the exact tool call arguments and results, the model version and sampling parameters -- is often subject to much shorter retention than a conventional system's logs, especially the raw model-call payloads discussed in agent observability's retention-policy section. The first action on any agent incident worth investigating, before anything else, is pulling and preserving the full trace for the affected run, because the window to do that can close in hours, not weeks, once standard retention policies rotate the payload layer out.

4. A failure pattern worth naming: the plausible non-explanation

A specific trap in agent postmortems is accepting a plausible-sounding narrative in place of an evidenced one. "The model probably got confused by the long context" is a sentence that sounds like an explanation and explains nothing verifiable -- it doesn't point at a specific piece of context, doesn't suggest a specific fix, and can't be checked against the actual trace. A postmortem that settles for this kind of narrative closes the incident without actually reducing the odds of a repeat. The discipline worth enforcing: every causal claim in the postmortem should point at a specific line in the trace -- this tool result, this context window content, this decision record -- or be labeled explicitly as speculation the evidence couldn't settle, not stated as if it were a finding.

5. A minimal checklist

  • The full trace for the affected run is pulled and preserved before retention policy can rotate it out.
  • What the agent believed is reconstructed from its actual context, separate from what was actually true.
  • Reproducibility is checked where it's safe to re-run, to distinguish a systematic cause from a tail case.
  • Every causal claim points at specific trace evidence, or is explicitly labeled as unresolved speculation.
  • The eval-coverage gap is named explicitly — could a test have caught this, and if so, is it being added now.
  • Any control that should have stopped this is checked, and its failure is treated as its own finding, separate from the model's behavior.

6. Where this fits

This guide assumes the observability groundwork in agent observability is already in place -- a postmortem can only reconstruct what was actually captured. Findings from this process are also the most reliable source of new eval-driven development test cases: a real incident is a test case nobody has to invent.