Merit AC™
Guide · systems engineering

Agent observability: what to log, trace, and alert on

Most teams discover their agent has an observability gap at the worst possible time: during an incident, when someone asks "what did it actually do" and the honest answer is a few megabytes of chat transcript with no structure, no correlation id, and no way to tell which of forty tool calls was the one that mattered.

1. A transcript is not a trace

A chat log answers "what did the agent say." A trace answers "what did the agent do, in what order, with what inputs, producing what outputs, and how do I find the one step that explains this result." Those are different artifacts, and a system that only has the first is flying blind in exactly the way that matters once something goes wrong. The transcript tells you the agent said it was going to check inventory before confirming an order; the trace tells you whether the inventory-check tool call actually happened, what it returned, and whether the agent's next message accurately reflected that return value or quietly ignored it.

The practical difference is structure. A trace is built from discrete, typed events -- a tool call with its arguments and return value, a model call with its token counts and latency, a decision point with the options the agent considered -- each carrying a correlation id that ties it to the request that started the chain. A transcript is prose. Prose is what a human reads at the end; structure is what makes the investigation possible before a human has to read anything.

2. The four things worth recording on every run

A correlation id, threaded everywhere. One identifier, generated at the start of a request, attached to every model call, every tool call, every sub-agent invocation, and every log line that chain produces. Without this, reconstructing "everything that happened for this one user action" means manually cross-referencing timestamps across however many systems the agent touched -- the model provider's logs, the tool's own logs, the orchestration layer's logs -- and timestamps alone are not a reliable join key once anything runs concurrently.

Every tool call, with its real arguments and real return value. Not a summary of what the agent intended to do -- the actual request sent and the actual response received, including errors. This is the single highest-value piece of structured data in the whole system, because it's the layer where an agent's internal reasoning meets the external world, and it's the layer every other debugging question eventually reduces to: did the tool get called with the right arguments, and did it return what the agent's next step assumed it returned.

Model call metadata, separated from model call content. Token counts, latency, model version, and whether the call hit a cache are metrics, not content -- they belong in a system built for aggregation and alerting, not buried inside a transcript where answering "did our p99 latency regress after last week's prompt change" requires parsing prose to find out.

A decision record at genuine branch points. When an agent chooses between materially different paths -- which tool to use, whether to ask for human approval, whether to retry or give up -- recording which path it took and a short machine-readable reason is worth far more after the fact than the chain-of-thought text usually associated with that decision, which is verbose, inconsistent in format, and expensive to search across many runs.

3. What actually deserves an alert

Logging everything and alerting on nothing produces a system nobody looks at until it's already broken; alerting on too much produces a system everyone mutes. The signals worth a real-time alert, as opposed to a dashboard someone checks when curious, share one property: they indicate the harness itself is behaving outside its designed bounds, not that the business result was merely disappointing.

Worked example. A coding agent's tool-call error rate holding at 2% is probably fine -- real-world inputs produce real edge cases, and a tool that fails cleanly and gets retried is the system working as designed. The same agent's error rate jumping to 40% in the last ten minutes is not a quality signal to review next sprint; it's very likely a dependency outage, a credential expiring, or a schema change upstream, and it should page someone now, specifically because the harness has a documented failure-handling contract this violates (see the checklist in AI harness engineering) and the alert is how you find out that contract's assumptions just broke.

Concretely, alert on: a sudden shift in tool-call error rate or latency relative to its own recent baseline (not a fixed threshold picked once and forgotten); any action that bypassed an approval gate that should have caught it; token spend per request moving sharply off its baseline (often the first visible sign of a prompt regression, a retry loop, or a context-window leak); and any tool call whose arguments fall outside the authority that tool is supposed to have, which should be rare enough that every occurrence is worth a human look.

4. The cardinality trap

A common failure in agent observability isn't under-logging, it's logging in a shape that can't be queried at the volume agents actually produce. An agent that calls tools in a loop can generate orders of magnitude more events per user action than a traditional request-response service does, and a trace schema designed by copying a web-service's logging conventions tends to choke on that volume -- too many unique field values (a raw user prompt as a tag, for instance) blow up the cost and usability of most tracing backends. The fix is deciding, before volume becomes a problem, what's a high-cardinality field that belongs in payload storage (full prompts, full tool arguments) versus a low-cardinality field that belongs in the indexed, queryable trace (tool name, status, model version, correlation id) -- the second set is what you filter and alert on; the first set is what you open once you already know which trace to look at.

5. Privacy and retention aren't an afterthought here

Agent traces routinely contain more sensitive material than traditional application logs, because the whole point of an agent is that it reads things -- documents, emails, customer data -- as part of doing its job, and a faithful trace captures what it read. A tracing system built without a retention policy and a redaction pass for sensitive fields will, with enough runs, become a second copy of whatever sensitive data the agent ever touched, outside whatever access controls protect the original. This isn't a reason to log less; it's a reason to decide, up front, which fields get redacted or truncated before storage, and how long the full payload layer is retained versus the indexed trace layer, which usually needs to persist far longer.

6. A minimal observability checklist

  • A correlation id on every event, threaded through every model call, tool call, and sub-agent invocation in a request.
  • Full, structured tool-call records — real arguments and real return values, not a paraphrase.
  • Model metadata separated from content, in a system built for aggregation, not buried in prose.
  • Decision records at real branch points, machine-readable, not just chain-of-thought text.
  • Baseline-relative alerting on error rate, latency, and spend — not fixed thresholds nobody revisits.
  • A cardinality plan: what's indexed and queryable versus what's in payload storage, decided before volume forces the question.
  • A redaction and retention policy for sensitive fields, decided before the first trace with real user data is written.

7. Where this fits

Observability is one of the control disciplines AI harness engineering names as part of the surrounding system a model needs, not an add-on. It's also the thing that makes eval-driven development possible at scale -- an eval suite tells you a change made things better or worse in aggregate; a trace is what tells you why, for the one run a human actually wants to understand. For systems built from more than one agent, see multi-agent orchestration patterns for why tracing matters even more once a failure can originate in any one of several agents or in the handoff between them.