Structured output reliability: getting agents to actually follow a schema
A schema violation that shows up once in two hundred calls looks like noise until it hits the one downstream system with no tolerance for a malformed field, at which point it's not noise, it's an outage with a two-hundred-to-one odds of happening on exactly the day you can't afford it.
1. "Mostly follows the schema" is not a property that composes
A model that follows a JSON schema correctly 99% of the time sounds reliable in isolation. Chain three such calls together, each depending on the previous one's structured output, and the chain's overall reliability is closer to 97%, not 99% -- and a system making thousands of calls a day will hit that remaining fraction often enough that "mostly works" has to be treated as a known, recurring condition to design around, not an edge case to shrug off. The practical implication: structured-output reliability isn't a property you achieve once with a good prompt. It's a property of the whole pipeline -- the request, the parsing, the validation, and what happens on failure -- and the request is only the first of those four places to get right.
2. Where it actually breaks, most common first
Near-miss formatting. The model returns valid-looking JSON with a type mismatch the schema didn't anticipate -- a number returned as the string "42" instead of 42, a null where the schema expects an empty string, a boolean rendered as "true" with quotes. These are the most common failures by far, and the most recoverable: a coercion layer that normalizes obvious near-misses before strict validation catches most of them without ever bothering the model with a retry.
Schema drift under complexity. As a schema grows more deeply nested, or includes more optional and conditional fields, model adherence degrades -- not linearly, but noticeably past a complexity threshold that varies by model and gets worse the less common the exact field combination is in the model's training distribution. A fifteen-field flat object is usually fine; a schema with three levels of nested optional arrays, each depending on a sibling field's value, is where adherence starts to slip even on frontier models.
Prose leakage. The model wraps its JSON in explanatory text -- "Here's the JSON you requested:" followed by the actual object, or a trailing note after it -- despite an instruction not to. This is a formatting failure, not a reasoning failure, and it's almost entirely preventable with the right generation mode rather than the right prompt wording (see below).
Semantic non-compliance. The hardest class, and the one a schema validator can't catch at all: the JSON is syntactically perfect and the schema is satisfied, but a field's content doesn't actually mean what it's supposed to -- a confidence field that's always 0.95 regardless of actual uncertainty, a category field correctly typed as a string but set to a plausible-sounding value that isn't one of the implied options the field was supposed to enumerate. Catching this requires validating semantics, not just shape, which is a different, harder layer than JSON schema validation and belongs in evals rather than runtime parsing.
3. The layered defense, in the order to build it
Use constrained decoding where it's available, before relying on prompting alone. Modern model APIs increasingly support generation modes that mechanically constrain the output to valid JSON, or to a specific schema, at the sampling level rather than hoping the model chooses to comply. Where this is available, it eliminates prose leakage and most syntax errors entirely -- it's a stronger guarantee than any amount of prompt engineering, because it changes what tokens the model is allowed to produce rather than asking it nicely. Where it isn't available for a given model or provider, the remaining layers below carry more weight.
Validate strictly, but coerce before you reject. A validation layer that rejects on the first type mismatch throws away calls that were one coercion away from fine. Normalize known near-misses (stringified numbers, stray whitespace, common enum case mismatches) before running strict validation, and only treat what survives coercion as a real failure.
Worked example: a schema expects
{"status": "approved" | "rejected" | "pending"}and the model returns{"status": "Approved"}. A coercion pass that case-normalizes enum values before validation turns this into a non-event. A validator with no coercion layer rejects it, triggers a retry, burns the latency and cost of a second model call, and has roughly the same odds of hitting a similarly trivial mismatch again.
Retry with the actual error, not a generic resend. When a call does fail validation, the most effective retry includes the specific validation error in the next prompt -- "the date field must be ISO-8601, you returned Jan 5" -- rather than simply resending the same request and hoping for a different sample. A targeted retry that tells the model exactly what was wrong converges far faster than a blind retry, and caps the number of blind retries at one or two before falling back to a deterministic default or escalating, since a model that fails the same schema twice in a row with no new information is unlikely to self-correct on a third identical attempt.
Have a deterministic fallback for the fields that can't be allowed to fail. For any field a downstream system treats as load-bearing -- a value that triggers a payment, a routing decision, an irreversible action -- decide in advance what happens if structured extraction fails entirely after retries: a safe default, a routed-to-human fallback, or an explicit "could not extract" state the downstream system is built to handle, rather than a null that a less careful piece of code downstream might silently treat as zero, false, or "proceed."
4. Semantic checks belong in evals, not in the request path
Because semantic non-compliance can't be caught by schema validation, it has to be caught by testing the field's actual meaning against known cases -- which is exactly what an eval suite is for, not something to try to prevent entirely at generation time. A periodic eval that checks whether the confidence field's distribution looks like actual confidence (rather than a constant) or whether category values are drawn from the expected set across a sample of real traffic will catch this long before a user complaint does. See eval-driven development for how to build that loop; the structured-output case is a specific, high-value instance of the general discipline that guide covers.
5. A minimal checklist
- Constrained decoding used wherever the provider supports it, not prompting alone as the first line of defense.
- A coercion layer runs before strict validation, normalizing known near-misses instead of rejecting on first type mismatch.
- Retries include the specific validation error, not a blind resend, and are capped before falling back.
- Every load-bearing field has a defined fallback for total extraction failure — a safe default, a human escalation, or an explicit failure state, never a silent null.
- Semantic compliance is checked in evals, on a schedule, against real traffic samples — not assumed just because the schema validator passed.
6. Where this fits
This is a specific instance of the eval loop eval-driven development builds generally, and the retry/fallback behavior described here is exactly the kind of per-tool failure-handling contract the checklist in AI harness engineering asks every stateful tool to have. If structured output feeds a decision an approval gate reviews, the fallback-on-failure behavior matters as much to that reviewer as the happy path does.