Merit AC™
Content From the Merit AC team

Guides

General writing on doing AI work well — governed agentic DevSecOps adapted from our own Enterprise Agentic DevSecOps Handbook, plus standalone field guides on how AI systems actually work, break, and get evaluated. Looking for cloud architecture or building with Claude? Those moved to their own sections.

Governed agentic DevSecOps

Adapted from our own Enterprise Agentic DevSecOps Handbook

AI systems & engineering

Independent field guides — how AI systems actually work, break, and get evaluated

AI system design patternsTwelve archetypes, six complex agent patterns, and the ML/AI software landscapeIncident response for AI agents: what a postmortem actually needs to captureThe questions a standard outage postmortem doesn't ask, and why the evidence window for an agent incident closes faster than for most outagesAgent memory architecture: short-term, long-term, and retrieval patternsWhat actually needs to persist across sessions, and why write policy is the hard partAgent observability: what to log, trace, and alert onThe difference between a transcript and a trace, and the specific signals worth alerting on before an agent's behavior becomes an incidentAI evaluation methods: rubrics, LLM-as-judge, and benchmarksRubrics, LLM-as-judge, and benchmarks — when to use which, and how judges failAI harness engineering: the discipline nobody namedWhy the system around the model, not the model itself, decides whether an agent is safe — and why that system deserves to be engineered on purposeContext engineering: what actually goes into the context windowWhat actually competes for space in the context window, and how to manage itCost and latency tradeoffs in LLM system designModel tiering, caching, streaming, and batching — the four levers, and their real costsEval-driven development: building an evaluation pipeline for AI featuresA regression suite for prompts and agent behavior, wired into the same pipeline as codeFine-tuning, prompting, and RAG: choosing the right leverBehavior, knowledge, and consistency at scale — three different questions, not oneHuman-in-the-loop design: where to put the approval gateWhy most approval gates stop working within weeks, and the design that keeps a reviewer actually reading what they're approvingMulti-agent orchestration patterns: when one agent isn't enoughThe four coordination shapes worth their cost, and the test for whether a second agent is solving a real problem or hiding onePrompt injection and agent security: a practical threat modelDirect vs. indirect injection, and why enforcement has to live outside the modelRAG failure modes: a debugging field guideA debugging field guide — retrieval failure, lost-in-the-middle, reranking, chunkingStructured output reliability: getting agents to actually follow a schemaWhy schema-following fails in specific, recurring ways, and the layered defense that catches what the model alone won'tWhy per-seat AI spend resists ROI measurementFour reasons the ROI question has no clean answer, and three things that are measurable instead

More guides land here over time — no invented statistics or a testimonial standing in for a real one, ever.