Merit AC
2026-09-18

Anthropic gives Accenture staff-level access to evaluate its own models

Embedded evaluators get more than an outside auditor's usual view -- direct interaction with staff and visibility into deployment decisions -- in a partnership both companies say they'll fund with at least $1 billion each over five years.

Anthropic announced a partnership on September 18 with Accenture's AI unit, Faculty, to embed evaluators directly inside Anthropic's own operations -- running model evaluation, red-teaming, alignment assessments, and safeguard testing with, in Anthropic's own description, "access comparable to an employee's." That means Accenture's evaluators can observe model development as it happens, follow along on deployment decisions, and interact directly with Anthropic staff, rather than reviewing finished artifacts from the outside after the fact.

A bet neither company is making alone

Both companies say they expect to invest at least $1 billion each in building out this capacity over the next five years. Anthropic is funding Accenture's work directly for now, but the announcement is explicit that it doesn't want that to be the permanent model: "long-term, we think funding should come from pooled or government sources" -- an acknowledgment that a lab paying its own evaluator, even an embedded and well-resourced one, isn't the end state anyone actually wants. Anthropic also frames Accenture as the first of several: the company says it's in discussions with METR and other nonprofit evaluators to do similar embedded work, positioning this as the opening move in a broader evaluator ecosystem rather than an exclusive arrangement.

The access is the point, and also the catch

"Embedded" evaluation is a meaningfully different claim than the external audits and published safety cards labs have leaned on so far -- an evaluator inside the building, talking to the team that shipped the model, can catch things a document review can't. But it's still Anthropic choosing, funding, and granting access to its own evaluator, which is exactly the dependency the "pooled or government sources" line is trying to get ahead of. For a company deciding how much to trust a vendor's safety claims, the honest read is that this is a genuine step up in rigor over self-attestation alone, paired with a funding structure Anthropic itself says isn't the answer yet.

Sources

← All news