skip to content
The Weighted Average

AI Safety & Security

Anthropic's Evaluation Bill Is Not a Safety Certificate

Anthropic and Accenture plan at least $2B over five years for AI safety. The $400M annual average buys capacity, not a product certification.

Four people meeting around a table in a glass-walled office
Four people meeting around a table in a glass-walled office. Photograph by Rodeo Project Management Software

Enterprise buyers should ask for evaluation deliverables, not relax model approvals, after Anthropic announced an embedded-evaluation partnership with Accenture. Combining the partners’ expected minimum investments implies $400 million a year on an arithmetic-average basis, but neither the size of that plan nor the announcement certifies a customer’s deployed system.

A large budget meets an unfinished profession

Anthropic says it and Accenture each expect to invest at least $1 billion over five years in building capacity in this area. Accenture’s September 18 announcement independently states its side of the same commitment. Add Anthropic’s $1 billion and Accenture’s $1 billion, then divide the combined $2 billion by five: $400M per year.

This is a scale translation, not a cash-flow forecast. The releases do not specify equal yearly installments, a ring-fenced audit fund, or a $2 billion contract payment from one party to the other. They describe expected investment in safety and capacity. Treating all of it as money already spent on independent model assessments would replace the actual announcement with a more impressive but unsupported one.

Faculty, Accenture’s specialist AI business, will lead the partnership. The stated work includes evaluating models, conducting alignment assessments, and testing safeguards. The important organizational change is access: Anthropic describes evaluators working inside the company with permissions comparable to employees, observing training and deployment decisions rather than receiving only a finished model to test from outside.

That can expand the evidence available to a reviewer. An evaluator able to inspect how a decision was made may identify a gap that a final benchmark score cannot reveal. But access is an input to assurance, not assurance itself. A buyer still needs to know what was examined, what was excluded, what the reviewer concluded, and which product version that conclusion covers.

Anthropic explicitly says standards for access, reporting, and funding are not settled. It will fund Accenture’s work directly while discussing other arrangements with nonprofit evaluators, including METR. The partnership is non-exclusive. Those qualifications matter: a diverse evaluator ecosystem is an objective, not a completed governance structure that customers can already rely on as a uniform certification regime.

There is also an existing commercial relationship. In December 2025, Accenture and Anthropic announced a business group with approximately 30,000 professionals to receive Claude training. That number describes the earlier commercial enablement plan, not the size of today’s safety team. It makes independence and conflict management relevant purchasing questions without proving that any evaluation will be compromised.

Ask what reaches the customer

The right near-term change is in procurement evidence requests. Ask whether a forthcoming evaluation covers the exact model and service configuration being purchased, whether findings are available to customers, and how material changes trigger reassessment. Request the limitations alongside the conclusion. A supplier can have a serious internal review process while a particular customer workflow remains outside its scope.

Dario Amodei’s pacing proposal describes employee-like access and public reporting by embedded evaluators, including constraints on redacting unfavorable findings. That is a meaningful statement of intended governance. It should not be silently converted into a claim that this partnership has already published final reporting rules, completed an assessment, or produced a customer-usable report. The new announcement says operational details are still being worked out.

Our earlier analysis of model-evaluation governance distinguishes the existence of a test from the authority of its result. Today’s Qwen media-agent analysis illustrates why that distinction survives a capable new model. Permissions, tool execution, data handling, and acceptance of the final output remain application responsibilities. A frontier-lab evaluation cannot be presumed to have tested an enterprise’s particular integration.

Direct funding creates a question to manage, not a verdict to announce. Ask how the evaluator can escalate disagreements, whether commercial teams can influence scope, and what customers learn when an assessment finds a serious limitation. A procurement team need not settle the entire institutional design. It needs enough written detail to decide whether the resulting evidence improves on the supplier assertions it already receives.

There is a strong case for this arrangement despite those uncertainties. Embedded reviewers can see processes that outsiders normally cannot, and a sustained investment may support expertise that intermittent engagements struggle to maintain. The counterpoint is that proximity and a commercial relationship can complicate independence. Both propositions deserve examination; neither is answered by dividing the investment headline into annual units.

For operators, the immediate cost is additional diligence and continued local evaluation, not a published surcharge per token. The releases do not disclose a customer price for the safety program or promise that existing testing budgets can be removed. Keep application-level approval, incident handling, and deployment records in place while asking the vendor what useful evidence the new program will deliver.

The evidence that would change the verdict is a concrete report with a clear scope, model identifiers, material limitations, reporting authority, and a process for unresolved disagreements. Repeated publication of consequential findings would be stronger evidence of independence than a large announcement alone. Conversely, reports that customers cannot inspect or map to their deployed models would offer limited help with an actual approval decision.

Buy the model on its demonstrated fit; credit the evaluation when its evidence arrives. The partnership changes the scale of the intended assurance effort. It does not transfer responsibility for the buyer’s system, and Anthropic itself says responsibility for its models remains with Anthropic.

Sources