Anthropic and Accenture announced on September 18 that they plan to build a team of outside evaluators inside Anthropic, the company behind Claude. Each company says it expects to invest at least $1 billion over five years in related AI evaluation and safety capacity. With access comparable to an employee’s, the reviewers could check Anthropic’s safety promises against decisions made while models are being built, not just information the company releases afterward.

Faculty, Accenture’s specialist AI business, will lead the work. The announcement advances Anthropic’s earlier pledge to embed outside evaluators by naming a partner and its intended responsibilities. The announcement describes a planned team, not one already operating or publishing findings.

At least $1 billionAnthropic’s expected five-year capacity investment
At least $1 billionAccenture’s expected five-year capacity investment

Those figures need careful reading. Anthropic describes investment in building evaluation capacity; Accenture describes investment in AI safety. Neither announcement specifies a cash payment, a budget reserved for this team, or how the spending will be divided among activities. There is no disclosed headcount, start date or milestone schedule. Anthropic separately says it will directly fund Accenture’s work, without stating an amount.

The proposed team would evaluate models and conduct red-teaming, deliberate attempts to expose failures or dangerous behavior. It would also perform alignment assessments, examining whether models behave as intended, and test safeguards designed to prevent harmful use. Evaluators would work alongside Anthropic’s internal teams and safety partners.

The most consequential promise is access comparable to an employee’s. Anthropic says embedded evaluators could observe training, follow decisions about how models are built and deployed, and speak directly with staff. That could let them compare a stated safety commitment with the decisions actually made inside the lab. For people deciding whether to trust AI in their work or public services, the potential benefit is evidence beyond a company’s own account.

That logic is consistent with research on rigorous third-party AI auditing, which emphasizes access to nonpublic systems and safety decisions. But the partnership announcement does not specify which records, systems or model versions Faculty will see. Access alone also supplies neither enforcement authority nor a guarantee that findings reach the public.

Who controls what the evaluators can say?

What would make the evaluation independent?

The voluntary AEF-1 standard for third-party AI evaluations offers concrete tests: sufficient access and time, evaluator control over scope and methods, and compensation that does not depend on the result. Publication protections should establish editorial control, timely disclosure rights and limits on redaction, including disclosure of who can authorize it. Conflict policies and organizational control matter alongside those rights. AEF-1 addresses third-party evaluations generally, and no public assessment of this arrangement against it was identified.

Anthropic acknowledges that standards for embedded evaluators’ access and public reporting do not yet exist, and that funding arrangements remain unsettled. It favors pooled or government funding over the longer term. Its June Advanced AI Framework also recommends funding disclosures and safeguards against companies shopping for favorable evaluators.

Here, Anthropic says it will directly fund Accenture’s evaluation work. Accenture also has a commercial interest in Claude’s adoption: the companies announced a dedicated Accenture Anthropic Business Group in December 2025, with approximately 30,000 Accenture professionals slated for Claude training. That was a training plan, not a reported completion count. These ties create a conflict that needs managing; they do not establish that anyone has compromised an evaluation.

In his earlier essay, Anthropic CEO Dario Amodei said evaluators should be able to publish key findings without Anthropic’s editorial control, subject to limited redactions, and disclose material redactions. The September 18 announcement does not establish that those protections are in Faculty’s agreement. It supplies no publication review deadline or terms governing disclosure after the relationship ends. That is an unresolved contractual question, not proof that protections are absent.

We would judge independence by whether evaluators can choose consequential tests, report unfavorable findings on a predictable schedule, and retain those rights if funding or the partnership ends. Disclosed reporting lines and separation from Accenture’s Claude sales work would help readers assess whose interests govern the team. Anthropic plans to keep training and releasing models alongside this work; the decisive test is whether its evaluators can tell the public something their paying client would prefer it did not hear.