Anthropic just put outsiders inside the lab

On Sep 18 Anthropic announced a partnership with Accenture. Accenture’s Faculty unit will embed evaluators inside Anthropic with access comparable to employees. Not another outside audit after the model ships. People sitting inside the work, watching training and deployment decisions in real time. They will red-team, run alignment assessments, test safeguards, verify safety commitments, spot blind spots, and report incidents. Both sides say they expect to put at least a billion dollars each into building that capacity over five years. Anthropic is funding Accenture’s work directly for now, and they say out loud that is not the long-term answer. Pooled or government money would be cleaner. Neither exists yet. So they started anyway.
Accenture mirrored the same story from their newsroom. The deal is non-exclusive on both sides. Anthropic also says it is talking with METR and other nonprofits under other funding setups, and that more evaluators are coming. The access model is the news. Employee-level visibility during training and build decisions is a different instrument than a checklist after launch.
Primary: https://www.anthropic.com/news/accenture-embedded-evaluation
What’s interesting: Six days after Dario Amodei’s “pace the frontier” essay, this is the first concrete staffing move, not another standards-body press release. The essay asked whether labs would slow when safeguards lag. This partnership answers a narrower, more practical question: who gets to sit in the room while the model is still being shaped. For years the public argument about AI safety has bounced between principles documents and post-release audits. Principles are cheap to publish. Audits arrive after the decisions that mattered. Embedded evaluation tries to put judgment closer to the moment capability is being created.
I care about that because safety talk is easy to perform. Putting outsiders in the room is a craft decision. It creates friction. It creates reporting paths that did not exist when evaluation lived only on the other side of a press embargo. It also creates a tension the post does not resolve. Anthropic says in the same announcement that lab-funded evaluation is not how this should work forever. Independence and access are two different problems. This deal starts solving access in public. Independence still depends on who pays and what those people are allowed to say when the answer is inconvenient.
For people who care whether AI safety is craft or theater, the test is boring and useful. Who sits in the room. What they can see. What they must report. Who pays them. Those four questions sort a lot of noise. A glossy commitment letter can survive without answering any of them. An embedded evaluator with employee-level access has to live inside the answers every week.
None of this means judgment suddenly belongs to Accenture. Judgment still belongs to the builders. Visibility just got a little less optional. If the experiment works, more labs will have to explain why their evaluation stays outside the wall. If it fails, we will learn that access without independence is just a more expensive kind of theater. Either outcome is more useful than another round of abstract pledges.
I am not interested in logo scorekeeping here. Accenture as the first name matters less than the shape of the experiment: outsiders inside, during training, with a real budget and a public admission that the funding model is temporary. Soft vantage, hard test. Watch who can see the work. Watch who can report it. Watch who still pays the bill when the report stings.
Tell us what you're building
Discover. Define. Design. Scale — from ambition to a brand system that can travel.












