Anthropic's Three Frontier AI Metrics: AI R&D, Agent Oversight and Compute Allocation
What do Anthropic's three frontier AI metrics mean for an agency?
They are a disclosure format, not a standard: they describe how a lab builds models, not what those models do in your client's stack. Published on 17 September 2026, they cover how much of a frontier lab's AI R&D its own models perform, how its agents are overseen, and where its AI R&D compute goes. The structure — coverage, review latency, escalation rate — is the part an agency can use.
What Anthropic published on 17 September 2026
On 17 September 2026 the Anthropic Institute published three measurements it wants the public to hold beside its internal reporting: “we’re sharing three measurements that will help the public track AI development inside frontier AI labs: how much of AI R&D is performed by AI itself, how well the actions of AI agents are overseen, and how compute is allocated”. Each came with an internal snapshot and a published methodology.
The reason given is a disclosure gap: “As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows.” Nothing about the format binds a buyer, a vendor or an agency, and no regulator has adopted it.
The sentence to carry into a client meeting: these measurements describe how a lab builds models, not what those models can do. What follows is the part an agency can act on — the three questions, turned into a screen for any vendor's deployment.
The three metrics, in Anthropic's own words
Three measurements, three scales, three denominators. A vendor who answers one has not answered the other two.
Metric 1: how much of the lab's AI R&D is done by AI
The R&D Automation Index rates AI involvement from AL0 to AL5, using Epoch AI's automation taxonomy. The levels an agency meets in delivery work are the middle ones:
“In AL3, AI “collaborates”: it can do large chunks of work under close human direction. In AL4, AI “leads”: it can complete most of the task end-to-end from a high-level prompt, while the human supervises.”
In August 2026 Claude sat at “leads” (AL4) for 26% of Anthropic's AI R&D work, over 90% of measured work was at or above “collaborates” (AL3), and the basket is frozen at July 2026. The ceiling is stated, not implied: “Claude is not operating fully autonomously for any measured subset of AI R&D work.”
Metric 2: how well the lab's agents are overseen
Three sub-metrics, three denominators: coverage (the share of an agent's actions passing a monitor), review latency (action to review, automated then human) and escalation rate (the share blocked, redirected or flagged). In August 2026 Anthropic's most-used internal platform ran roughly 30,000 agents; every action passed the online monitor before execution, blocked actions get human review within a week, about 100,000 transcripts a week are flagged and roughly 50 reach humans. The blocking rate carries its denominator: “We analyzed over a billion decisions from our research & engineering agents over August 2026 and found that 0.002% of them (about 1 in 47,000) were blocked by this monitor.”
Complete coverage means the actions went through a monitor — not that a human saw them. The 0.002% is a blocking rate over monitored actions, in one platform, over one month: not an error rate, not a misalignment rate, not a benchmark.
Metric 3: where the lab's AI R&D compute goes
The third is a share of one budget, not of total company compute: “Over the examined week, about 6% of compute that went to AI R&D was allocated toward safety, and about 12% of compute that went to AI-driven AI R&D was allocated toward safety.” That week is 13–20 July 2026, and Anthropic labels the estimate down rather than up: “These are deliberately conservative estimates.” It answers how much of the AI R&D budget went to safety work, not how much of the company's compute is spent on safety. Where the compute sits and what it costs are separate questions, carried by the Theseus/AWS build-out and the $45B compute deal.
What each metric misses: Anthropic's own caveats
The strongest critique of these numbers is in the same post. Quote it back rather than discover it in a client meeting.
- The judge is the lab's own model. “we’re using our own models to evaluate our systems, which could mean that the “judge” model could make the same kinds of errors as the model it is checking”. Exact agreement was 59%, against 35% human-to-human.
- The basket is frozen. The index rates Anthropic's own catalogued tasks, fixed at July 2026: a rising number says the work humans did then is being automated, not that new work appeared.
- One week is not a trend. “the measurement covers one week, which is enough to show that the measurement can be made, but not enough to show a meaningful trend”. The same limit applies to the coverage, latency and escalation figures.
- Spend is not effort. “Third, and most importantly, compute share measures only what is spent.” “Safety research tends to use less compute than frontier training runs by its nature, so compute is an imperfect proxy for how much a company focuses on safety.” And the line stays drawable by the lab: “Safety research is hard to distinguish from capabilities research, and each developer will be tempted to draw the line generously.”
Anthropic names what stands between these numbers and a comparable figure: “Two obstacles stand in the way of cross-lab comparison on this type of reporting. First is the lack of a common methodology.” The second obstacle is a lab grading its own work — which is why outside evaluation, not the metric, is the thing to watch. The compliance-audit version of that question starts at the AI safety compliance audit.
Where the metrics came from: the 12 September essay and the evaluator bet
Five days earlier, on 12 September 2026, Dario Amodei published “We Must Pace the Frontier”; of its three steps only the first is Anthropic's own commitment: “The first step is something Anthropic is unilaterally committing to (and calls on governments to require other frontier companies to match).” That step is an embedded-evaluator model — a standing team of third-party evaluators inside the lab, with employee-like access to internal processes.
The reception matters because clients will ask. It drew support from OpenAI's Sam Altman (“I agree with Dario that we need to pace the frontier”), and the strongest objection belongs in the same breath: “Tech CEOs banding together is an old ruse recycled from corporate America to get a pass from antitrust laws”. That antitrust argument is the counter-case to any “the industry is coordinating on safety” framing. The chronology lives on the frontier AI pacing roadmap.
One day later the evaluator step moved: “Anthropic and Accenture each expect to invest at least $1 billion in building capacity in this area over the next five years.” Announced is not operating, and Anthropic's own admission is the line to keep when a vendor claims independent evaluation: “There are, as yet, no standards for what information embedded evaluators should have access to, or how they should report what they find.” The buyer-side read of the same debate is AI slowdown: what it means for business.
The six-question procurement screen for agency clients
These six questions are adapted from Anthropic's own sub-metrics and apply to any agent deployment. Ask for the answers in writing; an answer without a denominator is not an answer.
- What share of agent actions is monitored before execution, over what denominator? A good answer names the platform, the action count and the window. “All agent activity is monitored” is a non-answer: it names a monitor, not a rate.
- Is review automated, human, or both — and within what window? Anthropic's split is the benchmark: automated review before execution, human review of blocked actions within one week.
- What share of actions were blocked, redirected or flagged last quarter? Insist on Anthropic's denominator — actions monitored — not a count of incidents, which can be driven to zero by counting nothing.
- Who signs off when a block is overridden, and where is that recorded? An override with no named approver and no log entry is an unmonitored action with extra steps.
- Can we see an escalation log for our account? One real redacted entry before signature tells you more than any policy document.
- What happens when the monitor is wrong? Ask how it was last tested, by whom, and what happened to the actions it missed.
One line makes the screen stick — the principle behind Anthropic's own safety-share caveat: “The burden of proof should sit with the developer to show that work is safety-related.” The client-facing version of the same conversation is the AI agent human checkpoints model.
Auditing a vendor's oversight claims
A screen produces answers; an audit produces evidence. Seven artefacts are worth asking for, each mapped to a sub-metric: the coverage denominator and whether the monitor sits before or after execution; the review window for blocked actions; block, redirect and flag shares over a stated period; the override sign-off and its log; a redacted escalation entry; the monitor's failure mode; the date each figure was last verified.
Outside checking has a precedent, and it is not a badge: METR red-teamed Anthropic's internal agent monitoring and found real problems — “The exercise discovered several specific novel vulnerabilities, some of which have since been patched, and none of which severely undermine major claims in the Opus 4.6 Sabotage Risk Report.” That is the test to ask whether anyone has run. And the accountability line to put in a contract review is Anthropic's own: “To be clear, independent embedded evaluators do not reduce our accountability, but help to make it more verifiable. The safety of our models remains our responsibility.”
Where that evidence is filed — who owns the agent identity, who can revoke it, what a reviewer may see — is the governance half: see the governance questions an agency already answers and, on the audit side, AI agent identity governance. To have the claims checked rather than described, run them through the free AI audit tool before the next client review.
The measurement stack: what to read next
Four boundaries, four different jobs. The pacing roadmap owns the policy chronology; the governance page owns the client-conversation questions these numbers now answer; Theseus and the AWS build-out and the $45B compute deal own compute economics, not the safety share. On the audit side, human checkpoints, agent identity governance and the safety compliance audit own the instruments; this page supplies the published sub-metrics behind them. To check a real deployment, start with the AI readiness assessment guide.
Frequently asked questions
What are Anthropic's three frontier AI metrics?
On 17 September 2026 the Anthropic Institute published three measurements for tracking AI development inside a frontier lab: how much of its AI R&D is performed by AI itself, how well the actions of its AI agents are overseen, and how compute is allocated. Each came with a snapshot from inside the company and a public methodology. Anthropic proposes the format as a shared standard; nothing about it is in force.
How much of Anthropic's AI R&D is done by AI?
As of August 2026, Anthropic's R&D Automation Index puts Claude at "leads" (AL4 on Epoch AI's scale) for 26% of its AI R&D work: the model completes most of a task from a high-level prompt while a human supervises. Over 90% of measured work sits at or above "collaborates" (AL3), and Anthropic says Claude is not operating fully autonomously for any subset it measured.
What do Anthropic's agent-oversight metrics measure?
Three things. Coverage: 100% of the actions of roughly 30,000 research and engineering agents pass through an online monitor before execution. Review latency: automated review runs before an action, humans review blocked actions within one week, roughly 100,000 transcripts a week are flagged, and about 50 a week escalate to human review. Escalation rate: of over a billion agent decisions analysed for August 2026, 0.002% (about 1 in 47,000) were blocked.
What share of Anthropic's compute goes to safety work?
In a snapshot week (13-20 July 2026), about 6% of the compute that went to AI R&D was allocated to safety work, and about 12% of the compute going to AI-driven AI R&D was. Anthropic calls these deliberately conservative estimates, counts work that serves capability as much as safety as capability, and warns that compute is an imperfect proxy because safety research is less compute-intensive by nature.
Do these numbers prove that a model, or a vendor, is safe?
No, and Anthropic says so itself. The measurements describe how models are built, not what they can do, and the post lists its own limits: the judge model agreed with humans exactly 59% of the time, the task basket is frozen at July 2026, the compute figure covers one week, and workload labels are best-effort. Our read (analysis): treat them as a disclosure format that makes claims checkable, not as an audit opinion.
Do Anthropic's metrics apply to my business or to an agency's clients?
Not as an obligation. They are voluntary disclosures about frontier-lab R&D, not rules that bind buyers, agencies or vendors, and no regulator has adopted them. They are useful as a template: the three questions they answer for a lab (how automated is the work, who oversees the agents, where does the compute go) are the questions a buyer should ask about any agent deployment. That is our recommendation (analysis).
What should I ask an AI vendor about agent oversight?
Six questions, adapted from Anthropic's own sub-metrics: what share of agent actions is monitored before execution? Is review automated, human, or both, and within what window? What share of actions were blocked or flagged last quarter, on what denominator? Who signs off when a block is overridden? Can we see an escalation log for our account? What happens when the monitor is wrong? Ask for the numbers in writing.
Sources
- Anthropic Institute, Measuring the pace of AI development (2026-09-17): anthropic.com
- Anthropic, independent evaluation partnership with Accenture (2026-09-18): anthropic.com
- Dario Amodei, We Must Pace the Frontier (2026-09-12): darioamodei.com
- Anthropic announcement post (2026-09-17): x.com
- CNBC on the three metrics (2026-09-17): cnbc.com
- METR, red-teaming Anthropic's agent monitoring (2026-03-25): metr.org
- Epoch AI, toward an ONET for AI R&D: epochai.substack.com
- The Guardian on labs coordinating on pacing (2026-09-16): theguardian.com
Accuracy note: every quoted span is verbatim from the sources above and was verified against saved copies on September 18, 2026. The figures are Anthropic's own, over stated snapshots (July and August 2026). Nothing here is a standard or a certification.