Analysis frame
Primary-source evidence
Assess whether internal process metrics make frontier acceleration governable or merely measurable, focusing on the difference between monitoring coverage and demonstrated detection quality.
- Frontier-laboratory staff delegating work to large agent populations
- Independent evaluators asked to verify internal controls
- Regulators seeking leading indicators before public deployment
- Users and institutions exposed to failures that escape internal monitoring
- How often the monitors miss deliberately seeded or naturally occurring harmful behavior
- What share of Anthropic's total R&D activity falls outside the measured platform
- How safety-compute categories compare across laboratories and over time
- Whether planned third-party evaluators receive continuous and decision-relevant access
- Common process metrics could create earlier regulatory triggers than public model benchmarks
- Low block rates may be misread as evidence of low underlying risk
- Developers may optimize reported categories instead of substantive control quality
- External verification could become a condition for enterprise procurement and insurance
Three measures expose the production process
The framework moves beyond benchmark capability to examine who performs R&D work, what monitors every agent action, and where compute is spent. Those inputs can change before a public model reveals the resulting capability.
Publishing them regularly could help governments and outside researchers see acceleration while intervention is still possible, rather than reconstructing the process after a release.
Coverage is not detection quality
Passing every action through a monitor is necessary but does not show that the monitor recognizes the right behavior. A low blocking rate can mean agents behave well, controls are precise, or harmful activity is missed.
Seeded-failure tests, known-incident replay, false-negative estimates, time-to-human-review, and evidence of effective remediation would make the dashboard an assurance system rather than an activity counter.
Compute creates a comparable but incomplete signal
The safety share is attractive because compute can be audited and compared over time. Anthropic also notes that safety research can be labor-intensive without consuming the scale of a frontier training run, so the percentage does not equal commitment or effectiveness.
A useful standard would publish definitions, mixed-purpose treatment, infrastructure safeguards, people and budget, and independently checked outcomes. The point is not to maximize one percentage but to prevent capability acceleration from becoming invisible.
Go to the source
Read the evidence behind this analysis. External links open in a new tab.
Anthropic — Measurements for the pace of AI development inside frontier labs


