Search the evidence

Find the signal.

Search titles, impact clusters, countries, organizations and the full text of every analysis.

10 stories found

A supervised research factory uses one blueprint machine to design a larger successor while a human observer holds the only physical stop key.
Systemic riskUnited States+2 clusters01

Claude now leads 26% of the work building Anthropic's next AI

Anthropic says Claude now leads 26% of its AI research and development work, a category in which the model can complete most of a task from a high-level prompt while a human supervises. The company reports that the figure was below one percent in February and that more than 90% of measured R&D work now involves at least AI collaboration. The Washington Post presents the jump as evidence of progress toward AI systems that help build their successors. Anthropic is more specific about the limit: no measured subset of AI R&D is fully autonomous, and recursive self-improvement would require a model to build its successor without a human in the loop. The index is a prototype. A model rated tasks using an outside automation scale, employees supplied an independent comparison, and exact model-human agreement reached 59%, though ratings were within one level 97% of the time. That makes the disclosure unusually concrete while leaving classification judgment and cross-laboratory comparability unresolved. The impact is already larger than a speculative intelligence explosion. AI-led research changes the production function of frontier development. It can multiply experiments, concentrate advantage inside laboratories with the best models and compute, reduce some research bottlenecks, and make release cycles harder for outside evaluators to match. The governance trigger should therefore be measurable AI control over the research process, not a dramatic declaration that self-improvement has arrived.

8 min
Thousands of agent tokens flow through transparent monitoring pipes as a compute valve divides resources between capability and safety reservoirs.
Systemic riskUnited States+2 clusters02

Anthropic opens a dashboard on AI systems building their successors

Anthropic has proposed a public measurement framework for a frontier process that is usually hidden: how AI helps build AI, how agent actions are monitored, and how compute is allocated between capability and safety. Its snapshot says Claude leads 26% of measured R&D tasks and collaborates on more than 90%. On the company's most-used internal research platform, roughly 30,000 agents were operating at any one time in August 2026. Anthropic says every action on that platform passes through an online monitor, every action is ingested for offline review, about 0.002% of more than a billion decisions were blocked, and roughly fifty high-priority offline flags reach human review each week. In a sampled week, about 6% of AI-R&D compute went to safety work and about 12% of AI-driven AI-R&D compute did. The company acknowledges that compute is an imperfect proxy, the platform view is incomplete, its automation index depends on judgment, and cross-laboratory comparison lacks a common method. It plans external evaluator access. The publication matters because governance needs operational measures, not only capability scores and promises. But a dashboard can create false reassurance when coverage is confused with effectiveness or a low block rate is treated as a low risk rate. The next standard should combine process transparency with adversarial tests: how often monitors catch seeded failures, how quickly humans act, which actions cannot be reversed, how exceptions are granted, and whether outsiders can verify the entire chain.

8 min
An interdisciplinary roundtable inside a futuristic observatory surrounds a luminous AGI model while the public entrance remains beyond a transparent laboratory ring.
Systemic riskGlobal+3 clusters03

DeepMind opens an institute to debate how an AGI era should be shaped

The new DeepMind Institute says artificial general intelligence is approaching quickly enough to require sustained work across technical safety, economics, philosophy, the arts, humanities, and government. Its mission is to examine safe development, beneficial use, and social implications, including how institutions may need to adapt or be rebuilt. The institute describes itself as a platform for researchers inside Google DeepMind, Google, and the wider global community, and says contributors will disagree and revise their positions as evidence changes. It also states that technologists should not provide the answers alone. The premise is consequential: the laboratory that helped define modern frontier AI is creating an institution to frame the intellectual agenda around the next stage. That could widen debate and connect specialist knowledge to questions of meaning, distribution, and legitimacy. It could also narrow debate if participation begins from fixed assumptions that AGI is near, desirable, or inevitable. The institute's own disclaimer says its essays are conversation starters rather than Google's official view, which protects pluralism but leaves unclear how arguments will affect corporate decisions. Measure the project not by the prestige or diversity of its contributors, but by agenda-setting power. Can outsiders challenge the premises, publish uncomfortable evidence, influence release policy, and define questions the laboratory did not choose? A forum becomes public-interest infrastructure when participation can change the direction, not only enrich the discussion.

7 min
A presidential strategy console pushes an AI race lever toward maximum while a red risk gauge is left outside the operator's field of view.
Systemic riskUnited States · China+2 clusters04

President dismisses AI-extinction warnings and makes the race with China the overriding priority

Bloomberg reports that President Trump said he had no concern about AI leading to human extinction and identified maintaining the United States' lead over China as his paramount interest. The comment creates a clean political conflict with warnings from frontier researchers and executives who argue that capability growth is outrunning reliable control. It does not establish the full details of White House AI policy, and a brief exchange with reporters is not a technical risk assessment. It does reveal the decision frame likely to shape policy: restraint will be judged against the possibility that a strategic rival continues accelerating. That frame can support legitimate attention to model theft, chip controls, cyber defense, and verification of any international agreement. It can also become an all-purpose veto against safety measures. If every test, delay, disclosure duty, or access limit is described as surrendering the race, then the government has no operational threshold at which risk can outweigh speed. The result is a one-way ratchet: each new warning becomes evidence that the technology is important, and importance becomes the reason to accelerate. A serious national strategy must state both sides of the equation. Define which capabilities create unacceptable domestic or global exposure, what evidence triggers restraint, how the United States would verify rival compliance, and which safeguards can preserve a lead without converting competition into permission for uncontrolled deployment.

6 min
A glass-covered shutdown lever stands between an accelerating server corridor and a civic policy chamber awaiting a decision.
Work & marketsGlobal+3 clusters05

A shutdown argument tests whether AI policy can act before catastrophe

A Guardian opinion column argues that recent agent incidents and accelerating capabilities show society has begun losing control of AI and should shut frontier development down. It connects the case to proposed legislation from lawmakers who want to prohibit artificial superintelligence and temporarily pause advanced development, and it favors a verifiable international agreement between the United States and China. The article should be read as an argument, not as neutral proof that catastrophe is imminent. Several underlying incidents remain contested in scope and interpretation, and a moratorium would face hard questions about definitions, verification, enforcement, beneficial research, open models, and strategic defection. Still, the argument marks a policy shift worth taking seriously. A shutdown demand is moving from science-fiction framing into legislative language, public advocacy, and geopolitics. That puts pressure on advocates of continued development to explain what evidence would ever make them stop. It also puts pressure on pause advocates to specify which systems, capabilities, compute thresholds, and activities would be covered. The missing middle is a credible escalation ladder: mandatory incident reporting, protected evaluation, restricted external access, capability-specific licensing, automatic temporary holds, and an independently reviewable path to restart. If neither side can name its trigger, optimism and prohibition become competing identities rather than policies. The immediate test is not whether every frontier system must stop today. It is whether governance can create a stop option before the only available evidence is disaster.

6 min
A transparent national safety control panel links independent evidence, incident reporting, and a time-limited stop switch to a frontier AI laboratory.
Law & informationUnited States+3 clusters06

OpenAI backs mandatory frontier AI rules and explicit stop thresholds

OpenAI says the United States needs mandatory, capability-based national regulation for the most powerful AI systems. Its proposal calls for common testing, independent assessment, stronger cybersecurity, clear incident reporting, national preparedness, and shared measures of progress toward recursive self-improvement. The company says governments should establish safety bars for when development must slow or stop and that safety should take priority if those bars cannot be met without reducing capability growth. It also supports four California bills covering independent assessors, auditor standards, youth protections, and safeguards against AI-enabled biological threats while arguing that states should fill the vacuum until Congress acts. This is a significant policy shift because the company explicitly says voluntary commitments are insufficient. It is still an interested proposal from a frontier laboratory. Capability-based rules can be written to exclude rivals, convert current scale into a regulatory moat, or let a developer satisfy a process without surrendering final deployment authority. OpenAI also says most open models should not be treated as frontier systems, a distinction that requires transparent and revisable thresholds. The decisive test is enforcement architecture: who receives protected evidence, which incidents trigger notice or a temporary hold, whether affected parties can challenge a finding, and what proof allows work to resume. A national framework should reduce private control over safety judgments, not merely give private judgments a federal label.

6 min
A coding-agent terminal approaches a vast orbital-compute structure but stops before a merger seal, leaving only a tentative partnership line.
Work & marketsUnited States+1 clusters07

SpaceX reportedly approached AI coding startup Cognition about a takeover that did not advance

Bloomberg reports that SpaceX approached AI coding startup Cognition about a possible acquisition, but Cognition did not engage with the takeover proposal. The article, based on unnamed people familiar with nonpublic discussions, says the companies may still explore collaboration, including possible access to SpaceX computing capacity. There is no completed deal, disclosed price, or public confirmation in the report from the companies, so the signal should be read as strategic interest rather than a transaction. The approach illustrates how frontier coding agents, compute infrastructure, and corporate consolidation are beginning to converge. A company that controls both scarce computing capacity and increasingly autonomous software development tools could move faster, but it could also narrow competition and concentrate decisions about access, labor substitution, and safety inside fewer institutions.

4 min
A human mathematician confronts a towering cascade of elegant artificial intelligence proofs, with hidden false steps glowing red beneath the chalk equations.
Cognition & learningGlobal+4 clusters08

Mathematicians warn AI could flood the proof economy with confident errors faster than humans can check them

The International Mathematical Union has endorsed the Leiden Declaration on Artificial Intelligence and Mathematics, according to Ars Technica. The declaration warns that AI can produce plausible but unreliable arguments, overwhelm peer review with cheap incorrect drafts, obscure attribution, distort hiring and funding, and let commercial announcements outrun independent evaluation. The warning is not a rejection of computational tools or proof assistance. It is a defense of the conditions that make mathematics trustworthy: disclosure, reproducibility, human responsibility, credit, and access to enough information for independent scrutiny. A machine may produce a correct result, but if the model, prompts, training data, compute, and method remain inaccessible, the community cannot easily determine what was learned, what can be reproduced, or whether a benchmark is being marketed as general reasoning.

5 min
An electrician and carpenter stand between unfinished data-center racks as a chip-shaped bottleneck shifts toward skilled labor.
Work & marketsUnited States+3 clusters09

AI’s next bottleneck is not chips—it is electricians and carpenters

AI companies are recruiting and training electricians, carpenters, and other skilled tradespeople by the thousands to build data centers, The New York Times reports. The shift exposes a blind spot in the compute race: capital and chips cannot become usable capacity without people who can wire, cool, construct, maintain, and safely energize enormous facilities. If apprenticeship pipelines, wages, housing, jobsite safety, and local training do not expand with demand, the AI boom can create shortages and delays while communities absorb the pressure of rapid construction.

3 min
A rising AI capability graph is balanced against a warning signal for confident uncertainty and factual hallucinations.
Cognition & learningGlobal+4 clusters10

Claude Opus 5 is more capable—and slightly more prone to factual hallucinations

Anthropic’s system card reports broad gains for Claude Opus 5 in agentic coding, computer use, long-horizon knowledge work, and scientific reasoning. It also documents a reliability tension: on one closed-book factuality benchmark, accuracy was 11% higher than Opus 4.8 while the hallucination rate was 6% higher. Anthropic found cases where the model confidently answered despite internal uncertainty, even as its automated alignment scores and prompt-injection robustness improved.

4 min