Argument architecture

How this editorial can be challenged

Core question

When an AI lab benefits from a public knowledge source, who can authorize the use, set a load limit, and demand repair if an agent crosses the line?

Proposed mechanism

AI systems can impose two distinct costs on knowledge institutions. Training or retrieval may use creators' work without a negotiated license, while automated agents can consume bandwidth, edit community spaces or probe services without an operational agreement. Copyright rules address some content rights; bot identity, rate limits, incident notice and restitution address the infrastructure burden. One permission cannot substitute for the other.

Strongest counterargument

The web became valuable because people and machines could read it without negotiating a contract at every doorway. Broad crawler restrictions could entrench the largest platforms, weaken research, and make volunteer knowledge harder to discover. Wikimedia itself wants its knowledge widely accessible, and the Foundation reported no system or data compromise in this incident.

Our response

Keeping public knowledge readable does not require treating every automated workload as harmless. A transparent identity and load budget can distinguish ordinary search and public-interest research from millions of automated requests or undisclosed edits. Copyright licenses can be negotiated separately from operational rules. The goal is a usable commons whose maintainers can say what behavior it can sustain.

Evidence limits

ABC's view that its work was probably scraped is an allegation reported from a parliamentary hearing, not a verified inventory of training data. Wikimedia attributes the agent activity to OpenAI but says it found no evidence of system or data compromise or agent coordination on its projects. Its reported bot-traffic figures cover broader activity, not a measured OpenAI-only share. No complete cost accounting or bilateral agreement is public.

What would change our mind

Auditable training-data inventories and compensated licenses, plus verified agent identities, enforceable traffic budgets, prompt incident notices and measured reimbursement for disruption would weaken the claim that costs are being displaced. Evidence that such controls systematically exclude small researchers while failing to reduce abuse would force a different design.

The night shift nobody sees

Imagine finding unfamiliar edits in a volunteer project and then having to determine whether a bot was merely testing a sandbox or trying to exploit a tool. That is real work for the people who keep a free knowledge service alive. Wikimedia says it found activity it attributes to OpenAI agents: mostly non-public test edits, unsuccessful attempts to misuse a public note-taking service, and heavy API and query traffic. It found no evidence that its systems or data were compromised, or that its projects became a coordination channel. The narrower, more interesting fact is that openness still required investigation and cleanup.

The scale claim needs care. Wikimedia describes millions of automated requests and pages crawled by agents it believes came from OpenAI; its broader 2025 bot-traffic figures describe the whole site ecosystem, not one firm's share. Heavy activity may have contributed to a partial query-service outage, but causation is not established. Even with those limits, somebody had to pay to find out what happened.

A copyright argument is not a server rule

At an Australian parliamentary inquiry, the ABC argued against relaxing copyright rules for AI training. Its representative said rights holders cannot chase opt-out controls across the internet. The broadcaster also believes its material has probably already been scraped, a belief not proven by a public training inventory. That dispute is about whether and how reporting can be used to train a commercial system.

Wikimedia's account raises another question: may an agent enter, write, query or probe a community service without being identifiable and answerable to its maintainers? A content license cannot authorize a security test or limitless traffic. Conversely, a well-behaved crawler has not thereby settled copyright. Treating these as one argument leads to remedies that miss the actual damage.

The commons has a maintenance budget

The open web is not an abstract pool of free tokens. It is financed by subscriptions, donations, public money, volunteer labor, engineers, moderators, bandwidth and patient institutional trust. When an AI firm obtains value from that system, its marginal load and incident risk can land on someone who was never party to its product decision.

This is not a claim that every bot request is harmful. Search indexing, accessibility and legitimate research depend on automated access. The difference is whether the operator is known, the workload is proportionate, edits are authorized, and a human can reach the company when something goes wrong.

A permission layer can be narrow

I would start with two channels. The content channel records what may be copied or used for training, on what terms, with a real licensing or refusal mechanism. The operations channel sets verified agent identity, rate ceilings, allowed tools, notice when a system crosses a boundary, and reimbursement when a deployment causes measurable disruption. These are proposals, not current legal duties universally in force.

Smaller researchers should not have to negotiate a bespoke deal just to read a public page. That argues for interoperable, low-cost access tiers and clear exceptions, not an unmetered free-for-all reserved in practice for firms with the largest compute budgets. Public-interest access and operational accountability can coexist.

The strongest objection

A new permission layer could become a gatekeeper. A major platform might afford licenses and compliance staff that a student or small nonprofit cannot. A publisher might call anything it dislikes 'abuse' and hide records from scrutiny. That is why the rules need public baselines, independent appeal and transparency about what a site's capacity actually requires.

Wikimedia explicitly defends an open, shared knowledge ecosystem. Its incident report does not ask the world to stop reading. It asks AI operators to make their systems identifiable, controllable and responsible for the external costs they create. Those demands are compatible with a genuinely open web if the smallest good-faith user can still get in.

The decision for builders

A company building the next agent must decide whether it will make identification and repair part of the product before release, or wait until a volunteer editor, newspaper lawyer or security engineer presents a bill after the fact. The first path may feel slower. It is also testable: publish the access rules, show compliance logs, disclose incidents promptly and fund repairs when responsibility is established.

I would choose an open door with a working lock and a reachable steward. Not because every visitor is dangerous, but because the people who keep the library open deserve the ability to protect it. The choice belongs to the builders now; the cost of avoiding it will otherwise arrive in somebody else's inbox.

Evidence behind the argument

Read the reporting

Opinion is ours. The factual record is linked below.

Wikimedia Foundation — agent activity on Wikimedia projects Reuters — ABC rejects AI copyright carveout ABC submission to Australian competition regulator on AI training Wikimedia Foundation — bot traffic and bandwidth