Catastrophic risk arguments and their critiques

One strand of AI safety argues that the systems now being built could produce harm of a kind there is no recovering from, and that this is different from the harms already documented rather than a larger version of them. An equally serious strand argues that the claim is unsupported and that acting on it is costly. Most writing in this field is positioned against that argument whether it says so or not, so it is worth being able to place.

What the argument from capability says

The starting observation is that capability has risen quickly, largely by increasing scale, and that the increases have not been easy to predict before they arrived. If that continues, the argument runs, systems will eventually do things no evaluation was written to look for, and waiting to prepare until then is preparing afterwards.

The word catastrophic rather than serious carries a second claim, about recovery. Ordinary software failures are bad and bounded. A harm that spreads faster than institutions can respond, or that removes the ability to correct it, is argued to sit in its own category and to justify precaution that normal risk management would not.

What autonomy and opacity add

Two further moves do most of the work.

The first is reach. A model producing text for a person to read is bounded by that person. A system given tools, budgets and the ability to act over many steps has effects with no review point in between, so an unchanged error rate produces a different consequence.

The second is verification. A trained model has no specification you can check it against, and its behaviour is known only from the cases somebody tried. The argument holds that you therefore cannot rule out a capability you did not test for, and that generality grows the untested space faster than testing covers it.

The critiques that land

These are not outside objections. They come from people working on the same systems.

It is extrapolation, not evidence. The case is built from a trend line and a chain of reasoning, and no link in the chain has been observed. Predicting where a technology goes from the curve it has traced so far has a poor record in every field it has been tried in.

It competes with harm that is already measurable. Discrimination in automated decisions, surveillance and fabricated claims about real people are happening now, to people with the least recourse. Critics argue that speculative risk absorbs funding, regulatory attention and seriousness that documented harm needs more.

It can serve the people making it. If a technology needs licensing, expensive audits and restricted release, the organisations best placed to comply are the largest incumbents. This is less an accusation of bad faith than an observation about whose position a framing happens to strengthen.

The track record is thin. Confident predictions about which capabilities would arrive when, and about what would follow from them, have been wrong often enough and in both directions that firm expectations about the next decade are hard to justify.

Why the two camps build the same things

The disagreement is about weighting and expectation, rarely about what to do next. Capability evaluation, interpretability, human oversight, red teaming and incident response are the practical programme on either account, because a system whose behaviour you cannot measure or interrupt is a problem whether the worst case is a wrongly refused loan or something larger.

That overlap is why a product person does not have to settle the argument before acting. Nearly everything in this certification is agreed ground.

Placing a paper rather than adopting a position

When you read something here, ask which kind of claim it is making. A harm already observed, an extrapolation from a trend, and a governance conclusion drawn from either are three different things, and papers move between them without marking the change while the confident tone stays constant.

The second question is what would change the author's mind. A position no evidence could revise is a commitment rather than an argument, and that applies to the confident dismissal exactly as it does to the confident warning.

Common misconceptions

The catastrophic risk argument is fringe, and no serious researcher holds it.

It is held by researchers at major laboratories and universities, and rejected by researchers of equal standing. The disagreement is internal to the field, and treating either side as fringe misreads most of what is published in it.

The argument is that AI will become conscious or turn hostile.

The standard version makes no claim about consciousness or intent. It is about a system pursuing an objective competently, with enough reach to matter, in a way nobody can verify in advance.

Taking future risk seriously means treating present harm as secondary.

That trade off is the critics' strongest point and it is not forced. The measures both positions ask for, which are evaluation, oversight and incident response, are the same measures that catch discrimination and fabrication happening now.

Where this is examined
AI Safety Foundation
Fundamentals of AI Safety, 16 per cent of the exam.
Related material
Book
Weapons of Math Destruction, On harm that is already measurable and already falling on people.
Book
Thinking, Fast and Slow, On confident prediction built from a coherent story.
Concepts