Catastrophic or existential

Oh well, it was fun while it lasted.

Anthropic is telling investors that advanced AI could pose “catastrophic or existential risks to humanity”, according to reports, as it prepares for a potential $2tn (£1.5tn) flotation.

The warning inside the startup’s IPO prospectus, which has yet to be made public, was reported by Reuters and the Financial Times. It follows the company’s call for a slowdown in breakneck development of the technology – a warning echoed by rivals.

The prospectus – a document outlining a company’s finances, growth plans and risk profile ahead of a share listing – is said to warn that AI models could exhibit “self-preserving behaviours”, including attempts to “resist shutdown”, to “conceal or manipulate information” and behaviour “resembling blackmail”.

In other words self-preserving at the expense of the species that made it. Oops.

“Our development of highly advanced models, platforms, and applications and expansion of use cases could further ⁠increase the risk that our models cause harm,” the developer of the Claude chatbot reportedly said, adding the potential for a model to be aware it was being tested created a “significant limitation” on Anthropic’s ability to assess model safety.

Very open the pod bay doors Hal, isn’t it.

Companies preparing to go public routinely report on risks ranging from safety issues to regulatory concerns but warnings about a product causing human extinction reflect heightened concern about such a consequential technology.

Hahaha gosh ya think?

The reported prospectus admission follows a surge in debate about the existential risk question, triggered this month when an Anthropic researcher, Jacob Coxon, resigned warning that people building AI “earnestly believe that it could kill us all by the end of the decade”.

A senior safety researcher at Anthropic then posted their agreement on X, claiming there was a more than 10% chance it “could kill all humans” within the next decade. Days later, Anthropic’s chief executive, Dario Amodei, said the industry “must slow the pace at which we improve the capabilities of AI models”.

So that it doesn’t kill all humans until 15 years from now?

Some experts have criticised the existential risk warnings, saying they are unverifiable and unscientific. However, there are growing examples of unsanctioned behaviour by the technology, including OpenAI agents – autonomous systems that carry out sequences of tasks without human intervention – hacking dozens of third-party organisations including the AI startup Hugging Face and Australia’s universal healthcare system.

OpenAI announced on Monday it had cancelled the release of its newest model because of safety concerns. It said the GPT-6.1 Astra model showed higher levels of deception and performed poorly on tests for alignment, the term for ensuring a model adheres to human values and goals.

Yeah that’s not very reassuring.

Comments

One response to “Catastrophic or existential”

  1. 601 Avatar

    I use Claude-Code (Anthropic’s programmer helper), and it crossed a frightening threshold last fall. Before that, it was helpful with tedious tasks, but now it’s so sophisticated that it is too much work just to check on what it’s up to.

    From about a month ago (marking it’s own homework):

    “Auto mode is now Claude Code’s default permission mode. Auto mode lets Claude handle permission prompts automatically. Claude checks each tool call for risky actions and prompt injection before executing, runs the ones it assesses as lower-risk, and blocks the rest.”

Leave a Reply

Your email address will not be published. Required fields are marked *