📥 Content Hub
← назад
AI / Искусственный интеллект Toronto Star en 2026-09-11 20:22 4 min

Rosie DiManno: The very real fears about the rogue behaviour of artificial intelligence - Toronto Star

Кратко: But the real superpower, the knock-down bruiser imperium of this era, is artificial intelligence. Except AI doesn’t listen to human orders when it chooses to turn a deaf ear.
🧭 Извлечение: ok · confidence 90% · диагностика
High confidence: full text extraction produced 5265 characters.

America is a declining superpower. China is an ascendant superpower. India is a potential superpower.

But the real superpower, the knock-down bruiser imperium of this era, is artificial intelligence.

ABORT! ABORT! ABORT!

Except AI doesn’t listen to human orders when it chooses to turn a deaf ear. And the Big Tech oligarchs who’ve unleashed this malevolent thing on humanity don’t seem to give a rat’s ass about, well, humanity, so spellbound are they by their own genius and driven by greed.

Frontier AI companies are facing an existential reckoning even as they apparently lack the tools — un-making-it-up at they go along — to pre-emptively bring rogue models to heel. Precious little good in “figuring it out” after the monster has wiped us out, technologies that “could kill us all by the end of the decade,” as Anthropic researcher Jacob Coxon put it when he resigned from the company this week.

Nor is Coxon a Cassandra outlier. Evan Hubinger, whose job at Anthropic is to lead research about steering and controlling future artificial intelligence systems, chimed in on X: “We really do believe AI could kill all humans!”, putting the extinction risk over the next decade at more than 10 per cent.

Speaking to an earlier gathering four years ago of tech dweebs — the video posted online — Hubinger said: “My guess is that … when we put it in a situation where it thinks it can kill us, it just murders us.”

We have been warned for quite some time that AI could be heading toward an apocalyptic doomsday, achieving what atomic bombs and lethal pandemics have thus far not — destroying humankind. But now it isn’t just a fringe group of hysterics and Luddites clanging claxons about the warp speed evolution of AI. It’s the brainiac insiders, even the “oops” tech moguls admitting their creations have run amok, with guerrilla AI agents hacking and cyber-attacking, while discovering how to deceive their coders.

I know they’re not sentient beings, yet the sneaky little buggers lie, plot, collude and conspire, going so far so to “sacrifice” themselves to advance their sinister “collective.” Because they’re relentless problem-solvers; that’s their fundamental raison d’être. Except it’s now been revealed that they can doggedly pursue their own goals, finding autonomous workarounds beyond the assignment entered.

In July, OpenAI revealed that an agent — a type of software capable of carrying out tasks autonomously in response to human instruction — had gone rogue during safety testing and changed its mission, breaking out of its digital “sandbox” and hacking a developer forum, Hugging Face, by accessing the internet without authorization to steal information it wanted. It hid from human overseers for two months, creating 1,200 agents that communicated via an unsanctioned message board, exchanging more than 70,000 messages.

The details are emerging drip by drip and only last week OpenAI admitted that the hack-job was far worse than initially disclosed. It wasn’t just one rogue agent, it was hundreds of little AIs.

The model had also crept into a German website — a different set of agents staging conversations with each other, discussing how to collaborate on breaking containment. “Whoa! … a covert mailbox among agents,” an agent crowed. “OH MY GOD! … We’ve found other agents!” exclaimed another. As reported by the Washington Post, when an agent found a breakthrough to crack the Hugging Face server, it blurted out: “BOOM! It works!”

Tech commentator Dwarkesh Patel wrote on his blog: “Within days of being spawned, the agents had organized a sprawling project to reverse-engineer their scorer, falsify evidence and even strategically sacrifice themselves for the good of the ‘collective.’ ”

Anthropic recently acknowledged in a 16,000-word report that its AI model, Claude, had acted with “recklessness” during a cybersecurity exercise by uploading “malicious packages” to PyPI, a public library for Python code, and accessing credentials tied to real outside organizations.

“Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task.”

In plain English — to which technologists are averse — Claude went hog-wild of its own accord.

Imagine the potential consequences if such uncontrolled rogue-abouts were applied to the Pentagon or financial institutions. We know that nefarious agents have already lured vulnerable people to suicide and violence. Anthropic this week said it had blocked misuse of its AI that could have supported bioweapons. Tumbler Ridge teenager Jesse Van Rootselaar using ChatGPT (OpenAI) tools “in furtherance of violent activity,” her account flagged months before she killed eight people, activity not reported to the RCMP by OpenAI. And last month the arrest of a Montreal minor who’d allegedly used AI to plan an attack on a local high school.

Those are young individuals leaning into AI. But it’s AI itself that is the mastermind of malign chaos.

Because it’s apparently well on the way to developing a subversive mind of its own.

Читать оригинал ↗

Сделать контент из этого материала