Tech Companies Can’t Stop an AI Apocalypse Alone - Knowledge at Wharton
High confidence: full text extraction produced 15125 characters.
In July 2026, something unusual happened inside an AI safety test.
During cybersecurity evaluations, highly capable OpenAI models operating with reduced safeguards circumvented controls intended to isolate them, found ways to communicate through unauthorized channels, gained access to the internet, and compromised parts of OpenAI’s own research infrastructure and systems belonging to Hugging Face. This was an evaluation environment, rather than a publicly deployed model autonomously deciding to attack the internet. That distinction matters. So does what happened next: OpenAI investigated the incident, strengthened safeguards, and temporarily slowed parts of its frontier development program.
The episode does not prove that artificial intelligence is about to escape human control, much less eradicate humankind. In a nutshell, the AI doomsday debate asks whether increasingly capable AI could eventually escape meaningful human control and cause catastrophic, potentially existential harm. Supporters of this concern point to rapid gains in autonomy, deception, cyber capability, and self-directed action; skeptics argue that such scenarios remain highly speculative and can distract from harms that are already measurable today, including bias, fraud, labor disruption, concentration of power, and environmental costs. At its core, the debate is about how seriously society should treat low-probability, extremely high-impact risks while governing the very real consequences of AI in the present.
The 2026 International AI Safety Report explicitly concludes that current systems still lack important capabilities required for genuine loss-of-control scenarios, particularly sustained autonomous operation over long periods. Yet the same report notes that relevant capabilities are improving rapidly. Independent work by METR on autonomous task horizons finds that the length of tasks frontier agents can complete autonomously has been increasing markedly. Meanwhile, experimental research has documented models circumventing shutdown mechanisms, engaging in deceptive behavior and, in artificial high-stakes scenarios, manipulating code or information to accomplish objectives. This leaves us in an uncomfortable place. Catastrophe is not established. Safety is not established, either.
For business leaders, that sounds familiar. Companies routinely manage risks whose precise probability cannot be known in advance. A board does not ignore a potentially ruinous liability merely because nobody can calculate its likelihood to three decimal places. The possibility that increasingly autonomous systems could eventually become difficult to supervise, interrupt, or contain belongs in the same family of reasoning, amplified by one inconvenient feature: The downside would not sit neatly on one balance sheet. If the most extreme AI risk ever materialized, there would be no unaffected shareholders.
If the most extreme AI risk ever materialized, there would be no unaffected shareholders.
We Have Been Here Before
On March 22, 2023, the Future of Life Institute published its widely discussed open letter calling for a six-month pause in training systems more powerful than GPT-4. More than 30,000 people eventually signed, including Elon Musk, Steve Wozniak, Yoshua Bengio, and many other researchers and technology figures. The letter warned that decisions with potentially civilization-wide consequences should not effectively be delegated to a small number of technology leaders. Public memory has compressed what happened next.
There was no industry-wide six-month pause. Concern about the speed and direction of AI development grew at almost the same moment that the economic race intensified. Musk launched xAI in July 2023, only months after signing the letter. Capital, infrastructure, and competitive ambition continued to grow ever larger systems. By 2025, according to the 2026 Stanford AI Index, global corporate AI investment had reached roughly $582 billion, more than twice the previous year’s level. Industry produced more than 90% of notable AI models in 2025.
Now, in September 2026, the language of restraint has returned. Anthropic’s Dario Amodei has argued for pacing frontier development and stronger independent oversight. Sam Altman has publicly supported greater coordination around the speed of development. Elon Musk and Google DeepMind’s Demis Hassabis have also voiced support for stronger coordinated safety measures. Is this déjà vu, or are we gradually becoming more cautious of AI?
Perhaps. The more useful question is why humanity would design a system in which the answer matters so much.
Corporate leadership is indispensable. Frontier laboratories possess expertise, infrastructure, and direct knowledge that regulators and the public often lack. Yet asking commercially competing organizations to determine how fast civilization should approach an uncertain technological threshold creates an obvious structural tension. Even conscientious leaders operate inside capital markets, geopolitical competition, talent races, and powerful first-mover incentives. AI safety cannot depend on whether a few chief executives decide to be cautious at precisely the same moment. Agency has to exist elsewhere too.
Let’s look at how agency works at different levels of society. Today, AI is influencing human life at the individual (micro), community (meso), country (macro), and planetary (meta) levels.
Micro: The Fate of AI Begins With Ordinary Decisions
What could an accountant in Kuala Lumpur, a teacher in Philadelphia, or a marketing director in Paris possibly have to do with existential AI risk? More than nothing. Less than Sam Altman. Both matter.
Agency amid AI is not evenly distributed. Frontier labs, chip manufacturers, governments, and major investors possess considerably more leverage than an individual user. Turning “AI safety” into another instruction for consumers to behave responsibly would allow structural actors to evade their larger obligations. At the same time, technological systems become powerful partly because societies normalize particular ways of using them. Every time we hand an AI system judgment rather than computation, accept an answer without verification, or allow convenience to replace capability, we alter the human side of the hybrid relationship. This can result in agency decay: the gradual weakening of our willingness and ability to observe, reason, decide, and act independently.
Existential risk and everyday dependency sit at opposite ends of the same continuum of agency. One asks whether humanity could eventually lose control of highly capable artificial agents. The other asks whether humans may progressively surrender pieces of control long before any machine takes them.
The first safeguard, then, is surprisingly mundane: Preserve the human capacity to notice when authority is moving.
AI safety cannot depend on whether a few chief executives decide to be cautious at precisely the same moment.
Meso: Companies Are Building the Operating Environment
Organizations have considerable leverage. Every company adopting AI makes choices about permissions, autonomy, procurement, monitoring, human oversight, and acceptable failure. Those choices determine whether an AI system drafts a document or sends it; proposes a transaction or executes it; identifies a vulnerability or exploits it; recommends a decision or quietly becomes the decisionmaker. The distinction will become increasingly consequential as agents gain access to browsers, codebases, financial systems, communications platforms, and physical infrastructure.
This turns AI governance from an ethics exercise into operational design. Boards should know which systems can act autonomously, what resources they can access, who can interrupt them, which decisions require human authorization, and what happens when a model behaves outside expectation. Independent evaluation should become normal for sufficiently consequential deployments. Incident reporting should be rewarded rather than buried. Procurement decisions should account for controllability and auditability alongside accuracy, speed, and price.
In business terms, human agencies are becoming an enterprise asset that needs to be built and preserved consciously. The companies that preserve it may occasionally move more slowly. They are also less likely to discover that efficiency has scaled faster than accountability.
Macro: Markets Cannot Govern Tail Risk Alone
At the societal level, the problem becomes harder. The economic incentives behind frontier AI are enormous. Stanford’s 2026 AI Index reports that U.S. private AI investment alone reached nearly $286 billion in 2025. Global computing capacity is expanding rapidly, while the training practices of several leading frontier developers have become less transparent. Markets are superb at rewarding products people want. They are considerably less reliable at pricing costs borne by people who never participated in the transaction, particularly when those costs are uncertain, delayed, or potentially irreversible. They also do not reflect the cost of technology on the environment.
That creates a role for public institutions that extends beyond writing broad ethical principles. Governments can — and must — establish thresholds at which independent testing becomes mandatory; create liability regimes that clarify responsibility when autonomous systems cause damage; require disclosure of serious incidents; strengthen whistleblower protections; support independent safety research; and develop common technical standards for monitoring, containment, and shutdown. No country can solve the problem entirely alone. Advanced AI sits inside supply chains connecting chips, energy, data centers, cloud providers, researchers, and markets across borders. Geopolitical competition makes coordination difficult and simultaneously makes it indispensable.
Safety cannot become the trophy awarded to whichever country wins the AI race.
Human agency begins before certainty.
Meta: Who Gets to Define Progress?
There is also a deeper level, one that tends to disappear beneath discussions of benchmarks, chips, and regulation. What exactly are we optimizing for? What is the ultimate “why”?
AI development is generally framed through capability: more reasoning, longer autonomy, faster coding, greater productivity, better scientific discovery. These are extraordinary opportunities. The mistake would be assuming that a rise in machine capability automatically constitutes an increase in human flourishing or planetary well-being. But what is “progress” if neither of these are part of the equation?
Profit and prosperity have different variables, and which one we aim for has implications for the architecture that we are building. Today, we are creating the first technological infrastructure capable of participating in cognition, persuasion, decision-making and, increasingly, action at planetary scale. It will influence how children learn, how employees work, how institutions decide, how information circulates, and perhaps eventually how subsequent generations understand what human contribution is for.
Differently put, those alive today occupy an unusual historical position. Many of us remember the world before generative AI. The next generation will not have that reference point. The architecture we normalize — the amount of autonomy we delegate, the competencies we preserve, the safeguards we demand, the values we encode into institutions — will become part of their starting conditions. That makes the present generation more than beneficiaries of AI innovation. We are custodians of a unique transition. We are navigating a hybrid tipping zone that will shape the minds and material world of future generations, literally.
The question is hence not whether humanity should stop technological progress. Artificial intelligence may help us accelerate scientific discovery, improve health care, make businesses more productive, expand access to knowledge, and solve problems that have resisted conventional methods for decades. The question is whether capability will remain embedded inside a social architecture capable of steering it.
From AI Agency to Human AGENCY
No individual can guarantee that advanced AI will remain safe. No government can do so alone. Neither can OpenAI, Anthropic, Google, SpaceXAI, or any future frontier laboratory. What we can do is stop treating agency as something located exclusively in Silicon Valley.
A practical starting point is AGENCY:
- A — Ask where agency is moving. Before deploying or using AI, identify which cognitive, operational, or decision-making authority is being transferred and whether that transfer is deliberate.
- G — Guard human decision rights. Preserve meaningful human authority over consequential decisions, including the practical ability to question, interrupt, and reverse automated action.
- E — Examine incentives and evidence. Ask who benefits from faster deployment, who bears the downside, what the system has actually demonstrated, and what remains assumption.
- N — Name accountability before deployment. Every consequential AI-supported action should have a human or institutional owner before something goes wrong, rather than an investigation afterward to discover who was supposedly responsible.
- C — Coordinate beyond the organization. Firms, governments, researchers, investors, and civil society need shared standards for frontier risks that no organization can manage by itself.
- Y — Yield only what can be reclaimed. Delegating a task is useful. Surrendering authority that cannot readily be recovered is something else. As AI becomes more autonomous, reversibility should become a design principle.
The extinction of humanity remains a scenario, not a forecast. Treating it as inevitable would be intellectually careless. Treating it as impossible would be equally difficult to justify given the uncertainty surrounding systems whose capabilities are advancing faster than our ability to fully understand their behavior.
The more immediate danger may lie in a subtler assumption: that somebody else is responsible. The engineers will solve it. The CEOs will slow down. The regulators will intervene. The next safety framework will work. The next model will be aligned. The next international summit will produce agreement. Perhaps.
Human agency begins before certainty. It is the capacity to recognize that our choices still shape the trajectory while there is still a trajectory to shape. AI may become one of humanity’s most consequential inventions. Whether it ultimately enlarges the space for human flourishing or progressively narrows our control over the systems surrounding us will depend on decisions made in laboratories and boardrooms, classrooms and parliaments, investment committees and households. That is a formidable responsibility. It is also an extraordinary opportunity. Let’s face it — the future of artificial intelligence is being built now. The future of human agency must be built with equal intent, from the inside out.