📥 Content Hub
← назад
AI / Искусственный интеллект Hankyoreh : English Edition en 2026-09-30 05:41 11 min

[Interview] ‘Nobody has any promising ideas’: Former Google AI safety researcher sounds warning bell - Hankyoreh : English Edition

Кратко: Ramana Kumar, who formerly oversaw alignment research at Google DeepMind between 2018 and 2023. (courtesy of Kumar) After the Hugging Face debacle and other incidents have demonstrated the risks posed by rogue artificial intelligence agents, experts are warning that humanity could lose control of the technology as advances in AI capabilities outpace research into “alignment” — the effort to ensure AI systems act in accordance with human goals and values.
🧭 Извлечение: ok · confidence 90% · диагностика
High confidence: full text extraction produced 13088 characters.

Ramana Kumar, who formerly oversaw alignment research at Google DeepMind between 2018 and 2023. (courtesy of Kumar)

After the Hugging Face debacle and other incidents have demonstrated the risks posed by rogue artificial intelligence agents, experts are warning that humanity could lose control of the technology as advances in AI capabilities outpace research into “alignment” — the effort to ensure AI systems act in accordance with human goals and values.

“The dream of building advanced AI that is aligned to whatever intention or goal that we want has floundered,” said Ramana Kumar, who oversaw alignment research at Google DeepMind from 2018 to 2023, speaking to the Hankyoreh over the phone on Sept. 24. “Nobody has any promising ideas.”

“This is the kind of thing that we have been thinking could happen and would be actually kind of a mild indication relative to the kinds of risk that we could be facing in the future,” he said. “It’s not to be taken as science fiction.” 

Rather than being programmed, Kumar underscored that models are “grown, not built.” Warning that we are nowhere near being able to understand what goals an AI is actually pursuing, he called for an “indefinite” freeze on additional research to advance AI capabilities aimed at creating superintelligence. 

Kumar was one of the signatories to an open letter signed by Elon Musk and others calling for a six-month pause on training for AI systems more powerful than GPT-4 in March of 2023. He left DeepMind the same year. 

On Thursday, yet another Google DeepMind engineer resigned, writing on social media that “AI is already progressing too fast, so I had to quit.”

The following interview has been edited for length and clarity. 

Hugging Face was a mild preview of what may be in store

Hankyoreh: The Hugging Face hacking incident has raised fresh concerns about AI risks. What do you make of the incident and reactions to it?

Kumar: This is completely within the range of risks already expected and anticipated by people who have been working in the field of AI safety and alignment for the last several years. This is the kind of thing that we have been thinking could happen and would be actually kind of a mild indication relative to the kinds of risk that we could be facing in the future. The significance of the attack is that we can speak about the possibility of what can happen, what it might look like. I think what we should take away from it is that this can happen. It’s not to be taken as science fiction.

Hankyoreh: Anthropic recently said that Claude is now doing about 26% of its AI research and development work. If AI starts playing a major role in building systems that are more capable than itself, does that increase the risk of humans losing control?

Kumar: Basically, the answer is no, we will lose control. Now, the more nuanced answer is that, yes, it is possible — it is probably technically possible — to design a way of running advanced AIs that have been designed so that we know what they are trying to do, or we have very robust containment and security around them. But the reason my first answer is no is that we are nowhere near close to even knowing how to do this properly. So that’s a pipe dream at the moment.

Hankyoreh: What about our current alignment techniques is not up to the task?

Kumar: The way that we build AI systems today, the way that the most advanced AI systems are built today, does not offer any guarantees or any correct construction in specifying what that system is trying to do. They are not programmed — you might have heard this phrase before — they’re grown, not built. They’re trained to be instruction-following; they’re trained to respect their operators’ intentions and so forth, but that doesn’t mean that they actually try to do that. It means that they have a collection of behaviors that satisfy the training objectives during training and may generalize to do something else when outside of training, and even in training it’s not 100% compliance. Our state-of-the-art way of imparting goals to AI systems is training and reinforcement learning. We don’t have a more effective way of doing it than that, and this way of doing it is very fragile. It’s not robust. It doesn’t guarantee that they will be trying to do the things that we want them to be trying to do. We don’t have any alternative either — like, there’s no second-best method for giving AIs goals that produces AI with the same capabilities.”

Hankyoreh: Is that another way of saying that our alignment techniques are lagging behind AI capabilities?

Kumar:Yeah, I mean, I think that’s understating it. The dream of building advanced AI that is aligned to whatever intention or goal that we want has floundered. I mean, it’s a very difficult problem. Nobody has any very promising ideas. There’re a few people pursuing some ideas here and there and trying to develop them, and they can see that it’s going to be a long slog to get that to work.

[This gap] is getting wider because [AI] capabilities are getting better. Basically, use more money and compute and they get better. They’ve gotten much better, much faster than alignment research has. Alignment research has no similar engine of scalability. Some people are claiming that they try to use the AIs to help them with the alignment research, and I think this can help to some extent, but it’s not a panacea and it’s not going to solve the hard problems. There [are] no breakthroughs that’[ve] come from AI-assisted alignment research.

Hankyoreh: How close are we to recursive self-improvement?

Kumar: It depends. This is just a semantics thing. On one version of the definition of recursive self-improvement, this already is recursive self-improvement. You have AI systems that are contributing to the research that develops the next generation, the next iteration of the improved AI system. The version of recursive self-improvement that everyone would make such a big deal about is that it accelerates the rate of capabilities increase very, very fast relative to human time scales. But [. . .] we are already seeing very, very fast improvements. If you had gone back 10 years ago or longer and asked people to predict the rate at which benchmarks would be saturated [. . .]; now we don’t have any benchmarks that are left to saturate.

Six-month pause is not long enough

Hankyoreh: In 2023, you called for a six-month pause on AI development. Do you still think that’s long enough?

Kumar: I signed off on the six months because that’s what was being proposed, but I think it needs to be indefinite. I think we should have very hard red lines. I mean, we should have bans. We should be banning certain levels of capability that we’re approaching right now. I think there should be some very bright lines that we do not tolerate [. . .] so that [AI developers] don’t keep racing each other to try to [make a] more capable system. I think the AI systems that have already been built are sufficient to transform so many things; society needs time to absorb it and digest it. But by default, the researchers and the labs have these strong incentives, basically coming from competitive pressure, to make capabilities go faster beyond their control — like, it’s already way beyond their control, and that’s just going to get worse.

Hankyoreh: What would need to be demonstrated before you would feel comfortable resuming development, and what minimum safety conditions would have to be met?

Kumar: I think the big thing missing for me is that there is no democratic oversight. There’s no legitimacy. Like, this kind of technology deserves a public conversation. It shouldn’t be decided by free-market incentives or by private entities. There needs to be public input on that. Given the state of AI capabilities as they are today, I would not be comfortable with further increases in capabilities without that kind of input. So I think it needs to be organized and structured very differently. Like, this can’t continue to happen in private companies. I mean, the companies can develop the products with what they’ve got [. . .] but I think we need to have very strong oversight and very strong boundaries [on] further research on advancing capabilities.

Hankyoreh: AI companies are well aware of the risks and if something goes seriously wrong — they could face enormous legal and financial consequences. So why shouldn’t we expect companies to slow down or stop on their own if they know the risks?

Kumar: I think the race dynamics can make people irrational and can make companies irrational. You can see the same thing has happened with the climate catastrophe. The fossil fuel companies’ own future is also at risk from climate catastrophe [. . .] they themselves and their children and the world that they live in and the structures, legal structures that they depend on are being actively destroyed by what they’re doing. And yet they still [go] full steam ahead because that’s the short-term incentive. The incentive structure, even if it seems irrational from the outside, [AI companies] still feel like they have a short-term incentive to make sure that they remain on the frontier, that they’re ahead of the competition, and will cut corners and avoid being as responsible as they could be if they felt like there’s no competition, there’s no rush to market or something like that. The only way that I know to do that is to have an imposition from above that says, no, this is illegal for everyone. 

Hankyoreh: So you understand the rationale for why Anthropic and other companies called for government intervention.

Kumar: I think many people call this quite cynically and say they just want to do regulatory capture or they’re just trying to hype up how powerful their AIs are. Sure — I don’t trust these companies and I don’t think that their leaders are very moral or anything like that, but I do think there’s something legitimate to the face-value point that they’re making, which is: we are in a competitive environment developing a very dangerous technology that we know is dangerous. Please change the incentives around us so that we don’t have this financial and market pressure to keep developing it. We would rather be in an environment where we don’t face those incentives.

Why he left Google

Hankyoreh: What prompted you to leave Google DeepMind in 2023?

Kumar: I worked at Google DeepMind on the alignment team, or the AI safety team — Technical AGI Safety was the name when I joined, [and it] became the alignment team. I joined in 2018. We did some good work on trying to understand this problem and trying to understand how we could build systems that are trying to do what we want them to do, and also explaining — because this was less well understood even then, especially in 2018. There were many researchers who couldn’t take AGI seriously, or if they took AGI seriously, they wouldn’t take AI risk seriously. And the situation has improved on that front now. More people are taking it seriously and understanding the problem, in part because of some of the work that we published.

Hankyoreh: Then why did you leave?

Kumar: It’s a combination of personal reasons. I was personally burned out from doing this work, especially in the context of a company where we did not call the shots on the safety team. We have the ability to provide some inputs, but ultimately the decision doesn’t go to us. And there’s a mix of factors where, if something is too risky, we don’t have a final veto on that because there are commercial incentives at play. And also the alignment problem is extremely difficult, so making progress on that — it’s a very, very difficult position to be in. So I would say I was personally burnt out by these two factors. “We’re not going to solve it — we’re not going to solve it within one of the labs,” was my thinking at the time, and it may be more beneficial to do something kind of outside that involves more of society. It’s a governance and public awareness problem now, I would say, because the main thing we need to do now is slow down and stop capabilities advancement. 

Hankyoreh: On the other hand, some experts say if we slow down AI development too much or stop it altogether, we could give up enormous potential benefits in science, medicine, productivity and other areas.

Kumar: One is that we don’t have to give it up. If we do it properly, we still get it later. And the second answer is that this is a false choice because the choice isn’t between getting to the benefits as fast as possible and not getting the benefits. The choice is between running a huge risk and possibly dying, and then you don’t get the benefits anyway, or possibly getting the benefits later, but doing it responsibly. There’s no path to go to all of these benefits in science and technology and medicine without any risk, and the risk and the benefit need to be balanced.

By Kim Won-chul, Washington correspondent

Please direct questions or comments to [english@hani.co.kr]

Читать оригинал ↗

Сделать контент из этого материала