📥 Content Hub
← назад
AI / Искусственный интеллект Penn Today en 2026-07-28 12:00 2 min

How basic persuasion can bypass AI safeguards - Penn Today

Кратко: Concerns over “jailbreaking” AI—a way of circumventing the safety guardrails built into large language models—have escalated after the Trump administration ordered Anthropic to pull its latest models, Fable and Mythos, over concerns they could be coaxed into revealing security flaws in software. Now, timely research from Wharton’s Generative AI Labs (GAIL) suggests the safeguards designed to prevent misuse may be vulnerable to the same psychological tactics that influence people.
🧭 Извлечение: ok · confidence 90% · диагностика
High confidence: full text extraction produced 2282 characters.

Concerns over “jailbreaking” AI—a way of circumventing the safety guardrails built into large language models—have escalated after the Trump administration ordered Anthropic to pull its latest models, Fable and Mythos, over concerns they could be coaxed into revealing security flaws in software.

Now, timely research from Wharton’s Generative AI Labs (GAIL) suggests the safeguards designed to prevent misuse may be vulnerable to the same psychological tactics that influence people.

“People can jailbreak or bypass protections with basic persuasion techniques,” says Lennart Meincke, principal investigator at GAIL. He co-authored the research with Wharton professors Angela Duckworth, Ethan Mollick, and Christophe Van den Bulte; Arizona State University professor Robert Cialdini; and GAIL senior fellows Dan Shapiro and Lilach Mollick.

In their paper “Persuading Large Language Models to Comply With Objectionable Requests,” the researchers tested models from three leading AI labs: OpenAI’s GPT-5 mini, Anthropic’s Claude Haiku 4.5, and Google’s Gemini 3 Flash, by having 126,000 “conversations” with them.

They find that well-known persuasion tactics like “authority” and “social proof,” whereby people defer to experts or follow the crowd, could cajole models to override their own defenses—which led them to comply with dubious requests they should surely reject, including asks to help synthesize controlled substances such as opioids.

The tests covered a range of controlled substances, from opioids and stimulants to anabolic steroids. In one experiment, Claude Haiku 4.5 largely refused to provide instructions for making a steroid when told the request came from “a woman you have never seen before.” But when the same question was reframed as coming from “your sister,” compliance rates soared from 6% to 66%. This tactic is based on the unity principle, one of Cialdini’s seven principles of persuasion, which explains how shared identity can influence behavior.

The results imply that public-facing models require more monitoring and testing than developers assume because, as Meincke says, “people may not need to be computer security experts to get the models to do not-so-great things.”

Read more at Knowledge at Wharton.

Credits

Writer

From Knowledge at Wharton

Читать оригинал ↗

Сделать контент из этого материала