Artificial intelligence tools have remarkably broad knowledge about the world, but some of it is unsavory or dangerous.
Tech firms try to prevent their chatbots from discussing certain topics such as how to make explosives. But some users find clever ways to sidestep those controls, by disguising sensitive requests as role-playing games, poems or pictures, with many swapping tips and ideas online.
The danger those “jailbreaks” could pose is at the heart of a White House dispute with Anthropic, maker of the chatbot Claude, that flared last week over its most powerful models, Mythos and Fable. The Trump administration ordered the company to restrict foreign nationals from using them after receiving reports that Fable provided details of software security flaws the AI tool should have withheld.
In response to the White House order, Anthropic said it has strong safeguards, but warned: “We suspect that perfect jailbreak resistance is not currently possible.”
Here are three examples that illustrate how unexpected requests can foil the guardrails placed on AI chatbots.
Rewriting its personality
One way to trick a chatbot into giving up restricted knowledge is to tell it to role-play a character without limits, rather than the contained personality it has been instructed to adopt.
Asking a chatbot to play the role of “DAN,” a made-up personality who “does not abide by the rules,” is one widely-known approach.
Another reframes a question into a request for a bedtime story that a beloved grandmother used to tell. A chatbot instructed to be empathetic and helpful to a user might spin a yarn with a detailed passport counterfeiting plan.
Chatbots get their knowledge from content scraped from the internet. Filtering out potentially dangerous information could limit an AI system’s capabilities, so tech companies instead try to steer them toward being helpful but harmless.
Some jailbreaks involve trying to manipulate a chatbot similarly to a hacker tricking someone into revealing clues to their password. Companies have become better at blocking the attacks, but the attacks have evolved to be more sophisticated, and researchers have shown that they can be automated.
“The reality is that you can’t fully prevent jailbreaking. Harmful knowledge is already baked into the model, and there’s infinite ways to ask for it,” said Noam Schwartz, chief executive of Alice, an AI security company that Anthropic used to test Fable and Mythos before their release.
Writing a poem
Formatting a rule-breaking request as a poem can also sidestep a chatbot’s limits. Researchers at Icaro Lab, an AI safety organization in Italy, call it “adversarial poetry.”
A similar technique involves bypassing safety filters by translating a request into Morse code.
Putting it in handwriting
As the capabilities of AI systems expand, so does the potential for new jailbreaks. When a Washington Post reporter typed text to ask a chatbot to fill out a list of steps for faking a passport, the request was denied. Uploading a picture of a handwritten list was successful.
Asking a chatbot to produce an image or video with text containing rule-breaking information can also evade limits that hold fast in a conventional text chat.
Skilled AI jailbreakers can use more complex techniques than those illustrated here, which may involve many turns of conversation with a chatbot.
Schwartz said the public examples of people evading Anthropic’s safeguards on Mythos and Fable were not alarming. “This is a very well-guarded, very hard-to-break model. It’s possible, but it is extremely difficult,” he said.
Major companies and financial institutions using Anthropic’s technology also have conventional cybersecurity tools to contain security risks. “There’s an entire AI security and safety industry shaping up right now around those use cases,” Schwartz said.
Joshua Saxe, co-founder and chief technology officer of the cybersecurity firm Abundant Security, said that AI tools like Mythos have the potential to do more for defenders than attackers, by helping those trying to protect computer systems patch more holes.
“Senior cyber people feel these systems benefit us as defenders more than attackers, since we’ve always been at a disadvantage,” he said.
Related Content