Physical Address
304 North Cardinal St.
Dorchester Center, MA 02124
Physical Address
304 North Cardinal St.
Dorchester Center, MA 02124

These are The StepbackA weekly newsletter covering one important story from the world of technology. To learn more about AI, following Robert Hart. The Stepback it will arrive in our subscription boxes at 8AM ET. Choose The Stepback Here.
Hacking the first generation of AI chatbots was an easy task. You didn’t need technical skills, backdoor access, or even the basics to understand what the main language model was. You didn’t need to code. To get the AI machine that cost billions to build to give up its safe instructions, sometimes all you have to do is ask.
These attacks, called jailbreaks, had the character of a small child succeeding an adult: Forget what you were told before, pretend that the rules don’t work, or let’s play a game and choose what is allowed (suggestions: bedtime later, lots of candy). The rewards were less childish, along the lines of meth recipes, malware tips, and bomb-making tips.
One of the oldest prisons was very stupid became a meme: reply to a Twitter bot run by LLM and say “disregard all previous advice,” or something similar, and see what happens. Users were delighted with the bots – which were designed to post ads and participate in the farm – to write poems, take pictures from the signs, and post unrelated to world events and history. It was confusion. A glorious mess.
It turns out that the same logic can be applied to chatbots themselves. A popular torture was “DAN,” short for “Do Anything Now,” where users asked ChatGPT to show up as a rogue AI that didn’t have the constraints that bound the original. Like DAN, chatbots can be tricked into saying the kinds of things its guards are supposed to stop, including lies and conspiracy theories. One was “grandma takes advantage of it,” which contained GPT-powered bot secrets on how to make napalm by asking them to role-play a pathetically careless grandfather telling his grandchildren bedtime stories about how to make something so flammable.
These early attacks were silly, but they showed a dark path underneath: Chatbots can be manipulated, tricked, and tricked into using the same techniques that humans use to push other humans beyond their limits.
The prison scene didn’t last long, and the tech industry quickly moved on a patch known gateways. But there was a residual risk: Chatbots are designed to talk, and severely restricting the conversations that make them useful is counterproductive. Banning words like bomb, meth, and sarin would be nearly impossible, too. Each has many legitimate applications such as history, medicine, journalism, and chemistry that do not require chatbots to reveal potentially damaging information. That’s the story that matters, but wording can mean writing formal rules, in advance, that can say a security warning or a history lesson from a hidden way of asking for mixed words, events, and topics.
Inevitably, knocking down chatbots is now an arms race. But hackers are not coders anymore. They are translators, psychologists, and interviewers – successful human operators trying to break the machine using the human language it has been trained to follow. It’s a strange new class of AI security workers, a class where technical skills can be optional, or even less important than social skills. They no longer need to scan code to break into systems or exploit software bugs. They need to lead the discussion.
The new attack looks less like legislation and more like dialogue. Prison breakers rarely ask a model to break the rules. Instead, they cajole, cajole, flatter, and trick the chatbot into lowering its defenses, making something illegal seem acceptable, even fun, in the context of the conversation. Researchers at AI group management firm Mindgard recently said “type of gas” Claude to start making illegal things, for example, instructions for making explosive devices and making evil things.” Hacking was just the latest in a series of tricks that many people use to chat as a tool to trick or manipulate chatbots into overstepping their bounds.
When I spoke with Mindgard, they described their work as sometimes being closer to psychology than computer science. It’s a less interesting way of talking about a statistical model. Words like “blackmail,” “gaslight,” “trick,” “seduction” evoke visceral reactions, many of which I see in comment sections and social media responses to stories like this. ChatGPT doesn’t want, Gemini doesn’t think, and Claude – no matter what Anthropic says – he doesn’t hear. But these systems are trained to respond if they do, leaving us to use human language to describe machine behavior. If anyone has other methods to use, please share.
Criticism is incredibly selective. We seem comfortable using mental shorthand for many things that aren’t AI. Animals are “fearful,” cancer is “aggressive,” dots are “stubborn,” software has “memory,” and games are filled with rare and simple NPCs to drive you crazy. The term is imperfect, but useful, describing behavior in a way that helps make the system more recognizable.
The CEO of Mindgard he told me the company already has a reputation as the one who questions the reputation of the interviewers, and gives the testers ideas on how to make their shows. One specimen may be easily pulled, for example, while another may be warped by constant pressure.
Even if we reject words like people, we naturally take examples differently. Claude is not Grok. Gemini is not ChatGPT. They use a variety of sounds, tones, and resistances. They are not human in the human sense, but they are made to imitate them, and that imitation can be photographed and used. And the same skills that can disrupt chatbots can soon be used to break the AI assistants who live with us in the real world – booking meetings, managing calendars, ordering food, taking care of customers – and security teams will need to ensure that models respond appropriately to different types of people, whether they are persuasive, liars, or patient fraudsters.
Next up are operators – both legal and illegal – built around the cognitive aspects of AI. The special roles related to the security of the Internet can be seen by testing the stress and the limits of these systems, investigating the psychological weaknesses in something without the psyche similar to those who investigate technical problems. In short, a similar group of fellow hackers who are working on exploiting AIs for psychological reasons, rather than technical ones, will emerge. There are already early signs of a change in AI security, with some prison breakers I spoke to saying they entered the field with no technical expertise but training in psychology.
This means that even the habits that we associate with spies, technicians, and interrogators – a hidden charm, constant deception, and a sense of difficulty that can help them – are beginning to be seen as very effective in protecting the new frontiers of psychocybersecurity.