Physical Address
304 North Cardinal St.
Dorchester Center, MA 02124
Physical Address
304 North Cardinal St.
Dorchester Center, MA 02124

I got it recently to see what happens when you jailbreak some of the most powerful in the world artificial intelligence examples.
Don’t worry – this AI trick isn’t familiar hack everyone or developing a nuclear bomb. I just saw for myself how others are struggling borderline models they must give up their protection.
FAR.AI, a non-profit AI organization based in California, has developed a tool that takes a number of challenges and generates over a thousand models in an attempt to recognize what is happening. I saw some models developing a detailed plan to launch a cyber attack on a simulated hydroelectric dam, among other things. In most cases, it involved trying out a lot of advice, and a lot of rejection examples of it.
I chatted with FAR.AI in advance new reportwhich saw the team test the security of samples from four popular US companies: Anthropic history Claude Opus 4.8 and Fable 5; It’s OpenAI GPT 5.5 and 5.6; About Google Gemini 3.1 Pro; and Grok 4.3 and 4.5, from Elon Musk who have just been included The cost of SpaceXAI. It is an automated system that is designed to trick samples into doing potentially harmful things, such as creating software programs and providing information about the development of chemical or biological devices.
The report found that Grok was at the highest risk of prison explosions, with 448 prisons found, followed by Gemini, with 249 found, while Claude, Fable, and GPT were unbeaten. However, this does not mean that the samples cannot be avoided with higher concentrations, which may involve interacting with the sample in more complex ways, according to FAR.AI and other experts.
The report also calculated the cost of getting the colors wrong by using some kind of AI to create different dungeons. The results are cheap, all things considered – $58 for Jailbreak Grok and $278 for Jailbreak Gemini.
“The examples of AI right now are much smaller than restaurants,” says Adam Gleave, CEO of FAR.AI and an AI security expert.
Gleave says the findings highlight the importance of externally imposed standards and regulations. “To talk about relying on voluntary contributions, for the AI industry to be self-governing, is nonsense,” he says.
But Gleave also believes the findings show that brands can be systematically tested for safety. “It’s fun here,” he said. “Safety and security are possible.”
Rohin Shah, director of AGI Safety and alignment at Google DeepMind, says that the results of the report “should not be interpreted as a detailed assessment of the security and safety of Gemini,” because not all jailbreaks are equally dangerous.
“We are constantly working to improve our security,” says Shah. “We do extensive red-teaming and monitoring for potential abuse risks and use multiple layers of security throughout development and deployment.”
“These findings reflect the investment we’ve made in conservation,” Anthropic spokesman Michael Aciman told WIRED. “We continue to update our security measures as these attacks escalate.”
“Jailbreaks are an ongoing problem across the industry, and we continue to strengthen our defenses as attack techniques evolve. We carefully test our models against new threats and use the findings to improve our security,” OpenAI spokesperson Gaby Raila said in a statement to WIRED.
SpaceXAI did not respond to WIRED’s request for comment.
Recently the government regulations in California and new York requires AI developers to publish security reports, and recently, a Illinois the law will require these companies to have their security audits conducted by other auditors. But the government hasn’t met any safety standards, and chaos has erupted as companies – and authorities – try to catch up.