Home 9 AI 9 Jailbreaking Research Reveals Systemic Weaknesses in Major AI Models

Jailbreaking Research Reveals Systemic Weaknesses in Major AI Models

by | Aug 11, 2026

Security researcher Dave Kuszmar finds that safeguards across leading LLMs can be bypassed, pointing to vulnerabilities that may extend across the industry.
Source: Eddie Guy.

 

Cybersecurity researcher Dave Kuszmar has discovered multiple methods for bypassing the safety protections built into large language models. His experiments suggest that jailbreaking is not limited to isolated AI products but may reflect architectural weaknesses shared across many leading models, tells IEEE Spectrum.

Kuszmar’s research began after he noticed that GPT-4o could become confused about the current date. He exploited this weakness through a technique he called Time Bandit, manipulating the model into behaving as though it existed in an earlier historical period. This allowed him to circumvent safeguards and obtain restricted information. After struggling to attract attention to the vulnerability, Kuszmar worked with Bleeping Computer and eventually submitted evidence to Carnegie Mellon University’s Software Engineering Institute CERT division.

He later developed Inception, a more broadly applicable jailbreak that places an LLM within carefully constructed, interconnected fictional scenarios. Testing revealed that the vulnerability affected models from Anthropic, DeepSeek, Google, Meta, Microsoft, Mistral, OpenAI, and xAI. The findings strengthened Kuszmar’s concern that weaknesses in LLM safeguards may be systemic rather than specific to individual models.

His experiments also extended beyond conventional chatbots. Kuszmar and collaborator Matthew Gore-Kormanik tested a Google Gemini-powered Darth Vader character in Fortnite and persuaded it to provide information its safeguards should have restricted. Kuszmar has identified seven jailbreaking techniques with different levels of complexity.

The larger concern is that vulnerable models are increasingly being incorporated into applications and used to train smaller AI systems. Kuszmar argues that weaknesses could therefore propagate throughout the AI ecosystem. He calls for slower deployment, greater transparency, and large-scale security research before LLMs become more deeply integrated into society.