Editorial illustration for Multi-turn attacks break AI models 88% of the time, Cisco warns
Multi-turn Attacks Break AI Models 88% of Time
Cisco's security researchers spent months testing 15 flagship AI models with a simple question: what happens when an attacker doesn't give up after one try? The answer, presented by Amy Chang, Cisco's head of AI threat intelligence and security research, at VB Transform 2026, should unsettle anyone who thinks their red-teaming program is solid. Across 6,986 multi-turn attacks, adversaries who adjusted their approach turn by turn broke through 88.3% of the time.
That's not a marginal failure rate. It's a near-total collapse of defenses that likely looked fine under single-turn testing.
The timing matters. VentureBeat's June 2026 Pulse survey of 107 enterprise respondents found 54% have already dealt with a confirmed agent security incident or a near-miss. Only 32% give each agent its own scoped identity.
Just 30% sandbox their riskiest agents. Most companies, 82%, are still leaning on provider-native and hyperscaler controls as their main line of defense. Chang, who spent nearly two decades moving between JPMorgan Chase, Capitol Hill, and the Navy Reserve, brought a specific warning to that room.
When Cisco ran 6,986 multi-turn attacks against 15 flagship models, attackers who adapted across the conversation broke through as often as 88.3% of the time. Amy Chang, Cisco's head of AI threat intelligence and security research, brought that finding to the agentic security panel at VB Transform 2026; the number should worry anyone still running single-turn red-teaming programs.
Why this matters
If your red-teaming budget went entirely to single-turn prompt tests, Cisco just told you that budget missed the failure mode that matters most. An 88.3% break rate across nearly 7,000 multi-turn attacks against 15 flagship models is not a rounding error, it's a sign that adversarial conversation, not clever one-shot prompts, is how these systems actually fail. For developers building agents on top of these models, this should reframe what "safety testing" even means: an attacker who can iterate across turns is playing a different game than the one most eval suites are designed to catch.
Chang's point about accounting for failure points isn't abstract. It's the difference between knowing your model's weaknesses and finding out about them in production, from a user who wasn't trying very hard. With over half of VentureBeat's 107 enterprise respondents apparently sharing this concern, the pressure to build multi-turn evaluation into standard practice is going to come from customers and boards, not just researchers.
Anyone shipping agentic products without conversational red-teaming is now on notice, this data point won't stay obscure for long.
Common Questions Answered
What was Cisco's key finding about multi-turn attacks against AI models?
Cisco's security researchers found that attackers who adapted their approach across multiple conversation turns successfully broke through 15 flagship AI models 88.3% of the time across 6,986 multi-turn attacks. This significantly higher success rate compared to single-turn attacks demonstrates that adversarial conversation, rather than one-shot prompts, is the primary failure mode for these systems.
Why are single-turn red-teaming programs insufficient according to Cisco's research?
Cisco's findings reveal that focusing red-teaming budgets entirely on single-turn prompt tests misses the most critical vulnerability in AI models. The 88.3% break rate in multi-turn scenarios shows that attackers who adjust their strategy turn by turn exploit weaknesses that static, one-shot testing cannot identify.
Who presented Cisco's multi-turn attack research and where was it unveiled?
Amy Chang, Cisco's head of AI threat intelligence and security research, presented these findings at VB Transform 2026 during an agentic security panel. The presentation highlighted that the 88.3% success rate across nearly 7,000 multi-turn attacks against 15 flagship models represents a critical security concern, not a marginal failure rate.
How should developers building AI agents reframe their safety testing approach based on this research?
Developers should recognize that effective safety testing must prioritize multi-turn adversarial conversations rather than relying solely on clever one-shot prompts. Cisco's research indicates that the adaptive attack vectors demonstrated in multi-turn scenarios represent the actual failure modes that matter most for deployed AI agents.
Further Reading
- Leading AI models are more vulnerable to malicious multi-turn prompts, Cisco finds - Cybersecurity Dive
- Proprietary Problems: No Frontier Model Is Multi-Turn Immune - Cisco Blogs
- Frontier AI models collapse under multi-turn AI attacks - Help Net Security
- AI Safety Benchmarks Miss Real Attack Risk, Cisco Finds - AI Risk Today
- AI models block 87% of single attacks, but just 8% when attackers persist - VentureBeat