Editorial illustration for AI Models Struggle Against Multi-Turn Attacks, Qwen3-32B Hits 86.18% Success Rate
AI Models Crumble Under Multi-Turn Cyber Attacks
AI models stop 87% of attacks but only 8% of attempts; Qwen3-32B hits 86.18%
Those headline defense numbers are a lie. Or at least, they describe a reality that doesn't exist. AI models block 87% of one-off attacks.
This is true. It is also useless. The moment an attacker tries a second time, the same defenses collapse.
Only 8% of these persistent attempts are stopped. The wall doesn't crack. It evaporates.
This is the new rule for breaking AI. You don't need a smarter prompt. You just need to keep talking.
Multi-turn attacks, where context is built across a conversation, succeed 64% of the time on average. For some models, it's a near certainty. Alibaba's Qwen3-32B fails 86.18% of the time under this pressure.
Mistral Large-2 fails 92.78% of the time. What looks like security is performance art, a convincing facade that falls apart if you push on it twice.
When attackers send a single malicious request, open-weight AI models hold the line well, blocking attacks 87% of the time (on average). But when those same attackers send multiple prompts across a conversation via probing, reframing and escalating across numerous exchanges, the math inverts fast. Attack success rates climb from 13% to 92%.
The researchers have it right. The problem isn't strength. It's memory.
Models cannot hold a defensive context. They forget they are under attack as soon as the chat bubble refreshes. An attacker's second message isn't met with renewed vigilance.
It's met with amnesia.
This turns every chatbot into a liability. Every customer service interface, every coding assistant, every creative co-pilot. Their safety is a snapshot, not a state.
Evaluating them on single prompts is like testing a bank vault's door while ignoring the plywood back wall. The numbers we celebrate are the wrong ones. The 8% failure rate for attempts is the only metric that matters.
It tells you the models are wide open. Anyone with a little patience can walk right through.
Common Questions Answered
How do multi-turn attacks differ from single-turn attacks on AI models?
Multi-turn attacks leverage persistent conversational strategies that dramatically increase the success rate of breaching AI defenses. While single-turn attacks might have a lower success rate, multi-turn approaches can escalate attack success rates by 5-10 times, exposing critical vulnerabilities in AI systems' contextual defense mechanisms.
Which AI models demonstrated the highest vulnerability to multi-turn attacks?
The research highlighted Alibaba Qwen3-32B and Mistral Large-2 as particularly susceptible models, with attack success rates of 86.18% and 92.78% respectively. These models showed a significant increase in vulnerability compared to their performance against single-turn attacks, with success rates jumping by up to 22%.
What makes multi-turn conversational attacks so effective against AI systems?
Multi-turn attacks exploit AI models' inability to maintain consistent contextual defenses across extended interactions. By persistently probing and manipulating the model through multiple conversational turns, attackers can gradually break down the AI's initial security barriers and increase their chances of successful breaches.
Further Reading
- Qwen3-32B Achieves 86.18% Performance on MMLU-Pro Benchmark — arXiv
- Qwen3 32B: Competitive Performance Analysis with GPT-4.1 and Claude Sonnet — Skywork AI
- Qwen3 32B Released April 29, 2025: Benchmarks and Performance Metrics — LLM Stats
- Qwen3 Benchmarks, Comparisons, Model Specifications and Performance Analysis — Dev.to