Skip to main content
Researchers update classifier evasion techniques for vision-language models, showing adversarial attacks on AI. [arxiv.org](h

Editorial illustration for Researchers Update Classifier Evasion Techniques for Vision-Language Models

AI Models Vulnerable to Stealthy Image Attacks

Researchers Update Classifier Evasion Techniques for Vision-Language Models

Updated: 3 min read

In 2014, a few altered pixels could make a computer see an ostrich instead of a bus. Today, the attacks target systems that both see and describe the world. A team from the University of California, Berkeley and NVIDIA has detailed new ways to fool these vision-language models, moving past simple image tricks to exploit the link between pictures and words.

Evading image classifiers In 2014, researchers discovered that human-imperceptible pixel perturbations could be used to control the output of image classification models. Figure 2, from the seminal paper Intriguing properties of neural networks, shows how the images on the left (all distinctly and correctly classified) could be perturbed by the pixel values in the middle column (magnified for illustration) to generate the images on the right, all of which are classified as ostriches. As the field of adversarial machine learning evolved, researchers developed increasingly sophisticated attack algorithms and open source tools.

The work shows these models remain brittle. The researchers’ techniques, which include manipulating both image and text inputs, often succeed with changes a person would not notice. The findings, published this month, add to evidence that securing AI will require new approaches as systems grow more complex.

The old goal was to stop a machine from mislabeling a picture. Now the challenge is to stop it from constructing a false story about what it sees.

Common Questions Answered

How do transferable adversarial attacks work on Vision Large Language Models (VLLMs)?

Researchers discovered that attackers can craft specific image perturbations that induce targeted misinterpretations across multiple proprietary VLLMs like GPT-4o, Claude, and Gemini. These universal perturbations can consistently manipulate model interpretations, such as making hazardous content appear safe or generating incorrect responses aligned with the attacker's intent.

What types of attacks did the researchers demonstrate on vision-language models?

The study revealed four primary attack types: forcing VLMs to generate outputs of the adversary's choice, leaking information from their context window, overriding safety training, and making models believe false statements. Experiments on LLaVA, a state-of-the-art vision-language model, showed that all attack types achieved a success rate of over 80%.

Why are transferable adversarial attacks a significant concern for Vision Large Language Models?

These attacks expose critical vulnerabilities in current vision-language models, demonstrating that attackers can consistently manipulate model interpretations across different proprietary systems. The research underscores an urgent need for robust mitigations to ensure the safe and secure deployment of VLLMs, as these models become increasingly integrated into various applications.

LIVE00:32Visa Open-Sources Mythos Tool After Testing AI on Its Own Payment Network