hidden-switches:-new-attack-plants-undetectable-backdoors-in-vision-transformers
Hidden Switches: New Attack Plants Undetectable Backdoors in Vision Transformers

Hidden Switches: New Attack Plants Undetectable Backdoors in Vision Transformers

Vision Transformers have quietly become the backbone of modern computer vision, powering everything from image classifiers and object detectors to medical imaging pipelines and autonomous driving prototypes. Their self-attention mechanism, which lets the model weigh relationships between every patch of an image, has displaced convolutional networks in many high-stakes applications. But a new study from researchers at the University of Information Technology, Ho Chi Minh City, and Vietnam National University Ho Chi Minh City suggests that the very architecture making these models so powerful may also make them dangerously easy to compromise. In a paper published in the International Journal of Machine Learning and Cybernetics, Dung Minh Do and Khang Nguyen describe a backdoor attack that achieves attack success rates of up to 99 percent while leaving the model’s performance on clean, unmodified images essentially untouched.

Backdoor attacks are among the most insidious threats in machine learning security. Unlike adversarial examples, which exploit a model at inference time by perturbing individual inputs, a backdoor is baked into the model itself during training or fine-tuning. A model with a backdoor behaves perfectly normally on ordinary data, matching the accuracy and reliability its developers expect. But when a specific trigger, a subtle pattern chosen by the attacker, appears in an input, the model’s behavior flips, producing whatever output the attacker has designated. For a deployed system, this means an adversary could, for example, cause a traffic sign classifier to misread a stop sign or a security screening system to wave through a prohibited item, all without any visible degradation in everyday performance that might raise alarms.

What makes the new attack notable is where it operates. Rather than modifying the weights of the transformer or poisoning the training labels in conventional ways, the method injects two specially designed tokens into the model’s token stream, progressively working from the shallowest layers to the deepest ones. Vision Transformers process images by splitting them into patches, embedding each patch as a token, and then passing those tokens through a stack of transformer blocks in which self-attention layers let tokens exchange information. By inserting crafted tokens at multiple depths, the attack ensures that the malicious signal is reinforced and refined as it travels through the network, rather than being diluted or overwritten by the model’s normal processing.

The progressive, layer-by-layer nature of the injection is central to the attack’s stealthiness. Defenses that inspect a single layer, or that look for anomalous attention patterns at one point in the network, can miss a signal that is distributed across the entire depth of the model. Because the injected tokens interact with the image tokens through the standard attention mechanism, they can steer the model’s internal representations toward the attacker’s chosen target class only when the trigger is present. On clean inputs, the tokens remain effectively inert, which is why the authors report that the method maintains competitive clean accuracy even as it delivers near-perfect attack success rates across multiple visual datasets.

The word switchable in the study’s title points to another dimension of the threat. The attack draws on a growing line of research into switchable backdoors against pre-trained vision transformers, in which a single compromised model can carry multiple behaviors that an attacker can toggle. This means a defender who discovers one trigger and filters it out cannot assume the model is safe; another trigger may remain dormant, waiting to be activated. The authors position their work as an analysis and exploitation of ViT vulnerabilities, and the breadth of prior work they survey, from BadNets and weight-poisoning attacks to attention hijacking and prompt-based backdoors, underscores how rapidly this attack surface has expanded as transformers have spread through computer vision.

The connection to prompt tuning is particularly significant for the current state of the field. Prompt-based methods, in which small learnable tokens are prepended to a model’s input or inserted into its layers, have become a popular way to adapt large pre-trained models to new tasks without expensive full fine-tuning. Visual prompt tuning and related techniques are widely used because they are efficient and effective. But the same mechanism that makes prompts useful, namely the ability to inject learned tokens that influence the model’s attention and representations, is exactly what this attack weaponizes. A malicious actor with access to a fine-tuning pipeline, or able to distribute a poisoned adapter or checkpoint, could embed a backdoor that looks indistinguishable from a legitimate prompt-based adaptation.

The supply chain implications are sobering. Modern machine learning practice relies heavily on pre-trained models downloaded from public hubs, fine-tuned adapters shared between teams, and third-party datasets. Each of these channels is a potential delivery mechanism for a backdoor. Earlier research has shown that weight poisoning attacks on pre-trained models can survive downstream fine-tuning, and that backdoors can be hidden in ways that evade standard inspection. The new study adds to this picture by showing that the token-based machinery of vision transformers offers attackers a particularly clean injection point, one that does not require the crude modifications of earlier attacks that made them easier to detect.

The authors evaluated their method on multiple visual datasets, measuring not only attack success rate and clean accuracy but also stealthiness and robustness. The headline figure, an attack success rate of up to 99 percent, is alarming enough, but the more troubling result is the combination of that success with preserved clean performance, since it means conventional accuracy-based validation would reveal nothing amiss. The paper also situates the work against existing defenses, which include fine-pruning approaches that remove rarely activated neurons, neural attention distillation that tries to erase trigger-related attention patterns, and input-level detection methods that look for inconsistencies in a model’s predictions under image transformations. Because the new attack distributes its signal across shallow and deep layers through dual tokens, many of these defenses, which were designed with convolutional networks or single-point injections in mind, face a harder problem.

The researchers are explicit about the warning their findings carry for real-world deployment. Vision Transformers are increasingly used in settings where a silent failure mode could have serious consequences, including medical diagnosis support, surveillance, industrial inspection, and safety-critical perception systems. A backdoored model in any of these contexts could be remotely triggered by an input crafted to contain the attacker’s pattern, and the compromise would be invisible in routine testing. The authors make their code publicly available, which serves the defensive side of the field as well: reproducible attack implementations are essential for developing and benchmarking countermeasures, and the history of adversarial machine learning shows that security research advances fastest when attacks are fully documented.

For the broader community, the study is a reminder that architectural progress and security progress have been badly out of step. The references in the paper trace a decade of deep learning breakthroughs, from early convolutional networks through EfficientNet and the original Vision Transformer, alongside a parallel literature of backdoor learning surveys, prompt injection analyses, and defense proposals. Yet the authors note that security threats to ViTs, particularly backdoor attacks, have not received research attention commensurate with the architecture’s adoption. Closing that gap will likely require defenses designed specifically for token-based architectures: methods that audit injected tokens, verify the provenance of fine-tuned checkpoints, test models against a family of triggers rather than a single known pattern, and treat the entire depth of the network, not just its input layer, as a potential attack surface. Until such defenses mature, the near-perfect stealth and effectiveness demonstrated by progressive dual-token injection stands as a stark warning that the models powering tomorrow’s vision systems may harbor switches that only their attackers know how to flip.

Subject of Research: Backdoor attack vulnerabilities in Vision Transformers via progressive dual-token injection

Article Title: Switchable backdoor attack in vision transformers via progressive dual-token injection from shallow to deep layers

Article References: Do, D. M., & Nguyen, K. (2026). Switchable backdoor attack in vision transformers via progressive dual-token injection from shallow to deep layers. International Journal of Machine Learning and Cybernetics, 17(10), Article 489. https://doi.org/10.1007/s13042-026-03326-8

Image Credits: AI Generated

DOI: 10.1007/s13042-026-03326-8

Keywords: Vision Transformers, backdoor attack, machine learning security, self-attention, prompt tuning, adversarial machine learning, model poisoning, supply chain security, deep learning, cybersecurity, image classification, AI safety

Cite Scienmag News

APA
MLA
Chicago

Copy citation
Download RIS

Tags: adversarial attacks on computer vision modelsadversarial machine learningAI safetybackdoor attackbackdoor attack in machine learningbackdoor detection in vision transformerscovert backdoor triggers in deep learningcybersecuritycybersecurity threats in AIdeep learningdeep learning model robustnesshidden switches in neural networkshigh-stakes AI application vulnerabilitiesimage classificationmachine learning securitymedical imaging model securitymodel integrity in autonomous systemsmodel poisoningprompt tuningself-attentionself-attention mechanism exploitationsupply chain securityvision transformer security vulnerabilitiesVision Transformers