Quantum computers may still be years away from breaking today’s encryption, but a new study shows that the cryptographic algorithms designed to survive that quantum threat can already be undone by something far more mundane: the electricity a chip consumes while doing its math. Researchers at Tsinghua University have demonstrated a profiling-based side-channel attack that recovers the secret decryption key of ML-KEM, the post-quantum encryption standard recently finalized by the US National Institute of Standards and Technology, from an unprotected hardware accelerator. Their attack, published in the journal Cybersecurity, needed an average of only about 27 power measurements to reconstruct an entire secret polynomial, a core component of the decryption key.
ML-KEM, standardized in NIST’s FIPS 203 and derived from the CRYSTALS-Kyber scheme, is the flagship lattice-based key-encapsulation mechanism intended to replace vulnerable key-exchange methods in the quantum era. Its mathematical security rests on the hardness of the Module Learning With Errors problem, a lattice problem believed to resist attacks even from large-scale quantum computers. But mathematical security is only as strong as the hardware that implements it. Side-channel analysis exploits physical observables, such as power consumption, electromagnetic emissions, or timing variations, to extract secret information without ever attacking the underlying mathematics directly.
The Tsinghua team, led by Munkhbaatar Chinbat with Liji Wu, Xiangmin Zhang, and Munkhnar Khulan, targeted the Number Theoretic Transform, or NTT, the arithmetic engine at the heart of ML-KEM. The NTT is a fast polynomial multiplication technique that converts secret coefficients into a transformed domain where multiplication becomes efficient. During decryption key generation, the secret vector is sampled from a centered binomial distribution and immediately transformed into the NTT domain before being stored. That transformation, the researchers realized, is a rich source of exploitable leakage because it processes every secret coefficient through well-defined, repetitive butterfly operations.
The structure of the first NTT round is what makes the attack possible. The transform operates on 256 coefficients arranged into 128 pairs, and in the first round every pair is processed with the same public twiddle factor. This means the power signature of each butterfly operation depends almost entirely on the secret input values, not on public data that could mask the leakage. Because the team’s accelerator processes coefficient pairs sequentially, the leakage from each pair occupies a distinct, predictable window in time, allowing an attacker to cleanly separate the fingerprints of all 128 pairs from a single recording.
To build their experimental platform, the researchers implemented a deliberately naive NTT accelerator on a Xilinx Artix-7 FPGA hosted on a CW305 board, using a ChipWhisperer-Lite for power trace acquisition. The design uses a single butterfly unit, a finite state machine controller, shared memory, and standard Montgomery and Barrett reduction techniques for modular arithmetic, mirroring the reference implementation of Kyber. The accelerator runs at 8.5 megahertz while the ADC samples power at roughly 105 million samples per second, with each trace capturing 20,000 sample points covering the entire first NTT stage. The minimalist design required only 743 lookup tables, 403 registers, one DSP block, and one block RAM, trading speed for simplicity and analyzability.
The attack methodology unfolds in two phases. In the profiling phase, the attacker trains on a clone device with known inputs. The team collected 100,000 power traces, using 80,000 for building statistical templates and 20,000 for the attack itself. Before any modeling, they applied the Test Vector Leakage Assessment, a statistical t-test comparing fixed versus random inputs, to pinpoint the exact time samples where secret-dependent information leaks. Any sample where the t-statistic exceeded 4.5 in absolute value, corresponding to a confidence level of 99.999 percent, was flagged as a point of interest. For each of the 128 coefficient pairs, these points were then compressed using principal component analysis, reducing thousands of raw samples to a handful of informative features.
Classification then proceeds via Mahalanobis distance, a statistical measure that accounts for the covariance structure of the leakage data. For every possible coefficient value pair, the profiling phase stores a mean vector and covariance matrix describing what its power signature typically looks like. During the attack, an unknown trace is compared against all templates, and the closest match reveals the secret coefficients. The results were striking: classification accuracy ranged from 96.93 to 97.97 percent across all 128 coefficient pairs, with an average near 97.4 percent. Notably, this classical statistical approach outperformed the team’s earlier deep-learning experiments with multilayer perceptrons, convolutional networks, and recurrent networks, which topped out at 96.64 percent accuracy.
Single traces, however, were not always sufficient. The team’s overlap analysis showed that only a small fraction of individual traces classified all 128 pairs correctly in one shot, since certain coefficient combinations, such as symmetric pairs like (1, -1) and (-1, 1), produce nearly identical power signatures. The solution is likelihood aggregation: by accumulating classification evidence across multiple decapsulation operations, the correct values emerge statistically. The guessing entropy analysis, which measures the expected rank of the correct answer among all candidates, dropped from 11.15 with 150 traces to just 1.09 with 10,000 traces, confirming that the templates converge rapidly toward certainty. Across 734 successful recovery trials, the number of traces needed for full polynomial recovery ranged from 7 to 92, with a mean of 27.3 and a median of 23.
The findings place ML-KEM hardware in a sobering context. Previous research had already demonstrated that software implementations on Cortex-M4 processors leak enough information for full key extraction with as few as 8 to 960 traces, and that masked and shuffled implementations can be broken with single traces when fault injection is combined with side-channel analysis. The new work extends this vulnerability picture to lightweight FPGA accelerators, showing that even a minimal, sequential NTT design without any countermeasures is effectively defenseless against a well-resourced profiling attacker. Because ML-KEM’s secret key consists of multiple polynomials, three for ML-KEM-768 and four for ML-KEM-1024, full key recovery simply requires repeating the polynomial-level attack for each component.
The researchers emphasize that their attack targets an unprotected design, and that more sophisticated architectures using parallelism or pipelining may present a smaller attack surface by blurring the temporal correlation of leakage. Still, the message for the post-quantum transition is clear: theoretical quantum resistance means nothing if the hardware leaks. The team proposes countermeasures including physically unclonable function-based address shuffling and higher-order masking, and points toward hybrid attacks that would feed template-attack likelihoods into belief-propagation frameworks such as Soft Analytical Side-Channel Attacks to reduce the trace requirements even further. As ML-KEM deployment accelerates across browsers, banks, and national infrastructure, the study serves as a warning that securing the mathematics is only half the battle; the silicon running it must be hardened against the whispers of its own power supply.
Subject of Research: A template-based side-channel attack recovering ML-KEM decryption keys from power leakage of an FPGA NTT accelerator
Article Title: A Template attack for decryption key recovery on the NTT accelerator of PQC ML-KEM
Article References: Chinbat, M., Wu, L., Zhang, X., & Khulan, M. (2026). A Template attack for decryption key recovery on the NTT accelerator of PQC ML-KEM. Cybersecurity, 9(1), Article 226. https://doi.org/10.1186/s42400-026-00675-3
Image Credits: AI Generated
DOI: 10.1186/s42400-026-00675-3
Keywords: ML-KEM, post-quantum cryptography, side-channel attack, template attack, Number Theoretic Transform, FPGA, power analysis, TVLA, principal component analysis, CRYSTALS-Kyber, decryption key recovery, NIST FIPS 203
