Every mass spectrometry experiment in modern proteomics ends with the same deceptively simple question: which proteins do these measured peptide fragments come from? Answering it means searching thousands, sometimes hundreds of thousands, of short amino acid sequences against enormous protein databases that are riddled with ambiguous characters. Those ambiguity codes, the wildcards of the protein world, have long been a computational headache, forcing researchers to choose between slow exhaustive searches and index structures that must be painstakingly precomputed. A new algorithm called Wild-AC, described in BMC Bioinformatics by Chris Bielow of Freie Universität Berlin, now promises to dissolve that dilemma with a clever twist on a fifty-year-old classic.
The classic in question is the Aho-Corasick algorithm, a cornerstone of computer science first published in 1975. Its genius lies in building a finite-state automaton from an entire set of search patterns at once, so that a single left-to-right pass over the text can locate every occurrence of every pattern simultaneously. For exact matching, it remains remarkably hard to beat: no preprocessing of the text is required, the automaton is built once from the patterns, and the scan proceeds at a pace independent of how many patterns are in play. That property, known as index-free operation, is precisely what makes it attractive for proteomics workflows, where the pattern set, the peptides, changes with every experiment while the text, the protein database, may be reused.
The trouble begins with wildcards. Protein databases routinely contain the ambiguity codes B, Z and X, which stand for ‘asparagine or aspartic acid’, ‘glutamine or glutamic acid’ and ‘any amino acid’, respectively. When a peptide pattern collides with one of these characters in the text, a standard Aho-Corasick automaton simply has no transition to follow, and the match is lost. Existing solutions have been unsatisfying in different ways. FM-indexes, the compressed full-text indexes popularized by genomics, can handle wildcards but demand substantial preprocessing of the text and become awkward when the wildcard sits in the searched text rather than in the pattern. Other multi-pattern methods, such as the widely used Wu-Manber algorithm, handle exact matching quickly but offer no native wildcard semantics at all.
Wild-AC’s central idea is disarmingly elegant: when the scan encounters a wildcard character in the text, it does not halt or fall back to a slow secondary routine. Instead, it branches the search into parallel ‘scout’ paths, one for each possible interpretation of the ambiguous residue, while the unmodified primary search continues unaffected through the common, wildcard-free case. Each scout carries its own state within the automaton and explores the subtree of possibilities that the wildcard opens up. Because wildcards are relatively rare, typically a few percent of database positions after standard masking, the scouts remain short-lived excursions rather than an exponential explosion. The primary automaton, meanwhile, behaves exactly like a textbook Aho-Corasick machine, so the overwhelming majority of matching work pays no wildcard tax whatsoever.
The engineering details matter as much as the concept. Wild-AC is implemented in C++ and released as open source, with support for multi-threading so that large peptide sets can be partitioned across cores. Because no index of the text is built, the algorithm starts working immediately on any database in plain sequence format, a meaningful advantage in laboratory pipelines where databases are swapped frequently, for example when searching against species-specific proteomes or contamination databases. The memory footprint stays modest, and the automaton construction scales gracefully with the number of patterns, which in practice ranges from a thousand peptides in a targeted experiment to half a million in a deep discovery run.
The benchmark results are where the algorithm earns its headline. Bielow tested Wild-AC across realistic proteomics workloads: pattern sets of 1,000 to 500,000 peptides with an average length of roughly 18 amino acids, searched against protein databases ranging from 3 million to 210 million characters. In the wildcard case, Wild-AC outperformed the FM-index outright. In pure exact matching, it matched or exceeded Wu-Manber, which is notable because Wu-Manber was designed specifically for fast multi-pattern exact search and has few peers. The advantage held across the entire realistic range of pattern counts, suggesting the algorithm is not a niche performer tuned to one benchmark configuration but a genuine workhorse.
One nuance deserves attention. When the databases were artificially masked to a wildcard rate of 5 percent, a scenario representing heavily annotated or deliberately ambiguous reference proteomes, the picture became more conditional. Wild-AC retained its advantage for large peptide sets, but the FM-index became the preferable choice for small ones. This makes intuitive sense: an index amortizes its construction cost over many queries, so if you plan to run only a handful of small searches against the same text, paying once for a prebuilt index can win. For the typical proteomics pattern of use, many peptides per search and databases that change often, Wild-AC’s index-free design tips the balance decisively the other way.
Why does this matter beyond the benchmark tables? Peptide-to-protein mapping sits at the foundation of nearly every downstream analysis in shotgun proteomics: protein identification, quantification, quality control and the detection of sequence variants all depend on it being fast and correct. Ambiguous amino acids are not an exotic edge case. They appear wherever sequences are incompletely characterized, in genomes assembled from noisy data, in databases that merge paralogous proteins, and in deliberately degenerate searches for modified or variant peptides. An algorithm that treats wildcards as a first-class citizen, without imposing index construction or sacrificing exact-search speed, removes a persistent friction from those pipelines. The author credits discussions with Sandro Andreotti on the algorithmic design and implementation input from Simon Gene Gottlieb, whose FM-index codebase and fuzzy amino acid matching implementation served as the benchmark comparison, with reviewer feedback from Ragnar Groot Koerkamp among others helping sharpen the final manuscript.
There is also a broader lesson for computational biology in how Wild-AC was built. Rather than inventing an entirely new data structure, the work takes a mature, well-understood algorithm and extends it minimally to cover a missing capability. The scout-path mechanism preserves the automaton’s behavior in the common case and pays for wildcard handling only where it is actually needed. This kind of surgical extension, validating performance against the strongest existing tools rather than straw men, is exactly the pattern that tends to produce software that gets adopted. The open-source C++ implementation, hosted publicly on GitHub, lowers the barrier for integration into existing search engines and quality-control tools, where Aho-Corasick machinery is often already present in some form.
For the proteomics community, the immediate takeaway is practical: if your pipeline maps large peptide sets against protein databases containing ambiguity codes, Wild-AC now offers the fastest known route, with no index to maintain and threads to spare. For small searches against a fixed, heavily masked database, the FM-index remains a sensible choice. For everyone else, the paper is a reminder that some of the most impactful bioinformatics advances still come from revisiting the classics of stringology and asking what happens when the text, not just the pattern, refuses to be unambiguous. As protein databases continue to swell with sequences of uneven certainty, algorithms that embrace that uncertainty at full speed will only grow in importance.
Subject of Research: A wildcard-enabled multi-pattern string matching algorithm for peptide-to-protein mapping in proteomics
Article Title: Wild-AC: A fast, index-free multi-pattern string matching algorithm with wildcard support for proteomics
Article References: Bielow, C. (2026). Wild-AC: A fast, index-free multi-pattern string matching algorithm with wildcard support for proteomics. BMC Bioinformatics. https://doi.org/10.1186/s12859-026-06686-8
Image Credits: AI Generated
DOI: 10.1186/s12859-026-06686-8
Keywords: Wild-AC, Aho-Corasick, string matching, wildcards, proteomics, peptide mapping, protein databases, FM-index, Wu-Manber, algorithms, bioinformatics, open source
