Nature Communications

K-MARVEL: K-Mer-based antimicrobial resistance virtual exploration lab

2026-09-03

The rapid spread of antimicrobial resistance (AMR) necessitates new computational surveillance tools. Current methods for analysis of next-generation sequencing data have trade-offs: assembly-based approaches are computationally intensive, and direct long-read mapping is hampered by high error rates that obscure resistance-conferring mutations. Here, we present K-MARVEL, an open-source method that captures antimicrobial resistance genes (ARGs) and resistance-conferring mutations from both short- and long-read datasets. Operating in protein k-mer space, K-MARVEL tolerates nucleotide-level sequencing errors. We benchmarked K-MARVEL on 209 long-read and 205 short-read datasets across 22 bacterial species and achieved F1-scores of 0.976 (short-read) and 0.958 (long-read), outperforming assembly-based methods in speed and memory usage. K-MARVEL had higher F1-scores than seven widely used short-read-based ARG classifiers (0.979) and two widely used long-read-based classifiers (0.961) for homology-model-based ARGs. K-MARVEL can accurately identify both homologous ARGs and structural genes containing resistance-conferring mutations, including multiple variants, directly from raw sequencing data.

Full text

DOI https://doi.org/10.1038/s41467-026-77438-8