Logan Search
In collaboration with Téo Lemane from CEA Genoscope, with our partner Rayan Chikhi from the Sequences Bioinformatics team at Pasteur Institute, and the Artem Babaian‘s group we introduced “Logan Search” [1] which allows you to search for any DNA sequence in minutes, bringing Earth’s largest genomic resource to your fingertips.
Under the hood, we built a 1 petabyte k-mer index for all 27 million sequencing datasets in the SRA up to 12-2023.
Logan Search transforms your query to its k-mers (k=31), and in the time it takes to brew a coffee, it will retrieve every dataset containing your k-mers. It’s the only service working at this scale.
The output datasets are easily visualized with custom plots in Logan Search, which accesses a harmonized set of query and SRA meta-data including sequencing technology, type of molecule, geographic distribution, and sample origins. Learn more about your sequence.
Logan Search returns a list of SRA accessions, not alignments. To bring you closer to the data we’ve also created a microservice to instantly retrieve Logan contigs matching your search.
Kaminari
With Kaminari [2] explore a new index design based on minimizers and integer compression methods. We show that a careful implementation of this design outperforms previous solutions based on Bloom filters by a wide margin.
Back to sequences
We proposed a simple yet useful tool when dealing with large genomic datasets. This is “Back to sequences: Find the origin of k-mers [3]”.
The backpack quotient filter
A dynamic and space-efficient data structure for querying k-mers with abundance [4]
kmcomp
Lossless compression of k-mer matrices enabling random row access [5]. Used for compressing kmindex indexes.
kmhelpers
A Python toolkit for managing, compressing, and querying kmindex indices efficiently [6]
Interpolating and Extrapolating Node Counts in Colored Compacted de Bruijn Graphs for Pangenome Diversity
In response to the evolution of pangenome representation, we introduce a novel method for comparing pangenomes by their node counts, addressing two main challenges: the variability in node counts arising from graphs constructed with different numbers of genomes, and the large influence of rare genomic sequences. [7]
Publications
[1] Chikhi, R., Lemane, T., Loll-Krippleber, R., Montoliu-Nerin, M., Raffestin, B., Camargo, A. P., … & Babaian, A. (2024). Logan: planetary-scale genome assembly surveys life’s diversity. bioRxiv, 2024-07.
[2] Levallois, V., Shibuya, Y., Le Gal, B., Patro, R., Peterlongo, P., & Pibiri, G. E. (2025). Kaminari: a resource-frugal index for approximate colored k-mer queries. bioRxiv, 2025-05.
[3] Baire, A., Marijon, P., Andreace, F., & Peterlongo, P. (2024). Back to sequences: Find the origin of k-mers. Journal of Open Source Software, 9(101), 7066, https://doi.org/10.21105/joss.07066
[4] Levallois, V., Andreace, F., Le Gal, B., Dufresne, Y., & Peterlongo, P. (2024). The backpack quotient filter: A dynamic and space-efficient data structure for querying k-mers with abundance. Iscience, 27(12).
[5] Regnier, A., Lemane, T., Bellenous, S., Chikhi, R., & Peterlongo, P. (2026). Lossless compression of k-mer matrices enabling random row access. bioRxiv, 2026-07.
[6] Bellenous, S., Jaron, K., Peterlongo, P. (2026). kmhelpers: A Python toolkit for automated management of genomic indexes. Submitted to JOSS.
[7] Parmigiani, L., & Peterlongo, P. (2026). Interpolating and Extrapolating Node Counts in Colored Compacted de Bruijn Graphs for Pangenome Diversity. Journal of Computational Biology, 15578666261486474.