What is it about?
Finding which DNA tags track male or female usually starts from restriction site-associated DNA sequencing (RAD-seq): short reads next to restriction sites, cheap enough to sequence many individuals of a non-model species. RADSex is the reference workflow that builds a marker-by-individual depth table and tests which tags are sex-biased. Once a panel reaches millions of tags, those table-building commands grow hungry for memory, the output is a frequentist call with no posterior, and there is no Python or C interface. rsx is a Rust implementation of the full RADSex command set that keeps the same marker-table semantics and the same command line. After the FASTQ files are ingested, writable allocations stay bounded by the number of individuals or by an explicit buffer; false-discovery-rate ranking is the one deliberate exception. Conjugate Beta-Binomial Bayes factors and directional posteriors grade each marker as a strict call, a posterior-supported hypothesis, or a Bayes-factor-only row. An optional CUDA backend batches the per-marker arithmetic on the GPU. On four published RAD-seq panels covering 41.9 billion sequenced bases, rsx reproduced the RADSex v1.2.0 calls, recovered every Bonferroni-significant positive-control marker, and was 8.38-fold faster in geometric mean across 56 paired timings. The CUDA path adds up to 29.86-fold on the p-value batch at one million markers on an NVIDIA A100. Python and C bindings drive the same core from notebooks and pipelines.
Featured Image
Why is it important?
Non-model sex-determination studies still run on reduced-representation data. A tool that forces you to downsample rare tags, or that emits one ungraded marker list, either drops the interesting candidates or over-calls sparse artefacts. rsx keeps the published RADSex interface so existing studies stay comparable, then adds the two things those studies lacked: full-table execution that does not grow with the number of tags, and explicit evidence grades. On the ayu and tench positive-control panels it recovers every matched Bonferroni-significant marker and adds seven posterior-supported candidates. On zebra danio, treated as a frequentist null in the source paper, it surfaces 30 W-linked hypotheses. On the Antarctic notothenioid panel it withholds 400 Bayes-factor-only rows from the sex-system call because their posteriors stay compatible with a low-prevalence null. That split is what a PCR or mapping decision needs: which rows are calls, which are validation targets, and which are only marginal evidence.
Perspectives
I kept hitting the same wall. RADSex is the right workflow, but the C++ table-building path does not stay in RAM once the panel has millions of tags, and a single chi-squared list is not enough when we need to decide which candidates are worth a PCR. I rewrote the command set in Rust so the memory bound is the number of individuals, not the number of tags, and so the same counts can carry a Bayes factor and a posterior. The compatibility constraint was non-negotiable. If the new code disagreed with RADSex on the positive-control panels, the speedup would be worthless. Matching v1.2.0 on every Bonferroni-significant marker, then adding graded extras rather than overwriting the frequentist result, is the design I wanted the first time I ran a large popmap. Python and C bindings exist because the biology does not stop at the command line. If you already have a RADSex table, rsx should drop in, stay in memory, and tell you how sure it is.
Rohit Goswami
University of Iceland
Read the Original
This page is a summary of: rsx: a high-performance streaming toolkit for RAD-seq sex determination, BMC Bioinformatics, September 2026, Springer Science + Business Media,
DOI: 10.1186/s12859-026-06628-4.
You can read the full text:
Contributors
The following have contributed to this page







