Bio Reads QC Mapping
Ingest, QC, and map reads with reproducible outputs. Use for raw read processing and coverage stats.
Instructions
Tool guides and versions: docs/README.md.
-
Parse and validate
sample_sheet.tsvagainstschemas/sample-sheet.schema.json. Use the executable driver for both planning and restartable execution:uv run --script skills/bio-reads-qc-mapping/scripts/run_reads_qc_mapping.py \ sample_sheet.tsv --out results/bio-reads-qc-mapping # Inspect run_manifest.json, then execute the same plan: uv run --script skills/bio-reads-qc-mapping/scripts/run_reads_qc_mapping.py \ sample_sheet.tsv --out results/bio-reads-qc-mapping --executeread_typemust bepaired_short,single_short, orlong. Mapping runs only for rows with a non-emptyreference; a missing reference is not a mapping failure. The driver reuses a stage only when its declared outputs are non-empty and the stage's.donemarker exists. -
For short reads: run QC and adapter/quality trimming with
bbdukorfastpv1.3.3+. -
For long reads: use current basecaller-aware QC first. For ONT, prefer Dorado summaries/trimming during basecalling or demultiplexing when starting from signal/BAM; for FASTQ-only filtering use
chopperfor quality/length/end trimming orfiltlongv0.3.1 when selecting reads for assembly (v0.3.0 renamed the short-read options to--short_1/--short_2; see docs/filtlong.md). UsePychopperfor full-length cDNA. TreatPorechop_ABIas a targeted legacy/fallback adapter-discovery tool, and record why it is needed.- For very large ONT FASTQ inputs, do not burn the first full read pass on raw
gzip -tor rawseqkit statspreflight unless the user explicitly asks for it. Record rawstatmetadata and, if needed, a small sampled sanity check; let the first full pass be the actual filtering/orientation step, then runseqkit statson produced outputs. - For ONT cDNA with
Pychopper, write outputs with plain.fastqsuffixes unless you explicitly pipe/compress them yourself.Pychoppercan write plain FASTQ even when the output path ends in.gz; avoidgzip -tonPychopperoutputs unless magic bytes confirm gzip. If legacy outputs have.fastq.gznames but plain FASTQ content, rename them to.fastqbefore resuming. Pychopperreport plotting can fail after the reads are already processed, for example from a pandas/statistics type-conversion error. On that failure, inspect whether the classified/unclassified/rescued/read-stats outputs exist and are non-empty. If they do, resume downstream from those outputs rather than rerunning the fullPychopperpass.
- For very large ONT FASTQ inputs, do not burn the first full read pass on raw
-
Map reads and produce coverage tables:
- Short reads, CPU:
bbmaporbwa-mem2v2.2.1+. Short reads, GPU node available: NVIDIA Parabricksfq2bam(wrapsbwa-mem2+ GATK markdup; typically 3–4× faster thanbwa-mem2on 8 cores and up to ~80× over a 96-core CPU pipeline). - Long reads, CPU:
minimap2v2.30+. AVX-512 hardware:mm2-fastas a drop-in replacement (~1.8× speedup). GPU node available:mm2-gbormm2-axfor CUDA-accelerated long-read alignment.
- Short reads, CPU:
-
Record the tool, version, and any GPU device used in the run log.
Input Requirements
Prerequisites:
- Tools declared in the project's pinned Pixi environment. See
docs/README.mdfor expected tools. - Sample sheet and reads are available. Inputs:
- sample_sheet.tsv
- reads/*.fastq.gz
- reference.fasta (optional)
Output
- results/bio-reads-qc-mapping/trimmed_reads/
- results/bio-reads-qc-mapping/qc_reports/
- results/bio-reads-qc-mapping/mapping_stats.tsv (
mapping_statusisplannedbefore--execute, thencompletedorreused, andnot_requestedfor rows with no reference) - results/bio-reads-qc-mapping/logs/
- stdout: the last line is one JSON envelope
{ok, skill, out, manifest, warnings}(driver stdout contract in AGENTS.md) - Sample sheet contract: schemas/sample-sheet.schema.json
Quality Gates
- [ ] Post-QC read count sanity checks pass.
- [ ] Mapping rate meets project thresholds.
- [ ] On execution failure, preserve logs and report the failed command; retry only after diagnosing the cause and recording the changed parameters. Report unmet biological thresholds as results; never tune parameters solely to pass a gate.
- [ ] Validate sample sheet schema and FASTQ integrity.
- [ ] The plan covers every sheet row exactly once, and mapping gates are applied only to rows that supplied a reference.
- [ ] For long-read QC, record whether trimming happened in the basecaller/demultiplexer,
chopper,filtlong,Pychopper, or a documented Porechop_ABI fallback. - [ ] For huge ONT inputs, avoid redundant full-file raw preflights; document raw file size/mtime and make the first full pass productive.
- [ ] For
Pychopperoutputs, verify actual file type by content, not suffix. Plain FASTQ with a.gzsuffix must be renamed or explicitly compressed before downstream tools that expect gzip. - [ ] Resume guards skip a stage only when its expected outputs are non-empty and its
.donemarker exists. Run a lightweight content check (seqkit stats, FASTQ header sniff, or gzip magic as appropriate) before accepting downstream data.
Examples
The runnable fixture at fixtures/sample_sheet.tsv covers paired-end, single-end, and long reads.
Troubleshooting
Issue: Pychopper failed during report/stat plotting but output FASTQs exist
Solution: Treat this as a recoverable post-processing failure. Confirm the classified FASTQ is non-empty and readable, fix any misleading .gz suffix, run seqkit stats, and resume downstream steps from the existing Pychopper outputs.
Scan to join WeChat group