nf-core/plasmodiumdrugres
Pipeline for analyzing drug resistance markers from Plasmodium microhaplotype data. It translates variants into amino acid changes at drug resistance loci and estimates allele frequencies and prevalences at both single-locus and multi-locus levels. Microhaplotype data can be supplied in the form of an allele table or a PMO file.
Introduction
nf-core/plasmodiumdrugres is a bioinformatics pipeline for analyzing drug resistance markers from microhaplotype data. It translates variants into amino acid changes at drug resistance loci and estimates allele frequencies and prevalences at both single-locus and multi-locus levels. Microhaplotype data can be supplied in the form of an allele table or a PMO file.
- Translate loci of interest (
PGEcore) - Split by population
- Estimate allele prevalence (
PGEcore) - Estimate multilocus allele frequency. Choice of method between:
- Estimate single locus allele frequency. Choice of method between:
- Merge prevalence and frequency outputs
- Concatenate population outputs into standardized summary tables, while preserving full tool-specific columns in
raw_summaries/
Usage
If you are new to Nextflow and nf-core, please refer to this page on how to set-up Nextflow. Make sure to test your setup with -profile test before running the workflow on actual data.
The pipeline accepts either a Portable Microhaplotype Object (PMO) or an allele table plus panel BED. You must also provide a loci of interest BED. Optionally provide loci groups for multi-locus estimates and a population assignment for multi-population runs.
Now, you can run the pipeline using a PMO:
nextflow run nf-core/plasmodiumdrugres \ -profile <docker/singularity/.../institute> \ --pmo input_file.pmo \ --loci_of_interest_bed loci_of_interest.bed \ --loci_groups loci_groups.tsv \ --outdir <OUTDIR>Or with an allele table:
nextflow run nf-core/plasmodiumdrugres \ -profile <docker/singularity/.../institute> \ --allele_table allele_table.tsv \ --panel_info_bed panel_info.bed \ --loci_of_interest_bed loci_of_interest.bed \ --loci_groups loci_groups.tsv \ --outdir <OUTDIR>Please provide pipeline parameters via the CLI or Nextflow -params-file option. Custom config files including those provided by the -c Nextflow option can be used to provide any configuration except for parameters; see docs.
For more details and further functionality, please refer to the usage documentation and the parameter documentation.
Pipeline output
To see the results of an example test run with a full size dataset refer to the results tab on the nf-core website pipeline page. For more details about the output files and reports, please refer to the output documentation.
Main results include standardized sl_summary.tsv and ml_summary.tsv tables, plus full tool-specific concatenated tables under raw_summaries/.
Credits
nf-core/plasmodiumdrugres was originally written by PlasmoGenEpi.
We specifically thank the following people for their extensive assistance in the development of this pipeline:
- Kathryn Murie
- Nicholas Hathaway
- Alfred Hubbard
- Jorge Amaya-Romero
A special thanks to everyone in the community who contributes to PGEcore and continues to do so.
Contributions and Support
If you would like to contribute to this pipeline, please see the contributing guidelines.
For further information or help, don’t hesitate to get in touch on the Slack #plasmodiumdrugres channel (you can join with this invite).
Citations
An extensive list of references for the tools used by the pipeline can be found in the CITATIONS.md file.
You can cite the nf-core publication as follows:
The nf-core framework for community-curated bioinformatics pipelines.
Philip Ewels, Alexander Peltzer, Sven Fillinger, Harshil Patel, Johannes Alneberg, Andreas Wilm, Maxime Ulysse Garcia, Paolo Di Tommaso & Sven Nahnsen.
Nat Biotechnol. 2020 Feb 13. doi: 10.1038/s41587-020-0439-x.