nf-core/deepmutscan
nf-core/deepmutscan is a reproducible, scalable, and community-curated pipeline for analyzing deep mutational scanning (DMS) data using shotgun DNA sequencing.
Define where the pipeline should find input data and save output data.
Path to comma-separated file containing information about the samples in the experiment.
string^\S+\.csv$The output directory where the results will be saved. You have to use absolute paths to storage on Cloud infrastructure.
stringEmail address for completion summary.
string^([a-zA-Z0-9_\-\.]+)@([a-zA-Z0-9_\-\.]+)\.([a-zA-Z]{2,5})$MultiQC report title. Printed as page header, used for filename if not otherwise specified.
stringNucleotide coordinates of the mutagenised open reading frame within the reference FASTA, in the format ‘start-stop’ (1-based, inclusive), e.g. ‘352-1383’.
string^\d+-\d+$Minimum number of counts for a variant to be retained. Variants observed fewer times are removed from the count tables and shown as dropouts in the count heatmaps.
integer1Minimum base quality (Phred) for a mismatch to be counted as a mutation, and for a base to contribute to sequencing coverage.
integer40Minimum distance (nt) a mutation or covered base must be from a read edge; also excludes those bases from coverage. Set to 0 to disable edge trimming.
integer2Sequencing-error correction of variant counts. ‘false_doubles’ (default) estimates the error rate of every single-nucleotide variant from reads that carry it together with a nearby programmed multi-nucleotide codon variant; ‘wildtype’ subtracts the background measured by additional deep sequencing of the unmutated template (requires samplesheet rows with type ‘wildtype’); ‘none’ disables correction.
stringEstimator used by the false-doubles correction. ‘mle’ (default) is a per-variant maximum-likelihood error rate; ‘eb’ is an empirical-Bayes estimate that pools across variants and shrinks per substitution class (steadier for sparse variants). Only applies when error_correction = ‘false_doubles’.
stringWindow (in codons, upstream and downstream) within which programmed multi-nucleotide codon variants are used to estimate the error rate of a single-nucleotide variant during false-doubles correction.
integer40Optional wildtype 3D structure (PDB). When supplied together with --fitness, an interactive variant effect inspection tool (self-contained HTML) is built that projects fitness, counts and error-correction biases onto the structure.
stringSliding-window size (in codons) used to smooth the positional coverage and count profiles in the library QC plots.
integer10Targeted number of counts per amino acid variant. Used to draw the sequencing coverage that would be required (assuming an even spread) in the positional coverage plots and the run report.
integer100Codons programmed at each position of the library. Choose from nnk, nns, nnh, nnn, nnk_nns, nnk_nns_nnh or custom. When using ‘custom’, also provide ‘–custom_codon_library’.
stringnnkPath to a .csv file defining a custom codon library. Required when ‘–mutagenesis_type custom’ is set. Provide either one global comma-separated list of codons without header (e.g. ‘AAA,AAC,AAG’), or a position-wise list with a header line containing ‘Position’, followed by one row per codon position with the position and its allowed codons (e.g. ‘1,ACG,AAA,ACA’ and ‘2,AAA,TTT,ACA’ on separate lines).
string/NULLAdditionally estimate fitness with DiMSum (Faure et al., 2020). Requires ‘–fitness’.
boolean,stringAdditionally estimate variant enrichment with mutscan (Soneson et al., 2023), using edgeR and limma. Requires ‘–fitness’.
boolean,stringEstimate variant fitness from matched ‘input’ and ‘output’ libraries of each sample, using the default log-ratio estimator.
boolean,stringEstimate sequencing-depth saturation of every library by closed-form hypergeometric rarefaction (on by default).
booleantrueReference genome related files and options required for the workflow.
Name of iGenomes reference. Not used for typical deep mutational scanning runs; provide --fasta instead.
stringPath to a FASTA file with the wildtype reference sequence of the mutagenised gene (exactly one sequence).
string^\S+\.fn?a(sta)?(\.gz)?$Do not load the iGenomes reference config.
booleanThe base path to the igenomes reference files
strings3://ngi-igenomes/igenomes/Parameters used to describe centralised config profiles. These should not be edited.
Git commit id for Institutional configs.
stringmasterBase directory for Institutional configs.
stringhttps://raw.githubusercontent.com/nf-core/configs/masterInstitutional config name.
stringInstitutional config description.
stringInstitutional config contact information.
stringInstitutional config URL link.
stringLess common options for the pipeline, typically set in a config file.
Display version and exit.
booleanMethod used to save pipeline results to output directory.
stringEmail address for completion summary, only when pipeline fails.
string^([a-zA-Z0-9_\-\.]+)@([a-zA-Z0-9_\-\.]+)\.([a-zA-Z]{2,5})$Send plain-text email instead of HTML.
booleanFile size limit when attaching MultiQC reports to summary emails.
string25.MB^\d+(\.\d+)?\.?\s*(K|M|G|T)?B$Do not use coloured log outputs.
booleanCustom config file to supply to MultiQC.
stringCustom logo file to supply to MultiQC. File name must also be set in the MultiQC config file
stringCustom MultiQC yaml file containing HTML including a methods description.
stringBoolean whether to validate parameters against the schema at runtime
booleantrueBase URL or local path to location of pipeline test dataset files
stringhttps://raw.githubusercontent.com/nf-core/test-datasets/Suffix to add to the trace report filename. Default is the date and time in the format yyyy-MM-dd_HH-mm-ss.
stringDisplay the full detailed help message.
booleanDisplay hidden parameters in the help message (only works when –help or –help_full are provided).
boolean