×

Nextflow Modules

Clear

Showing module(s) with keyword "catalogue"

Module Keywords Description
nf-core/custom/orfcollapse orf ribo-seq catalogue smorf deduplication Collapse small ORFs that share an amino-acid sequence cluster into a single catalogue entry. Pair with `custom/orfmerge` (coordinate-based catalogue), `bedtools/getfasta` + `seqkit/translate` (AA FASTA keyed by orf_id), and `mmseqs/easycluster` (AA clusters) upstream. The coordinate-based merge in `custom/orfmerge` only groups ORFs that overlap on the genome, so the same micropeptide encoded at several distinct, non-overlapping loci (typically repetitive regions) survives as separate rows. This adopts the peptide-level deduplication and 0.9 amino-acid-similarity threshold of the GENCODE Ribo-seq ORF consolidation (Mudge et al. 2022, Nat Biotechnol, doi:10.1038/s41587-022-01369-0; gencode-riboseqORFs collapse_cutoff 0.9), implemented here with MMseqs2 sequence-identity clustering rather than that tool's longest-shared-string / P-site-overlap metric. Small ORFs (`aa_length` <= `--smorf-max-aa`, default 100) are clustered by amino-acid identity upstream and this module folds each multi-member cluster down to one representative. Only small ORFs are collapsed; larger ORFs are passed through untouched. Eligibility is the catalogue's `is_smorf` flag, independent of `orf_class`, so a short uORF and a short novel ORF are both candidates; `--smorf-max-aa` re-derives the flag and aborts on disagreement. Among the members of a cluster the representative is chosen by class specificity, then longest aa_length, then orf_id, so the result does not depend on which sequence MMseqs2 labelled the cluster representative. Catalogue row order is preserved; dropped members fold their `called_by_<caller>` / `score_<caller>` evidence, `n_samples` / `samples` recurrence and gene mappings into the survivor.
nf-core/custom/orfmerge orf ribo-seq catalogue merge clustering Cluster normalised per-sample, per-caller ORF predictions into a single cohort-level catalogue. Pair with `custom/orfnormalise` upstream and (typically) `bedtools/getfasta` + `seqkit/translate` downstream to obtain the AA FASTA. Rows sharing an identical exon structure are grouped first and clustered as a single proxy chosen by class specificity, then the members are restored. Class selects the strategy, so without this an ORF that two callers class differently would be split across strategies and could never merge. Proxies are then partitioned by clustering strategy. The harmonised `orf_class` written by `custom/orfnormalise` selects the strategy but is never part of a grouping key, because callers disagree on class for the same ORF (Ribo-TISH reports `5'UTR` for both uORFs and CDS-overlapping uORFs) and keying on it would emit one catalogue row per disagreeing caller: - canonical_cds: grouped by (transcript_id, strand), then reciprocal-overlap clustered within the transcript so a short truncated variant is not folded into the full-length CDS. - uORF, uoORF, dORF, collapse by (transcript_id, strand, start, doORF, intORF, other: end). A single transcript can host multiple distinct uORFs / dORFs / internal ORFs, so keying on the outer span keeps them in separate clusters while still merging cross-caller calls that agree on coordinates. - novel_u: greedy reciprocal-overlap clustering on summed exon-block intersection at `--reciprocal-overlap` (default 0.8). Catches fuzzy cross-caller matches and exact-coordinate collapses in one pass. Order-dependent at the boundary: a chain A-B-C where A-B and B-C overlap at ~0.85 but A-C only at ~0.75 may cluster as {A,B,C} or {A,B}+{C} depending on iteration order. Rare in practice at 0.8. Cross-caller consensus is recorded in two column families on the catalogue TSV: - `called_by_<caller>`: 0/1 indicator per supported caller (ribotish, ribocode, ribotricer, rpbp, price). - `score_<caller>`: best score from that caller within the cluster. Score direction is per-caller (p-values are minimised; Bayes factors / phase scores are maximised). Cross-sample recurrence is recorded in two further columns: - `n_samples`: number of distinct samples contributing to the cluster (a cohort recurrence metric). - `samples`: sorted, comma-separated list of those sample ids. Emits a small MultiQC custom-content TSV (per-class counts) for inclusion in downstream MultiQC reports. Alongside the full catalogue, emits a consensus view (`*.consensus.*`) filtered to ORFs supported by at least `--min-callers` distinct callers and recurring in at least `--min-samples` samples (both default 1, i.e. no filtering, so the consensus view equals the full catalogue). Raising either threshold yields a higher-confidence catalogue without altering the full one.