A reproducible wrapper around nf-core/circdna v1.1.0
for detecting extrachromosomal circular DNA (ecDNA) and other circular DNA
from whole-genome sequencing (WGS). It pins the pipeline version, keeps all
parameters in versioned -params-file YAMLs (one per detection branch), and
keeps licensed/large inputs (Mosek license, AmpliconArchitect data repo,
containers) out of the repo.
The wrapper does not reimplement anything — it launches the upstream
nf-core/circdna pipeline reproducibly.
| Branch | Tool | Best for | On WGS |
|---|---|---|---|
ampliconarchitect |
AmpliconArchitect (AmpliconSuite) | Amplified ecDNA | ✅ validated WGS branch |
circle_map_realign |
Circle-Map Realign | eccDNA | |
circle_map_repeats |
Circle-Map Repeats | repetitive eccDNA | |
circexplorer2 |
CIRCexplorer2 | circular junctions |
For amplified ecDNA from WGS, use ampliconarchitect — the branch the
pipeline documents as WGS-only. The other three are included for smaller
eccDNAs; on WGS they can yield more false positives, so treat their calls with
caution.
- Nextflow ≥ 22.10 and a container engine (Singularity/Apptainer recommended, or Docker).
- For the
ampliconarchitectbranch: the AmpliconArchitect data repo, a Mosek license (mosek.lic), and the pipeline containers. If you keep these in a local folder (e.g.tsd_transfer), pointenv.shat them — see below. Get them via AmpliconSuite-pipeline.
# 1. one-time: tell the wrapper where your local resources live
cp env.sh.example env.sh
$EDITOR env.sh # set AA_DATA_REPO, MOSEK_LICENSE_DIR, NXF_SINGULARITY_CACHEDIR
# 2. list your WGS samples (edit with real absolute paths)
$EDITOR assets/samplesheet_bam.csv # sample,bam (or use samplesheet_fastq.csv)
# 3. run a branch (amplified ecDNA from WGS)
./run.sh ampliconarchitect
# ...or: make ampliconarchitect
# other branches:
./run.sh circle_map_realign
./run.sh circle_map_repeats
./run.sh circexplorer2run.sh pins -r 1.1.0, loads env.sh, injects --aa_data_repo and
--mosek_license_dir from your environment (only for ampliconarchitect), and
launches with -params-file params/<branch>.yaml -c nextflow.config.
Edit assets/samplesheet_bam.csv (WGS aligned BAMs):
sample,bam
tumor_wgs_1,/abs/path/tumor_wgs_1.bamOr start from FASTQ with assets/samplesheet_fastq.csv and set
input_format: FASTQ + input: assets/samplesheet_fastq.csv in the branch's
params file.
- Parameters live in
params/*.yaml(genome,circle_identifier,reference_build,aa_cngain, resource caps, …). circdna requires parameters to be passed via-params-file, not-c. - Resources / containers live in
nextflow.config, passed with-c(executor, retries, Singularity settings only — no parameters). - Machine-specific paths live in
env.sh(git-ignored).
Default genome/reference_build is GRCh38; for GRCh37 set both to
GRCh37 (and use a matching AA data repo build). For mouse use mm10.
- Pipeline version pinned to 1.1.0 in
run.sh(nextflow pull nf-core/circdnato refresh the cache). - All run settings are captured in the committed
params/*.yaml— commit the params file alongside your results. .gitignoreblocks the Mosek license, AA data repo, containers, BAM/FASTQ, andwork//results/so licensed or large data is never committed. Keep patient/controlled WGS data and any licensed files out of this public repo.
run.sh pinned launcher (./run.sh <branch>)
params/<branch>.yaml parameters per detection branch (-params-file)
nextflow.config executor/container/resource config (-c)
assets/samplesheet_*.csv input templates (BAM / FASTQ)
env.sh.example copy to env.sh; local paths to AA repo/Mosek/containers
ci/validate.py lint params + samplesheets (run by CI)
Makefile make ampliconarchitect | validate | clean
This wrapper runs nf-core/circdna, originally written by Daniel Schreyer (University of Glasgow). Please cite the pipeline (doi:10.5281/zenodo.6685250) and the underlying tools (AmpliconArchitect, Circle-Map, CIRCexplorer2). See the nf-core/circdna citations.
MIT — see LICENSE. (Applies to this wrapper only; nf-core/circdna and the
bundled tools carry their own licenses.)