This repository contains data-processing workflows, quality-control procedures, and analysis code for the Los Angeles County Fire Department (LACoFD) firefighter DNA methylation study, led by Thomas Sullivan (LACoFD) and Janine LaSalle (University of California, Davis).
The study is part of the California Firefighter Cancer Prevention and Research Program (Award: F01FF8767) and examines molecular changes associated with occupational exposure to products of combustion. Firefighters can be repeatedly exposed to potentially harmful chemicals during training and active duty. The study uses longitudinal samples collected before and after exposure to investigate whether these exposures are associated with changes in DNA methylation.
The project includes new firefighter recruits and experienced instructors, who differ in their histories and frequency of occupational exposure. Chemical exposure measurements currently include volatile organic compounds (VOCs) and polycyclic aromatic hydrocarbons (PAHs). Cell-free DNA (cfDNA) is isolated from plasma and characterized before whole-genome bisulfite sequencing (WGBS). DNA methylation profiles generated from WGBS will subsequently be examined in relation to exposure measures and participant characteristics.
The long-term goal of the project is to better understand molecular changes associated with repeated occupational exposures in firefighters and to identify potential blood-based biomarkers relevant to cancer risk.
The primary workflow is:
LACoFD firefighter recruitment
│
├── New recruits
└── Experienced instructors
│
▼
Pre-exposure plasma
│
▼
Training / occupational exposure
│
▼
Post-exposure plasma
│
▼
cfDNA isolation
│
▼
DNA quantity and quality QC
├── Qubit
└── Bioanalyzer
│
▼
WGBS
│
▼
Epigenerator processing
│
▼
Cytosine reports
│
▼
DNA methylation analysis
Downstream DNA methylation analyses are under development. Approaches being considered include region-based differential methylation analyses and DNA methylation network analyses using tools such as DMRichR and CoMethyL.
The repository separates original data, intermediate processing, analysis-ready datasets, computational pipeline runs, scripts, and generated results.
LACOFD/
├── config/
│
├── data/
│ ├── raw/
│ ├── interim/
│ ├── processed/
│ └── metadata/
│
├── docs/
│
├── pipeline_runs/
│ └── epigenerator/
│
├── results/
│ ├── sequencing_qc/
│ └── wet_lab_qc/
│
├── scripts/
│ ├── repo/
│ ├── slurm/
│ └── wet_lab/
│
└── tests/
Contains source files in their original form. These files should not be modified during analysis.
Current raw wet-lab data include:
- sample manifests;
- DNA isolation records;
- Bioanalyzer submission forms; and
- Bioanalyzer reports.
DNA isolation records are organized by processing batch.
data/raw/wet_lab/dna_isolation/
├── arizona/
│ └── DNAEXT_AZ_01/
└── la_county/
├── DNAEXT_LA_01/
└── DNAEXT_LA_02/
The arizona designation reflects sample-processing and shipping origin and does not represent a separate Arizona firefighter cohort. These samples originated from the LACoFD study but were handled through an Arizona-associated workflow. This distinction is retained in the repository to preserve sample-processing origin.
Contains cleaned or standardized datasets derived from the original source files.
Examples include datasets with:
- standardized variable names;
- harmonized timepoint definitions;
- standardized sample identifiers;
- extraction batch identifiers; and
- calculated DNA yield variables.
Intermediate files may be regenerated from the corresponding raw files and scripts.
Contains datasets prepared for statistical analysis or downstream computational workflows.
For example:
data/processed/wet_lab/dna_isolation/
contains analysis-ready DNA isolation datasets.
As WGBS processing is completed, final cytosine reports selected for downstream DNA methylation analyses will also be stored within the processed-data structure. Full pipeline working directories remain separate under pipeline_runs/.
Contains study metadata, data dictionaries, variable descriptions, and related supporting information.
Contains working directories generated during computational processing.
WGBS preprocessing is performed usingepigenerator, a reproducible workflow for processing bisulfite sequencing data.
Epigenerator project repository: vhaghani26/epigenerator
Individual processing runs are maintained separately:
pipeline_runs/epigenerator/
├── EPI_PILOT_01/
├── EPI_AZ_01/
├── EPI_LA_01/
└── EPI_LA_02/
These directories may contain raw-sequence links, trimmed reads, alignment outputs, methylation extraction files, cytosine reports, quality-control files, and computational logs.
pipeline_runs/ should therefore be considered a computational workspace rather than the permanent location of analysis-ready data.
Contains generated figures, tables, and quality-control summaries.
Current result categories include:
results/
├── sequencing_qc/
└── wet_lab_qc/
Wet-lab DNA isolation results are further organized by processing batch where appropriate.
For example:
results/wet_lab_qc/dna_isolation/la_county/
└── DNAEXT_LA_01/
Contains reproducible scripts used for data cleaning, quality control, plotting, and computational processing.
Scripts are grouped according to their purpose rather than by individual dataset whenever possible.
For example:
scripts/
├── repo/
│ └── reorganize_repo.sh
│
├── slurm/
│
└── wet_lab/
└── dna_isolation/
├── 01_clean_dna_isolation.R
├── 02_plot_dna_isolation_pre_post.R
└── submit_rscript.sbatch
The DNA isolation scripts accept command-line arguments so the same workflow can be applied across multiple extraction batches without modifying hard-coded file paths.
Detailed execution instructions should be documented within README files located in the relevant script directories.
Contains supporting documentation including study methods, wet-lab protocols, research plans, and reproducibility documentation.
Processing steps are assigned stable identifiers so that sample origin can be followed across wet-lab and computational workflows.
DNA extraction batches use:
DNAEXT_<GROUP>_<BATCH>
Examples:
DNAEXT_AZ_01
DNAEXT_LA_01
DNAEXT_LA_02
where:
DNAEXT= DNA extraction;AZ= Arizona-associated processing;LA= Los Angeles County processing; and01,02, etc. = sequential processing batch.
The batch identifier describes the DNA extraction event and should not be interpreted as a sequencing or computational batch.
Bioanalyzer processing uses:
BIOA_<GROUP>_<RUN>
Examples:
BIOA_LA_01
BIOA_LA_02
BIOA_LA_03
Each Bioanalyzer run can contain the associated submission information and instrument-generated reports.
Sequencing batches, when assigned, use:
SEQ_<GROUP>_<BATCH>
For example:
SEQ_LA_01
Sequencing batches are tracked independently from DNA extraction batches because samples from multiple extraction batches may be sequenced together.
Epigenerator processing runs use:
EPI_<GROUP>_<RUN>
Examples:
EPI_PILOT_01
EPI_AZ_01
EPI_LA_01
EPI_LA_02
where EPI identifies an Epigenerator processing run.
DNA extraction, Bioanalyzer, sequencing, and Epigenerator identifiers therefore describe different stages of sample processing:
DNAEXT_LA_01 DNA extraction batch
BIOA_LA_01 Bioanalyzer run
SEQ_LA_01 Sequencing batch
EPI_LA_01 Epigenerator processing run
These identifiers should be maintained as separate metadata variables rather than inferred from one another.
Cell-free DNA isolation data are processed using reusable R scripts located in:
scripts/wet_lab/dna_isolation/
The general data flow is:
Original extraction workbook
│
▼
data/raw/
│
▼
Cleaning and standardization
│
▼
data/interim/
│
▼
Analysis-ready dataset
│
▼
data/processed/
│
▼
QC figures and summaries
│
▼
results/wet_lab_qc/
Current DNA isolation QC includes evaluation of:
- starting plasma volume;
- DNA concentration;
- total DNA yield;
- DNA yield normalized to starting plasma volume; and
- paired pre- and post-exposure measurements where available.
Detailed instructions for executing these scripts are maintained separately within the DNA isolation scripts directory.
Whole-genome bisulfite sequencing data are processed using the vhaghani26/epigenerator workflow.
The full Epigenerator run is retained under:
pipeline_runs/epigenerator/<RUN_ID>/
A typical run produces sequential outputs including:
raw sequence links
↓
read trimming
↓
sequence screening
↓
alignment
↓
deduplication
↓
methylation extraction
↓
cytosine reports
↓
MultiQC
Large pipeline intermediates remain within the pipeline working directory. Final files selected for downstream analyses are copied or linked into the appropriate processed-data directory.
This distinction allows the complete computational workflow to be retained while keeping analysis-ready datasets separate from transient pipeline outputs.
The LACOFD study is a collaborative project between the Los Angeles County Fire Department, the University of California, Davis, and collaborating investigators.
Current team members include:
| Team member | Role |
|---|---|
| Thomas Sullivan | Los Angeles County Fire Department Co-Principal Investigator |
| Janine LaSalle | UC Davis Co-Principal Investigator |
| Shehnaz Hussain | UC Davis Co-Investigator |
| George Kuodza | UC Davis Postdoctoral Scholar |
| Logan Williams | Former UC Davis Graduate Student |
| Additional team members | To be added |
The team list will be updated as project roles and contributors are finalized.
This work is supported by the California Firefighter Cancer Prevention and Research Program, University of California Office of the President.
Project: Examining longitudinal changes in DNA methylation in firefighters exposed to products of combustion
Award: F01FF8767
Fire service Co-Principal Investigator: Thomas Sullivan, Los Angeles County Fire Department
UC Co-Principal Investigator: Janine LaSalle, University of California, Davis
Wet-lab processing, sample quality control, WGBS processing, and development of downstream DNA methylation analyses are ongoing. Repository organization, analysis scripts, and documentation will continue to be updated as additional samples and processing batches become available.