Welcome! This repository contains the source material for the open-source textbook "Reproducible Research Using R". This project serves as a comprehensive guide to performing statistical analysis and creating transparent, reproducible reports.
This book represents the applied companion to
Reproducible Research Using R.
Over the course of the semester, students transformed structured assignments into fully reproducible analyses using real datasets, transparent workflows, and public-facing communication. What began as individual homework submissions evolved into polished research chapters assembled into this cohesive volume.
This book treats reproducibility not as an enhancement, but as a default professional practice.
Each chapter reflects a progression of skills: - data import\
- cleaning\
- visualization\
- statistical testing\
- modeling\
- interpretation\
- communication
More importantly, each chapter reflects a shift in mindset — from “running code” to designing complete analytical workflows.
This is not just a collection of assignments.
It is a curated portfolio of applied reproducible research.
This volume showcases:
- Applied statistical reasoning\
- Transparent data cleaning and transformation\
- Reproducible modeling and interpretation\
- Clear, professional communication of results\
- Public-facing research using real-world datasets
The goal is not perfection. The goal is transparency.
The chapters follow the structure of the course and mirror the conceptual arc of the main textbook.
- Base R Foundations\
- mtcars Wrangling and Feature Engineering\
- NYPD Shooting Incidents — Cleaning, Insights, and Visualization\
- Merging Data and Comparing Means with t-tests\
- R Script → Quarto Report\
- ANOVA and NYC Camera Violations\
- Midterm: Exercise & Sleep Analysis\
- NBA Analytics — Correlation\
- Florida Crime — Regression Modeling\
- Streaming Analytics — Categorical Analysis\
- Wage Analytics — Logistic Regression\
- Final Project — Civic Data Research\
- Final Assignment — Building a Quarto Book
Each chapter applies a specific methodological framework to a real dataset.
Together, they form a complete analytical portfolio.
Reproducibility is treated as a default practice rather than an optional enhancement.
Each chapter:
- Loads data programmatically\
- Cleans variables transparently\
- Documents transformation decisions\
- Avoids hard-coded statistical results\
- Produces outputs directly from executable code
This ensures analyses are:
- inspectable\
- repeatable\
- defensible
These are the expectations of: - graduate research\
- industry analytics\
- public-facing data work
By assembling these projects into a Quarto book, students move beyond submission-based coursework and into publication-based thinking.
This book serves as:
- A professional portfolio artifact\
- A demonstration of applied statistical competency\
- Evidence of reproducible workflow mastery\
- A foundation for conference submissions and civic engagement
Completing this book is not simply a course requirement. It is a transition point — from student to analytical author.
A previous cohort of students transformed their projects into a public-facing publication:
👉 NYC Open Data Student Gallery - Brooklyn College
This work is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
https://creativecommons.org/licenses/by/4.0/
To build this book on your own machine:
- Clone this repository\
- Open the
.Rprojfile in RStudio\ - Run:
```bash quarto render
