Skip to content

Integrate real GOES-18 XRS data into existing evaluation pipeline - #36

Merged
dfeen87 merged 2 commits into
mainfrom
copilot/use-real-goes-data-in-experiments
Mar 13, 2026
Merged

dfeen87 merged 2 commits into
mainfrom
copilot/use-real-goes-data-in-experiments

Conversation

Copilot AI commented Mar 13, 2026 •

Copy link
Copy Markdown
Contributor

The pipeline previously had no path to run the existing eval_*.py experiment scripts against the real GOES-18 XRS 1-minute data (2024–2025) already in the repo as noaa_goes18_xrs_1m.csv.zip. This adds a preparation layer that converts that CSV into SWPC-format JSON cache files the existing data loader reads transparently — no changes to loaders or experiment scripts.

Data preparation

shared/prepare_real_data.py — run once before executing any _real script:

  • Reads noaa_goes18_xrs_1m.csv.zip, converting J2000 epoch seconds → UTC via vectorised arithmetic
  • Cleans: deduplicates, enforces uniform 1-min cadence via reindex + interpolate(method="time", limit=60)
  • Writes five SWPC-format JSON cache files per interval to data/raw/goes/<dataset>/<start>_to_<end>.json (gitignored)

Channel mapping from XRS CSV columns:

Pipeline dataset Source
xray_flux longwave_masked (energy "0.1-0.8nm")
xray_background 12-hour rolling median of longwave_masked
magnetometer He normalised longwave_masked → centred at 100 nT ±10 nT
euvs e_low shortwave_masked (EUV proxy)
flare_catalogue empty (no detection applied)

Four intervals from t₀ = 2024-01-01: +30d, +90d, +182d, +365d.

New experiment scripts

Four experiments/eval_<interval>_real.py scripts with fixed 2024 date ranges, writing to results/eval_<interval>_real.json (synthetic results untouched). Also adds the missing experiments/eval_three_month.py (rolling window, date.today() anchor, same pattern as existing scripts).

Julia loader

shared/RealDataLoader.jl — reads the same cache files, returns NamedTuples matching the DataLoader.jl interface.

Bug fix

shared/event_evaluation.compute_lead_times crashed with KeyError when the flare catalogue was empty — pd.DataFrame([]) has no columns, so the subsequent column-order selection failed. Fixed by returning an empty DataFrame with explicit columns when rows is empty.

# Before: KeyError on empty flare catalogue
result = pd.DataFrame(rows)
return result[["onset_time", "lead_time_first_crossing_hours", ...]]

# After
if not rows:
    return pd.DataFrame(columns=["onset_time", "lead_time_first_crossing_hours", ...])

Usage

python shared/prepare_real_data.py          # populate cache (~20 JSON files)
python experiments/eval_one_month_real.py   --n-shuffles 500 --random-state 0
python experiments/eval_three_month_real.py --n-shuffles 500 --random-state 0
python experiments/eval_six_month_real.py   --n-shuffles 500 --random-state 0
python experiments/eval_one_year_real.py    --n-shuffles 500 --random-state 0

Warning

Firewall rules blocked me from connecting to one or more addresses (expand for details)

I tried to connect to the following addresses, but was blocked by firewall rules:

  • services.swpc.noaa.gov
    • Triggering command: /usr/bin/python python -m pytest test/ -x -q (dns block)
    • Triggering command: /usr/bin/python python -m pytest test/ -q k/_temp/copilot-developer-action-main/dist/index.js (dns block)

If you need me to access, download, or install something from one of these locations, you can either:


🔒 GitHub Advanced Security automatically protects Copilot coding agent pull requests. You can protect all pull requests by enabling Advanced Security for your repositories. Learn more about Advanced Security.

…ipts + Julia loader)

Co-authored-by: dfeen87 <158860247+dfeen87@users.noreply.github.com>
Copilot AI changed the title [WIP] Add loading of real GOES‑18 XRS data for experiments Integrate real GOES-18 XRS data into existing evaluation pipeline Mar 13, 2026
Copilot AI requested a review from dfeen87 March 13, 2026 20:28
@dfeen87
dfeen87 marked this pull request as ready for review March 13, 2026 20:32
Copilot AI review requested due to automatic review settings March 13, 2026 20:32
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@dfeen87
dfeen87 merged commit af75d17 into main Mar 13, 2026
6 checks passed
@dfeen87
dfeen87 deleted the copilot/use-real-goes-data-in-experiments branch March 13, 2026 20:32

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds a “real data” path that converts the repo’s GOES-18 XRS 1-minute CSV ZIP into SWPC-range JSON cache files so the existing Python evaluation pipeline can run unchanged, plus a Julia loader for the same caches and a small robustness fix for empty flare catalogues.

Changes:

  • Add shared/prepare_real_data.py to generate SWPC-format JSON cache files under data/raw/goes/<dataset>/<start>_to_<end>.json.
  • Add real-data experiment entrypoints (eval_*_real.py) and the missing rolling-window eval_three_month.py.
  • Add shared/RealDataLoader.jl and fix compute_lead_times to handle empty flare catalogues.

Reviewed changes

Copilot reviewed 9 out of 9 changed files in this pull request and generated 10 comments.

Show a summary per file
File Description
shared/prepare_real_data.py New data-prep script: reads GOES-18 CSV ZIP, cleans/interpolates, writes SWPC-style JSON caches for multiple datasets/intervals.
shared/event_evaluation.py Fix for empty flare catalogue case in compute_lead_times (avoid KeyError).
shared/RealDataLoader.jl Julia loader to read the same long-range cache files and return DataLoader-like NamedTuples.
experiments/eval_one_month_real.py Real-data 30-day evaluation wrapper calling run_interval_eval.
experiments/eval_three_month_real.py Real-data 90-day evaluation wrapper calling run_interval_eval.
experiments/eval_six_month_real.py Real-data 182-day evaluation wrapper calling run_interval_eval.
experiments/eval_one_year_real.py Real-data 365-day evaluation wrapper calling run_interval_eval.
experiments/eval_three_month.py New rolling “most recent 90 days” synthetic evaluation entrypoint.
README.md Document new scripts and the real-data preparation workflow.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

You can also share your feedback on Copilot code review. Take the survey.

df["xrs_long"] = df["xrs_long"].interpolate(method="time", limit=60)
df.dropna(inplace=True)
df = df.reset_index()

Comment on lines +283 to +299
def write_interval_caches(df: pd.DataFrame, name: str,
start: datetime, end: datetime) -> None:
"""Slice *df* to [start, end) and write all five dataset cache files."""
label = _INTERVAL_LABELS.get(name, name)
mask = (df["timestamp"] >= start) & (df["timestamp"] < end)
subset = df.loc[mask].copy().reset_index(drop=True)
print(f"[prepare_real_data] {label}: {len(subset):,} rows "
f"({start.date()} — {end.date()}, excl.)")

for dataset_key, builder in _BUILDERS.items():
records = builder(subset)
path = _cache_path(dataset_key, start, end)
path.parent.mkdir(parents=True, exist_ok=True)
with open(path, "w", encoding="utf-8") as fh:
json.dump(records, fh, separators=(",", ":"))
print(f" ✓ {path.relative_to(_REPO_ROOT)} ({len(records):,} records)")

Comment thread shared/RealDataLoader.jl
Interval strings (start → end, exclusive)
------------------------------------------
1-month : "2024-01-01" → "2024-01-31"
3-month : "2024-01-01" → "2024-04-01"
Background is approximated by a 12-hour (720-sample) rolling median of
the long-wave flux, representing the quiet-Sun baseline level.
"""
bg = df["xrs_long"].rolling(window=720, center=True, min_periods=1).median()
Interval Start End (excl.)
============ =========== ===========
1-month 2024-01-01 2024-01-31
3-month 2024-01-01 2024-04-01
Comment on lines +79 to +89
# Interval definitions (start always = t0 = first 2024 timestamp)
# ---------------------------------------------------------------------------

_T0 = datetime(2024, 1, 1, 0, 0, 0, tzinfo=timezone.utc)

_INTERVALS: dict[str, tuple[datetime, datetime]] = {
"1m": (_T0, _T0 + timedelta(days=30)),
"3m": (_T0, _T0 + timedelta(days=90)),
"6m": (_T0, _T0 + timedelta(days=182)),
"1y": (_T0, _T0 + timedelta(days=365)),
}
=====================================
Run the precursor evaluation pipeline over the real GOES-18 3-month interval.

Uses the fixed date range 2024-01-01 — 2024-04-01 drawn from the real
Comment on lines +163 to +167
if not rows:
return pd.DataFrame(columns=_cols)
result = pd.DataFrame(rows)
# Ensure column order
result = result[
[
"onset_time",
"lead_time_first_crossing_hours",
"lead_time_max_signal_hours",
"first_crossing_time",
"max_signal_time",
]
]
return result
return result[_cols]
Comment thread README.md
Comment on lines +484 to +485
The script derives four aligned intervals from **t₀ = 2024-01-01** (the
earliest timestamp in the file):
-----------------------------
xray_flux ← longwave_masked (0.1–0.8 nm, energy key "0.1-0.8nm")
xray_background ← 12-hour rolling median of longwave_masked
magnetometer He ← normalised longwave_masked scaled to [90, 110] nT range
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants