Skip to content

Fix real Phase 1-4 perf hotspots found by profiling (~5x on a real session) - #22

Merged
hd152 merged 4 commits into
mainfrom
perf/robust-pca-and-phase4-passes
Sep 22, 2026
Merged

hd152 merged 4 commits into
mainfrom
perf/robust-pca-and-phase4-passes

Conversation

@hd152

@hd152 hd152 commented Sep 22, 2026

Copy link
Copy Markdown
Owner

Summary

Profiling a full pipeline run (not guessing) found several real hotspots across master calibration, DBE, post-processing and Gaussian blur (~30 call sites project-wide). Fixed in stages, each measured before moving to the next — same real profiled session throughout: 3m 7.8s -> 36.9s (~5x).

  • robust_pca_decompose fusion: two new native kernels (robust_pca_pre_svd_input, robust_pca_iterate) fuse the per-iteration IALM elementwise arithmetic that was outside the already-native SVD step; the L-update matmul is rerouted through the existing small_times_wide kernel instead of np.matmul (this machine's numpy has no optimized BLAS). 126.9s -> 45.4s.
  • Calibration timing visibility: _build_masters's timer was gated on real bias/dark/flat frames existing, so --flat-from-lights (which runs precisely when they don't) was invisible, hidden inside an unlabeled "Other (I/O)" bucket. Now timed unconditionally with its own line.
  • --flat-from-lights CFA-aware downsample: block-averages each of the mosaic's 4 Bayer sub-planes independently (never across colours) before decomposition, upsampling the recovered flat back to full resolution after. 45.2s -> 3.6s.
  • DBE compact-source dilation: scipy.ndimage.binary_dilation with a large disk structuring element was pathologically slow on one contiguous bright source (16.2s on a synthetic case matching that shape); replaced with a distance-transform threshold, the same result computed near-linearly.
  • Post-process hot-pixel median filter: was calling scipy directly instead of this project's own native per-channel median kernel — same class of miss already fixed once elsewhere, missed here. 5.1s -> ~0.1s.
  • gaussian_filter_ds's upsample: cubic -> bilinear interpolation (the input just came out of the function's own huge blur, so cubic recovers nothing real). ~3x/call.
  • New native Gaussian blur kernel, swept across every direct scipy.ndimage.gaussian_filter call site that uses a plain 2-D scalar sigma (~15 files) — correlate1d was the single largest self-time item in every profile taken. ~4x faster at small sigma, ~1.3-1.5x at large.

Also included

  • A separate, unrelated fix from the same session: per-frame cosmic-ray rejection's >=20-frame auto-skip is now sub-length-aware (Config.LACOSMIC_LONG_SUB_EXPTIME_S), not just frame-count-aware — keeps per-frame detection on for long-sub sessions where the noise win is real and the star-softening cost is negligible, instead of always deferring to the frame-count skip.
  • Real benchmark refresh: ran tools/bench_vs_siril.py for real against a local Omega Nebula session (Siril CLI installed locally) on this branch's code, found the cosmic-ray fix above now correctly triggers on that session's real 30s subs, and updated README.md/docs/index.html with the real new numbers (113s -> 132s, quieter: noise 0.99-1.19x Siril's -> 0.87-1.03x). Flagged Sunflower/Sculptor as not yet re-measured rather than leaving them silently stale next to freshly-updated Omega figures.
  • A real bug found and fixed while sweeping the Gaussian kernel: the wrapper didn't preserve input dtype, silently upcasting float32 callers to float64. Fixed, with a new test pinning it.
  • One attempted optimisation (finer-grained rayon parallelism in small_times_wide) was measured, found to make no difference (memory-bandwidth bound, not core bound), and reverted rather than shipped.

Test plan

  • Full test suite passes (1647 tests) on the merged result
  • ruff check . and tools/lint_conventions.py clean
  • Every native kernel has bit-exact or documented-tolerance parity tests vs. its numpy/scipy fallback
  • Real end-to-end runs against real Omega Nebula session data (both a synthetic dataset and the user's actual 114-frame session)
  • Two of the newly-wired opt-in paths (--trail-reject, --noise-validate) smoke-tested against real data

astro_native bumped 0.32.0 -> 0.34.0. Full narrative with every measurement in CLAUDE.md and CHANGELOG.md.

🤖 Generated with Claude Code

hd152 and others added 4 commits September 22, 2026 10:40
Profiling a full pipeline run (not guessing) found --flat-from-lights's
synthetic-flat build silently dominating wall time -- 126.9s of a 187.8s
total, hidden in an unlabeled "Other (I/O)" bucket because its timer was
gated on real bias/dark/flat frames existing, which is exactly what this
path runs without.

- Fuse robust_pca_decompose's per-iteration elementwise arithmetic into
  two new native kernels (robust_pca_pre_svd_input, robust_pca_iterate)
  and route the L-update matmul through the existing small_times_wide
  kernel instead of np.matmul (this machine's numpy has no optimized
  BLAS). 126.9s -> 45.4s.
- Time master-calibration building unconditionally, not just when real
  cal frames exist, and give it its own line in the timing summary
  instead of letting it vanish into "Other".
- --flat-from-lights block-averages each of the 4 CFA sub-planes
  independently (never across colours) before decomposition, upsampling
  the recovered flat back to full res after -- vignetting/dust is smooth
  well above the pixel scale this removes, unlike a real dark/bias
  master's per-pixel hot pixels. 45.2s -> 3.6s.
- DBE's compact-source dilation used scipy's binary_dilation with a large
  disk structuring element, pathologically slow on one contiguous bright
  source (16.2s on a synthetic case matching that shape); replaced with a
  distance-transform threshold, the same result computed near-linearly.
- Post-process's hot-pixel step called scipy's median_filter directly
  instead of this project's own native per-channel median kernel -- same
  class of miss already fixed once elsewhere, missed here. 5.1s -> ~0.1s.
- gaussian_filter_ds's large-sigma upsample switched cubic -> bilinear
  interpolation; the input just came out of the function's own huge blur,
  so cubic's extra curvature term recovers nothing real. ~3x/call.

Same real session end to end: 3m 7.8s -> 36.9s (~5x). Tried and reverted:
finer-grained rayon parallelism in small_times_wide measured no
improvement (memory-bandwidth bound, not core bound) -- kept simple.

Every change is bit-exact, measured-equivalent, or covered by a new
native-vs-fallback parity test. New tests in test_native.py,
test_robust_pca.py, test_dbe_gradient.py; tools/bench_robust_pca_scale.py
and expanded tools/bench_native.py entries for the new kernels. astro_native
bumped to 0.33.0. Full narrative in CLAUDE.md and CHANGELOG.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Ran tools/bench_vs_siril.py for real against C:\astro\Omega_Nebula (Siril
is installed locally) on this branch's code. Found the earlier
sub-length-aware cosmic-ray fix now correctly triggers on this session's
real 30s subs, which the previous commit's Phase 4 speedups don't touch:
total time 113s -> 132s (per-frame cosmic-ray detection now runs, at its
documented 39-75% cost), noise 0.99-1.19x Siril's -> 0.87-1.03x (its
documented 11-16% quieter benefit). Net: slower wall clock, quieter and
equally sharp output, on this specific showcase session.

Updated README.md and docs/index.html with the real new Omega numbers
throughout (table, star-width chart, Speed/Noise/Memory prose). Sunflower
and Sculptor's subs (20s, 10s) fall under the threshold and keep the old
default behavior, but their Phase 4 time should still be faster from the
same commit's fixes -- not re-measured (no local data for those sessions),
so both pages say so explicitly rather than leaving them silently stale.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
correlate1d (the C function behind scipy.ndimage.gaussian_filter) was
the single largest self-time item in every full-pipeline profile taken
this session, spread across ~30 call sites project-wide.

gaussian_filter_native is a from-scratch separable reimplementation of
scipy's default mode='reflect' behaviour (two rayon-parallel passes,
same reflect_idx boundary convention the hot-pixel kernels already use).
Not a port, so parity is numerical rather than bit-exact: verified to
double-precision rounding (rtol/atol 1e-9) across sigma 0.8-96 and two
shapes. ~4x faster at small sigma, ~1.3-1.5x at the large sigmas
gaussian_filter_ds's downsampled branch uses, measured in isolation.

Wired via a new _gaussian_blur wrapper (2-D float only, transparent
scipy fallback otherwise) into gaussian_filter_ds's both branches and
into background.py/denoising.py's highest-traffic direct callers
(reduce_chroma_noise, multiscale_local_contrast,
_structure_tensor_coherence, DBE's mesh-fallback surface fit).

Deliberately not a full sweep: registration.py, quality.py,
moving_objects.py, noise_validation.py and a few others still call
scipy directly, each with its own sigma/shape assumptions worth
checking before switching over. A 3-D per-axis-sigma call in
reduce_stars (sigma=(s,s,0)) was left alone on purpose -- the kernel
only supports scalar sigma on a 2-D plane.

Measured on a controlled 15-frame real-session subset, same machine,
back to back: 46.0s -> 44.7s. The full 114-frame real benchmark run
showed no reliable signal either way across repeated runs (system-level
timing noise on this machine outweighed the effect at that scale,
confirmed via idle CPU load between runs) -- not used as evidence here.
astro_native bumped to 0.34.0.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Follow-up to the previous commit, which wired gaussian_filter_native
into only the highest-traffic callers and flagged the rest for later.
Swept every remaining direct scipy.ndimage.gaussian_filter call that
uses a plain scalar sigma on a 2-D array: cli.py (master bias/dark/flat
smoothing), exposure_fusion.py, merge.py, moving_objects.py,
noise_validation.py, postprocess.py, quality.py, registration.py,
star_removal.py, trail_reject.py.

Left alone on purpose: call sites passing a per-axis sigma tuple
(denoising.py::reduce_stars's sigma=(s,s,0), exposure_fusion.py's
pyramid up/downsample, originvision_infer.py's resize prefilter with
its own tight numeric tolerances against a reference implementation)
-- the kernel only supports a scalar sigma on a 2-D plane. These
already fall back to scipy safely on their own (a tuple sigma raises
inside the wrapper's float(sigma) cast, caught by the existing
try/except), so nothing needed changing there.

Also fixed a real bug in _gaussian_blur found while auditing the
sweep: it didn't preserve the caller's dtype, so a float32 array
(several master-calibration and Phase 4 call sites pass float32)
silently came back as float64, doubling memory for no reason. The
native kernel is always float64 internally like every other kernel in
this file; the wrapper now casts the result back to the input's dtype,
matching scipy's own contract. New dtype-parity test in test_native.py
catches this going forward.

Full test suite (1647 tests) passes; smoke-tested --trail-reject and
--noise-validate (two of the newly-wired opt-in paths) against a real
15-frame session subset with no crash.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant