Fix real Phase 1-4 perf hotspots found by profiling (~5x on a real session) - #22
Merged
Merged
Conversation
Profiling a full pipeline run (not guessing) found --flat-from-lights's synthetic-flat build silently dominating wall time -- 126.9s of a 187.8s total, hidden in an unlabeled "Other (I/O)" bucket because its timer was gated on real bias/dark/flat frames existing, which is exactly what this path runs without. - Fuse robust_pca_decompose's per-iteration elementwise arithmetic into two new native kernels (robust_pca_pre_svd_input, robust_pca_iterate) and route the L-update matmul through the existing small_times_wide kernel instead of np.matmul (this machine's numpy has no optimized BLAS). 126.9s -> 45.4s. - Time master-calibration building unconditionally, not just when real cal frames exist, and give it its own line in the timing summary instead of letting it vanish into "Other". - --flat-from-lights block-averages each of the 4 CFA sub-planes independently (never across colours) before decomposition, upsampling the recovered flat back to full res after -- vignetting/dust is smooth well above the pixel scale this removes, unlike a real dark/bias master's per-pixel hot pixels. 45.2s -> 3.6s. - DBE's compact-source dilation used scipy's binary_dilation with a large disk structuring element, pathologically slow on one contiguous bright source (16.2s on a synthetic case matching that shape); replaced with a distance-transform threshold, the same result computed near-linearly. - Post-process's hot-pixel step called scipy's median_filter directly instead of this project's own native per-channel median kernel -- same class of miss already fixed once elsewhere, missed here. 5.1s -> ~0.1s. - gaussian_filter_ds's large-sigma upsample switched cubic -> bilinear interpolation; the input just came out of the function's own huge blur, so cubic's extra curvature term recovers nothing real. ~3x/call. Same real session end to end: 3m 7.8s -> 36.9s (~5x). Tried and reverted: finer-grained rayon parallelism in small_times_wide measured no improvement (memory-bandwidth bound, not core bound) -- kept simple. Every change is bit-exact, measured-equivalent, or covered by a new native-vs-fallback parity test. New tests in test_native.py, test_robust_pca.py, test_dbe_gradient.py; tools/bench_robust_pca_scale.py and expanded tools/bench_native.py entries for the new kernels. astro_native bumped to 0.33.0. Full narrative in CLAUDE.md and CHANGELOG.md. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Ran tools/bench_vs_siril.py for real against C:\astro\Omega_Nebula (Siril is installed locally) on this branch's code. Found the earlier sub-length-aware cosmic-ray fix now correctly triggers on this session's real 30s subs, which the previous commit's Phase 4 speedups don't touch: total time 113s -> 132s (per-frame cosmic-ray detection now runs, at its documented 39-75% cost), noise 0.99-1.19x Siril's -> 0.87-1.03x (its documented 11-16% quieter benefit). Net: slower wall clock, quieter and equally sharp output, on this specific showcase session. Updated README.md and docs/index.html with the real new Omega numbers throughout (table, star-width chart, Speed/Noise/Memory prose). Sunflower and Sculptor's subs (20s, 10s) fall under the threshold and keep the old default behavior, but their Phase 4 time should still be faster from the same commit's fixes -- not re-measured (no local data for those sessions), so both pages say so explicitly rather than leaving them silently stale. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
correlate1d (the C function behind scipy.ndimage.gaussian_filter) was the single largest self-time item in every full-pipeline profile taken this session, spread across ~30 call sites project-wide. gaussian_filter_native is a from-scratch separable reimplementation of scipy's default mode='reflect' behaviour (two rayon-parallel passes, same reflect_idx boundary convention the hot-pixel kernels already use). Not a port, so parity is numerical rather than bit-exact: verified to double-precision rounding (rtol/atol 1e-9) across sigma 0.8-96 and two shapes. ~4x faster at small sigma, ~1.3-1.5x at the large sigmas gaussian_filter_ds's downsampled branch uses, measured in isolation. Wired via a new _gaussian_blur wrapper (2-D float only, transparent scipy fallback otherwise) into gaussian_filter_ds's both branches and into background.py/denoising.py's highest-traffic direct callers (reduce_chroma_noise, multiscale_local_contrast, _structure_tensor_coherence, DBE's mesh-fallback surface fit). Deliberately not a full sweep: registration.py, quality.py, moving_objects.py, noise_validation.py and a few others still call scipy directly, each with its own sigma/shape assumptions worth checking before switching over. A 3-D per-axis-sigma call in reduce_stars (sigma=(s,s,0)) was left alone on purpose -- the kernel only supports scalar sigma on a 2-D plane. Measured on a controlled 15-frame real-session subset, same machine, back to back: 46.0s -> 44.7s. The full 114-frame real benchmark run showed no reliable signal either way across repeated runs (system-level timing noise on this machine outweighed the effect at that scale, confirmed via idle CPU load between runs) -- not used as evidence here. astro_native bumped to 0.34.0. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Follow-up to the previous commit, which wired gaussian_filter_native into only the highest-traffic callers and flagged the rest for later. Swept every remaining direct scipy.ndimage.gaussian_filter call that uses a plain scalar sigma on a 2-D array: cli.py (master bias/dark/flat smoothing), exposure_fusion.py, merge.py, moving_objects.py, noise_validation.py, postprocess.py, quality.py, registration.py, star_removal.py, trail_reject.py. Left alone on purpose: call sites passing a per-axis sigma tuple (denoising.py::reduce_stars's sigma=(s,s,0), exposure_fusion.py's pyramid up/downsample, originvision_infer.py's resize prefilter with its own tight numeric tolerances against a reference implementation) -- the kernel only supports a scalar sigma on a 2-D plane. These already fall back to scipy safely on their own (a tuple sigma raises inside the wrapper's float(sigma) cast, caught by the existing try/except), so nothing needed changing there. Also fixed a real bug in _gaussian_blur found while auditing the sweep: it didn't preserve the caller's dtype, so a float32 array (several master-calibration and Phase 4 call sites pass float32) silently came back as float64, doubling memory for no reason. The native kernel is always float64 internally like every other kernel in this file; the wrapper now casts the result back to the input's dtype, matching scipy's own contract. New dtype-parity test in test_native.py catches this going forward. Full test suite (1647 tests) passes; smoke-tested --trail-reject and --noise-validate (two of the newly-wired opt-in paths) against a real 15-frame session subset with no crash. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This was referenced Sep 22, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Profiling a full pipeline run (not guessing) found several real hotspots across master calibration, DBE, post-processing and Gaussian blur (~30 call sites project-wide). Fixed in stages, each measured before moving to the next — same real profiled session throughout: 3m 7.8s -> 36.9s (~5x).
robust_pca_pre_svd_input,robust_pca_iterate) fuse the per-iteration IALM elementwise arithmetic that was outside the already-native SVD step; the L-update matmul is rerouted through the existingsmall_times_widekernel instead ofnp.matmul(this machine's numpy has no optimized BLAS). 126.9s -> 45.4s._build_masters's timer was gated on real bias/dark/flat frames existing, so--flat-from-lights(which runs precisely when they don't) was invisible, hidden inside an unlabeled "Other (I/O)" bucket. Now timed unconditionally with its own line.--flat-from-lightsCFA-aware downsample: block-averages each of the mosaic's 4 Bayer sub-planes independently (never across colours) before decomposition, upsampling the recovered flat back to full resolution after. 45.2s -> 3.6s.scipy.ndimage.binary_dilationwith a large disk structuring element was pathologically slow on one contiguous bright source (16.2s on a synthetic case matching that shape); replaced with a distance-transform threshold, the same result computed near-linearly.gaussian_filter_ds's upsample: cubic -> bilinear interpolation (the input just came out of the function's own huge blur, so cubic recovers nothing real). ~3x/call.scipy.ndimage.gaussian_filtercall site that uses a plain 2-D scalar sigma (~15 files) —correlate1dwas the single largest self-time item in every profile taken. ~4x faster at small sigma, ~1.3-1.5x at large.Also included
>=20-frame auto-skip is now sub-length-aware (Config.LACOSMIC_LONG_SUB_EXPTIME_S), not just frame-count-aware — keeps per-frame detection on for long-sub sessions where the noise win is real and the star-softening cost is negligible, instead of always deferring to the frame-count skip.tools/bench_vs_siril.pyfor real against a local Omega Nebula session (Siril CLI installed locally) on this branch's code, found the cosmic-ray fix above now correctly triggers on that session's real 30s subs, and updated README.md/docs/index.html with the real new numbers (113s -> 132s, quieter: noise 0.99-1.19x Siril's -> 0.87-1.03x). Flagged Sunflower/Sculptor as not yet re-measured rather than leaving them silently stale next to freshly-updated Omega figures.small_times_wide) was measured, found to make no difference (memory-bandwidth bound, not core bound), and reverted rather than shipped.Test plan
ruff check .andtools/lint_conventions.pyclean--trail-reject,--noise-validate) smoke-tested against real dataastro_native bumped 0.32.0 -> 0.34.0. Full narrative with every measurement in CLAUDE.md and CHANGELOG.md.
🤖 Generated with Claude Code