Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
72 changes: 72 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,78 @@ match the `VERSION` file and `v*` git tags.

## [Unreleased]

### Changed

- **A native Gaussian blur kernel, swept across essentially every 2-D scalar-sigma call site.** `correlate1d`
(the C function behind `scipy.ndimage.gaussian_filter`) was the single largest self-time item in every
full-pipeline profile taken this session, spread across ~30 call sites project-wide. Added
`gaussian_filter_native` (from-scratch separable reimplementation, `mode='reflect'`, verified against scipy
to double-precision rounding) behind a `_gaussian_blur` wrapper (`src/background.py`) that preserves the
caller's dtype (float32 stays float32 -- a first version of the wrapper silently upcast to float64, caught
by a new dtype-parity test before it could inflate memory on every caller) and falls back to scipy for
anything the kernel doesn't cover. ~4x faster at small sigma, ~1.3-1.5x at large. Wired into every direct
`gaussian_filter` call site that uses a plain scalar sigma on a 2-D array: `background.py`,
`denoising.py`, `cli.py` (master bias/dark/flat smoothing), `exposure_fusion.py`, `merge.py`,
`moving_objects.py`, `noise_validation.py`, `postprocess.py`, `quality.py`, `registration.py`,
`star_removal.py`, `trail_reject.py`. Deliberately left alone: the handful of call sites that pass a
per-axis sigma tuple (e.g. `(sigma, sigma, 0)` to blur spatial axes only, or `--fix-atmospheric-dispersion`
-adjacent inference preprocessing with its own tight numeric tolerances against a reference implementation)
-- the kernel only supports a scalar sigma on a 2-D plane, and those calls already fall back to scipy
safely on their own (a tuple sigma raises inside the wrapper's `float(sigma)` cast, caught and handled).

- **Several real perf fixes, found by profiling a full run rather than guessing.** A synthetic-flat build
(`--flat-from-lights`, auto-triggered whenever no flat frames exist) was silently the single biggest cost
on a real profiled session -- 126.9s of a 187.8s total, hidden inside an unlabeled "Other (I/O)" bucket in
the timing summary. Fixed in stages, each measured before moving to the next:
- `robust_pca_decompose`'s per-iteration numpy arithmetic (soft-threshold, residual, `Y` update, norm) was
~12 full-array passes/iteration outside the (already-native) SVD step; fused into two native kernels
(`robust_pca_pre_svd_input`, `robust_pca_iterate`), plus routing the L-update's matrix reconstruction
through the existing `small_times_wide` kernel instead of a `np.matmul` call that hit this machine's
unoptimized reference BLAS. 126.9s -> 45.4s.
- The timing summary's "Other (I/O)" bucket was hiding master-calibration building specifically because its
timer was gated on real bias/dark/flat frames existing -- `--flat-from-lights` runs precisely when they
don't, so it was never timed at all. Now timed unconditionally as its own "Calibration" line.
- `--flat-from-lights` now block-averages each of the mosaic's 4 CFA sub-planes independently (never across
colours) by 4x before decomposition, upsampling the recovered flat back to full resolution after --
a flat's vignetting/dust signal is smooth well above the pixel scale this removes, unlike a real dark or
bias master's per-pixel hot pixels, which is why this is `--flat-from-lights`-only. 45.2s -> 3.6s.
- DBE's compact-source mask dilation (`_build_emission_mask`) called `scipy.ndimage.binary_dilation` with a
large disk structuring element, pathologically slow on one large contiguous bright source (a real galaxy
or comet core): 16.2s on a synthetic case matching that shape. Replaced with a distance-transform
threshold -- the exact same result (verified bit-exact), computed in near-linear time instead.
- Post-processing's first step (per-channel hot-pixel removal) called `scipy.ndimage.median_filter` directly
instead of this project's own native per-channel median kernel -- the same class of miss already fixed
once elsewhere in the codebase, just never caught here. 5.1s -> ~0.1s (3 native calls, one per channel).
- `gaussian_filter_ds`'s large-sigma upsample (used throughout Phase 4: chroma denoising, local contrast,
DBE, and more) used cubic interpolation on a grid that had just come out of the function's own huge
Gaussian blur, where cubic's extra curvature term recovers nothing real; switched to bilinear, ~3x
faster per call, measured difference negligible.

End to end, the same profiled real session: 3m 7.8s -> 36.9s (~5x). Every change is either bit-exact,
measured-equivalent (documented tolerance), or covered by a new native-vs-fallback parity test; one
attempted optimisation (finer-grained parallelism in `small_times_wide`) was measured, found to make no
difference, and reverted rather than shipped. Full detail, including what was tried and didn't help, in
[CLAUDE.md](CLAUDE.md).

- **A separate, unrelated fix from the same session moved the Omega benchmark numbers too.** The first change
of the session made per-frame cosmic-ray rejection's >=20-frame auto-skip sub-length-aware (see below) --
Omega's real 30s subs cross that new threshold, so a `tools/bench_vs_siril.py` re-run on the real session
found the README/website comparison table stale: total time 113s -> 132s (per-frame detection now runs,
same documented 39-75% cost), noise 0.99-1.19x Siril's -> 0.87-1.03x (same documented 11-16% quieter
benefit). Sunflower and Sculptor's subs (20s, 10s) are under the threshold and unaffected by it, but their
Phase 4 time should improve too from the fixes above -- not yet re-measured, flagged as such on both pages
rather than left silently stale. README.md and docs/index.html updated with the real Omega numbers.

- **Per-frame cosmic-ray rejection's auto-skip is now sub-length-aware, not just frame-count-aware.**
Previously any >=20-frame session with a rejection-based stack method skipped per-frame L.A.Cosmic
entirely, on the theory that stack-level sigma-clip catches cosmic rays just as well. True on average, but
the noise-vs-softening tradeoff of forcing it back on (documented in the 2.2.4 entry below) tracks sub
length: ~0% star softening at 30s subs, +6.8% at 20s, +11.8% at 10s, while the noise win (11-16%) holds
across all three. `Config.LACOSMIC_LONG_SUB_EXPTIME_S` (25s) now keeps per-frame detection on even at
>=20 frames when the session's median sub exposure is long enough that the softening cost is negligible,
instead of always deferring to the frame-count skip. Only three real sessions inform the exact threshold,
so it's a starting heuristic, not a calibrated cutoff.

### Added

- **Self-update check.** The CLI prints one line at the end of a run, and the desktop app shows a small
Expand Down
8 changes: 5 additions & 3 deletions CLAUDE.md

Large diffs are not rendered by default.

12 changes: 7 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -815,16 +815,18 @@ OriginStack first, Siril second. "Stack only" leaves out OriginStack's finishing

| Session | Stack only | OriginStack, finished image | Peak memory | Star width |
|---------|-----------|-----------------------------|-------------|------------|
| Omega, 114 frames | 57 s / 72 s | 110 s | 16.5 / 8.1 GB | 4.40 / 4.56 px |
| Omega, 114 frames | 102 s / 63 s | 132 s | 18.4 / 8.1 GB | 4.39 / 4.56 px |
| Sunflower, 158 frames | 106 s / 81 s | 159 s | 16.9 / 8.7 GB | 3.56 / 3.60 px |
| Sculptor, 532 frames | 258 s / 150 s | 270 s | 16.7 / 11.2 GB | 2.78 / 2.95 px |

- **Sharpness: OriginStack is as sharp or sharper.** Its stars are 3% narrower than Siril's on Omega, level on Sunflower and 6% narrower on Sculptor. It used every frame (114/114, 158/158 and 532/532); Siril used 111, 157 and 525.
- **Speed depends on the session.** OriginStack stacks Omega faster and the two larger sessions slower (1.3× as long on Sunflower, 1.7× on Sculptor). Timings move by up to 15% between runs of the same code: Siril's Omega ranged from 53 to 72 s across our runs, OriginStack's Sunflower from 83 to 106 s.
- **Noise: Siril is still cleaner.** OriginStack's per-pixel noise (after matching the flux scale per channel) is 0.99–1.19× Siril's on Omega, 1.09–1.20× on Sunflower and 1.05–1.33× on Sculptor: about level in green, up to a third higher in red and blue. Part of that is the demosaicing filter: after calibration and debayering, a single OriginStack frame is about 9–10% noisier than Siril's in green and blue and about 11% quieter in red (Malvar-He-Cutler against Siril's RCD). The rest traces to per-frame cosmic-ray/hot-pixel spikes: Malvar's debayer smears a single-pixel spike over more pixels in red/blue than in green, and the pipeline skips per-frame cosmic-ray detection once a session is large enough to rely on stack-level rejection instead — which runs after debayering, on the already-smeared pixels, and doesn't fully make up for it. Forcing `--cosmic-ray-rejection` back on closes most of the gap (measured 11–16% quieter across all three sessions) but softens stars slightly on shorter sub-exposures and adds 39–75% to the stacking time, so it stays opt-in rather than the default — see the changelog for the numbers.
- **OriginStack's memory does not grow with the session.** It stayed at 16.5–16.9 GB from 114 to 532 frames, set by the worker count (`-j` lowers it), while Siril's grew from 8.1 to 11.2 GB. Siril uses less at these sizes. Peak temporary disk was about the same for both (18–59 GB).
- **Speed depends on the session, and now on your sub length.** Omega's stack-only time is now 102 s against Siril's 63 s (1.6× as long) — slower than an earlier version of this section reported, because Omega's 30 s subs now cross this version's threshold for keeping per-frame cosmic-ray detection on through a big session (previously skipped by default above ~20 frames); that's most of the extra time, and it's also what closes most of the noise gap below. Sunflower and Sculptor's shorter subs (20 s, 10 s) fall under that threshold and keep the old default, but this version's finishing step (Phase 4) is roughly twice as fast for every session regardless of sub length — their figures below are from the previous measurement and have not been re-run since this pass. Timings move by up to 15% between runs of the same code: Siril's Omega ranged from 53 to 72 s across our runs, 63 s the latest; OriginStack's Sunflower from 83 to 106 s.
- **Noise: closer than it was, on sessions long enough to earn it.** On Omega, OriginStack's per-pixel noise is now 0.87–1.03× Siril's (was 0.99–1.19×) — the same 30 s-sub cosmic-ray fix above, not a separate change. Sunflower and Sculptor's subs are too short to trigger it by default, so their figures (1.09–1.20× on Sunflower, 1.05–1.33× on Sculptor) stand from the earlier measurement, about level in green, up to a third higher in red and blue. Part of the underlying gap is the demosaicing filter itself: after calibration and debayering, a single OriginStack frame is about 9–10% noisier than Siril's in green and blue and about 11% quieter in red (Malvar-He-Cutler against Siril's RCD). The rest traces to per-frame cosmic-ray/hot-pixel spikes: Malvar's debayer smears a single-pixel spike over more pixels in red/blue than in green, and stack-level rejection alone runs after debayering, on the already-smeared pixels, so it doesn't fully make up for it — per-frame detection catches it earlier, at the time cost above, which is why it's a threshold (`--cosmic-ray-rejection` / `--no-cosmic-ray-rejection` override it either way) rather than always on. See the changelog for the full numbers.
- **OriginStack's memory does not grow with the session.** It ranged 16.7–18.4 GB across the three sessions (Omega's cosmic-ray pass adds some of that), set by the worker count (`-j` lowers it), while Siril's grew from 8.1 to 11.2 GB. Siril uses less at these sizes. Peak temporary disk was about the same for both (18–59 GB).
- **DeepSkyStacker** took 14 min 8 s on Omega (defaults: plain average, no rejection), gave stars softer than both others (4.83 px against 4.56 for Siril on the same stars), and its stack showed horizontal hot-pixel streaks and a few misregistered frames. A tuned run would look better.
- **Where OriginStack fits.** It finishes the image in the same run with no setup, explains the night ([diagnostics](docs/advanced.html)), runs fully offline on request (`--offline`) and does much that the other tools leave to you. On a plain stack it is now competitive on sharpness and behind Siril on noise and, for larger sessions, on time.
- **Where OriginStack fits.** It finishes the image in the same run with no setup, explains the night ([diagnostics](docs/advanced.html)), runs fully offline on request (`--offline`) and does much that the other tools leave to you. On a plain stack it is now competitive on sharpness, close to Siril on noise on longer subs where per-frame cosmic-ray detection kicks in, further behind on shorter subs where it doesn't, and behind Siril on time.

Omega's figures above were refreshed on 2026-09-22 after a performance pass (Phase 4 roughly twice as fast; a new threshold that keeps per-frame cosmic-ray detection on through long sessions with 25 s+ subs, slower but quieter). Sunflower and Sculptor still reflect the prior measurement and are due for a re-run.

How star width is measured matters more than it looks. It is a Gaussian fit to each of up to 400 bright, unsaturated, isolated stars, **at the same positions in both stacks and on each stack's own pixel grid** (no resampling), and the median is reported. An earlier version of this section compared each stack's FWHM over its own detected star list and concluded OriginStack's stars were tighter; that came from which stars were picked and was withdrawn. Measured properly, OriginStack's stars were then genuinely 13% wider than Siril's on Omega, until a hot-pixel filter that was clipping the cores of bright stars was found, by switching Phase 1 steps off one at a time, and fixed (see the changelog). Registration accuracy (about 0.1 px between frames), the warp kernel, rejection, weighting and normalisation were all checked and are not the cause of anything above.

Expand Down
Loading
Loading