Skip to content

Faster Camera Raw preview, same pixels - #179

Open
Fvzion wants to merge 5 commits into
robbietilton:mainfrom
Fvzion:feature/camera-raw-speed
Open

Fvzion wants to merge 5 commits into
robbietilton:mainfrom
Fvzion:feature/camera-raw-speed

Conversation

@Fvzion

@Fvzion Fvzion commented Sep 29, 2026

Copy link
Copy Markdown

Builds on #178 (the color-noise fix, which the pixel hashes below depend on), so this branch includes its commit; the four commits after it are this PR's. Each commit builds and passes the suite on its own.

Why

Dragging a Camera Raw slider redrew the 2048 px preview in about a second in a Debug build (≈370 ms of it the main C pass on one core, ≈680 ms the histogram and vectorscope counted in unoptimized Swift over every pixel).

What

  1. Pin the kernels' output: FNV hashes of what the light/color, clipping, effects and detail kernels make of a test picture with partial and clear pixels (odd sizes, so no chunk divides evenly).
  2. Main pass on every core: adjust_camera_raw splits rows with dispatch_apply; opaque pixels read their linear value from a 256-entry table built with the loop's own expression. Same arithmetic per pixel, so the hashes don't change. 370 → 49 ms (standalone clang -O3 harness, 2048 × 1365).
  3. Effects and the box blur on every core: the effects' luma and final passes by rows; box_blur_plane's vertical pass in row order with a running sum per column, in bands of 64 columns: the same sums in the same order. 140 → 22 ms. The detail kernel stays single-threaded (its noise loop rewrites luma that neighbours read).
  4. Scopes on a small copy: the histogram and vectorscope count a copy no larger than 512 px, sampled nearest-neighbour so scattered clipped pixels still reach the ends of the histogram.

Whole preview with scopes, Debug build: ≈1030 ms → ≈115 ms.

Tests

  • Pixel hashes unchanged through every commit (also checked by an independent review over 1,350 size/alpha/scale combinations; TSan, ASan and UBSan clean).
  • CameraRawTimingTests (serialized) check the speed, but only when code coverage is off: Xcode's coverage counters are shared by all threads and make parallel code contend, so under coverage (CI's default) they're skipped.

🤖 Generated with Claude Code

Fvzion and others added 5 commits September 29, 2026 17:35
The saturation plane that Color noise reduction blurs was never written
under clear pixels (the luma plane beside it is), so the blur read
whatever that memory held, and the pixels next to a transparent area
came out differently from one run to the next.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hashes of what the light, color, effects and detail kernels make of a
test picture with partial and clear pixels, so a change that is meant to
leave the pixels alone (a faster kernel) can show that it does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rows are split across cores with dispatch_apply, and an opaque pixel's
linear value comes from a 256-entry table built with the loop's own
expression. Every pixel gets the same arithmetic as before, so the
result is identical (the kernel hashes are unchanged); a 2048 x 1365
preview takes about 50 ms instead of 370.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The effects' luma and final passes run by rows, and the box blur's
horizontal pass by rows and its vertical pass by bands of columns, row
after row, each column keeping its own running sum in the order a single
loop would. The pixels are unchanged (hashes, including a 1 x 1 picture,
one row and one column); Texture, Clarity and Dehaze on a 2048 x 1365
preview take about 22 ms instead of 140. The timing checks run in their
own serialized suite, since side by side each gets part of the cores.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
They show the spread of tones and colors, which a copy no larger than
512 px keeps as long as it picks pixels rather than averaging them (so
scattered clipped pixels still reach the ends of the histogram).
Counting every pixel of the preview on each slider step took longer
than grading it. With the kernels on every core, a 2048 x 1365 preview
with its scopes takes about 115 ms instead of about a second in a Debug
build.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant