Conversation
The saturation plane that Color noise reduction blurs was never written under clear pixels (the luma plane beside it is), so the blur read whatever that memory held, and the pixels next to a transparent area came out differently from one run to the next. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hashes of what the light, color, effects and detail kernels make of a test picture with partial and clear pixels, so a change that is meant to leave the pixels alone (a faster kernel) can show that it does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rows are split across cores with dispatch_apply, and an opaque pixel's linear value comes from a 256-entry table built with the loop's own expression. Every pixel gets the same arithmetic as before, so the result is identical (the kernel hashes are unchanged); a 2048 x 1365 preview takes about 50 ms instead of 370. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The effects' luma and final passes run by rows, and the box blur's horizontal pass by rows and its vertical pass by bands of columns, row after row, each column keeping its own running sum in the order a single loop would. The pixels are unchanged (hashes, including a 1 x 1 picture, one row and one column); Texture, Clarity and Dehaze on a 2048 x 1365 preview take about 22 ms instead of 140. The timing checks run in their own serialized suite, since side by side each gets part of the cores. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
They show the spread of tones and colors, which a copy no larger than 512 px keeps as long as it picks pixels rather than averaging them (so scattered clipped pixels still reach the ends of the histogram). Counting every pixel of the preview on each slider step took longer than grading it. With the kernels on every core, a 2048 x 1365 preview with its scopes takes about 115 ms instead of about a second in a Debug build. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Builds on #178 (the color-noise fix, which the pixel hashes below depend on), so this branch includes its commit; the four commits after it are this PR's. Each commit builds and passes the suite on its own.
Why
Dragging a Camera Raw slider redrew the 2048 px preview in about a second in a Debug build (≈370 ms of it the main C pass on one core, ≈680 ms the histogram and vectorscope counted in unoptimized Swift over every pixel).
What
adjust_camera_rawsplits rows withdispatch_apply; opaque pixels read their linear value from a 256-entry table built with the loop's own expression. Same arithmetic per pixel, so the hashes don't change. 370 → 49 ms (standaloneclang -O3harness, 2048 × 1365).box_blur_plane's vertical pass in row order with a running sum per column, in bands of 64 columns: the same sums in the same order. 140 → 22 ms. The detail kernel stays single-threaded (its noise loop rewrites luma that neighbours read).Whole preview with scopes, Debug build: ≈1030 ms → ≈115 ms.
Tests
CameraRawTimingTests(serialized) check the speed, but only when code coverage is off: Xcode's coverage counters are shared by all threads and make parallel code contend, so under coverage (CI's default) they're skipped.🤖 Generated with Claude Code