Skip to content

gpu-compute: stop decoding instructions after stopFetch - #1

Open
Basemism wants to merge 11 commits into
stagingfrom
staging-basem/gpu-fetch-stop
Open

gpu-compute: stop decoding instructions after stopFetch#1
Basemism wants to merge 11 commits into
stagingfrom
staging-basem/gpu-fetch-stop

Conversation

@Basemism

@Basemism Basemism commented Sep 2, 2026

Copy link
Copy Markdown

Fetch has two steps: request/cache instruction bytes, then decode buffered bytes into the wavefront instruction buffer. Wavefront::stopFetch() already treated branches, returns, and s_endpgm as boundaries, but decodeInsts() checked it
only when scheduling future fetch requests. Its inner loop continued consuming bytes already buffered after a boundary.

When s_endpgm and following bytes occupied the same fetched region, gem5 decoded those following bytes under the ending kernel's wavefront. Constructing the following instruction immediately mapped its operands against that kernel's
valid reservation and produced a false panic. The instruction would never have executed: s_endpgm erases later buffered instructions and terminates the wave.

Fix:
Honor Wavefront::stopFetch() before decoding either a split instruction or additional instructions from the fetch buffer. This prevents the decoder from continuing to populate the instruction buffer after the wavefront has entered a state that blocks further fetch and decode work.

v-ramadas and others added 11 commits August 26, 2026 17:31
This commit fixes the instruction execution latency of LDS instructions
based on what microbenchmarks suggest
The LDS uses 4 byte wide banks in most GPUs. This requires each 4B
address to map to a bank. The previous model mapped each byte to a bank
and introduced several bank conflicts. This commit fixes the bank
mapping to map each 4B word to an LDS bank

Change-Id: Iaca0bcd3525a31c4ae55e379bf341a1de3ef8b31
Change-Id: I6dc1c9118027434a9e0bd26cc3843d3f56d955ad
Previously, the LDS model assumed each instruction has the same bus data
transfer costs. This commit updates that to use the instruction dword
length instead

Change-Id: Id3decd6bb6fba3c50c84377ac5c7550660113092
Previously, `GlobalMemPipeline::getNextReadyResp` only checked the
absolute oldest request in the entire ComputeUnit. If this single
request was incomplete, it blocked all subsequent completed requests
from being processed, degrading global memory pipeline performance.

This commit updates the response selection logic to pick the oldest
completed request on a per-wavefront basis.

This prevents a single stalled wavefront from starving the entire CU,
improving overall global memory pipeline throughput.

Change-Id: Ia9bfde40a1940f450aee181ab816cc47dc677644
Introduce separate fabric_clk and memory_clk for the GPUFS,
replacing the single ruby_clock domain.

Change-Id: I817ab2e5a215b76728fd31dcc067535fd6590126
Previously, gpu clock was not applied to the model

Change-Id: Ic3e39f984bc79fc0787282bb73d004a1f474e904
Honor Wavefront::stopFetch() before decoding either a split instruction or
additional instructions from the fetch buffer. This prevents the decoder from
continuing to populate the instruction buffer after the wavefront has entered
a state that blocks further fetch and decode work.
@Basemism
Basemism requested review from TomXia, mattsinc and v-ramadas and a lite review from Copilot September 2, 2026 10:28
@Basemism Basemism self-assigned this Sep 2, 2026

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

The new stopFetch() usage in the while condition can introduce avoidable O(n²) behavior in a hot decode loop due to repeated full scans of the instruction buffer.

Pull request overview

This PR fixes a GPU fetch/decode boundary bug in the shader instruction fetch unit: once a wavefront has buffered a control-flow/end-of-kernel instruction (e.g., s_endpgm), decoding now stops immediately instead of continuing to consume already-fetched bytes and populating the instruction buffer with instructions that should never execute.

Changes:

  • Gate FetchBufDesc::decodeInsts() so it does not decode split or subsequent instructions after Wavefront::stopFetch() becomes true.
  • Prevent decoding of post-s_endpgm bytes that can incorrectly map operands and trigger false panics.
File summaries
File Description
src/gpu-compute/fetch_unit.cc Stops decode from consuming buffered bytes past branch/return/end-of-kernel boundaries by honoring Wavefront::stopFetch() during decode.
Review details

Suppressed comments (1)

src/gpu-compute/fetch_unit.cc:597

  • wavefront->stopFetch() is an O(n) scan over instructionBuffer (see Wavefront::stopFetch()), and calling it in the while condition can make decodeInsts() potentially O(n^2) per buffer fill (one full scan per decoded instruction). You can keep the same correctness behavior while avoiding repeated scans by checking stopFetch() once up-front, then breaking based on the newly-decoded instruction being a branch/return/end-of-kernel.
           hasFetchDataToProcess() && !wavefront->stopFetch()) {
        if (splitDecode()) {
            decodeSplitInst();
        } else {
            TheGpuISA::MachInst mach_inst =
  • Files reviewed: 1/1 changed files
  • Comments generated: 0
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants