gpu-compute: stop decoding instructions after stopFetch - #1
Open
Basemism wants to merge 11 commits into
Open
Conversation
This commit fixes the instruction execution latency of LDS instructions based on what microbenchmarks suggest
The LDS uses 4 byte wide banks in most GPUs. This requires each 4B address to map to a bank. The previous model mapped each byte to a bank and introduced several bank conflicts. This commit fixes the bank mapping to map each 4B word to an LDS bank Change-Id: Iaca0bcd3525a31c4ae55e379bf341a1de3ef8b31
Change-Id: I6dc1c9118027434a9e0bd26cc3843d3f56d955ad
Previously, the LDS model assumed each instruction has the same bus data transfer costs. This commit updates that to use the instruction dword length instead Change-Id: Id3decd6bb6fba3c50c84377ac5c7550660113092
Previously, `GlobalMemPipeline::getNextReadyResp` only checked the absolute oldest request in the entire ComputeUnit. If this single request was incomplete, it blocked all subsequent completed requests from being processed, degrading global memory pipeline performance. This commit updates the response selection logic to pick the oldest completed request on a per-wavefront basis. This prevents a single stalled wavefront from starving the entire CU, improving overall global memory pipeline throughput. Change-Id: Ia9bfde40a1940f450aee181ab816cc47dc677644
Introduce separate fabric_clk and memory_clk for the GPUFS, replacing the single ruby_clock domain. Change-Id: I817ab2e5a215b76728fd31dcc067535fd6590126
Previously, gpu clock was not applied to the model Change-Id: Ic3e39f984bc79fc0787282bb73d004a1f474e904
Honor Wavefront::stopFetch() before decoding either a split instruction or additional instructions from the fetch buffer. This prevents the decoder from continuing to populate the instruction buffer after the wavefront has entered a state that blocks further fetch and decode work.
There was a problem hiding this comment.
🔵 Needs a closer look
The new stopFetch() usage in the while condition can introduce avoidable O(n²) behavior in a hot decode loop due to repeated full scans of the instruction buffer.
Pull request overview
This PR fixes a GPU fetch/decode boundary bug in the shader instruction fetch unit: once a wavefront has buffered a control-flow/end-of-kernel instruction (e.g., s_endpgm), decoding now stops immediately instead of continuing to consume already-fetched bytes and populating the instruction buffer with instructions that should never execute.
Changes:
- Gate
FetchBufDesc::decodeInsts()so it does not decode split or subsequent instructions afterWavefront::stopFetch()becomes true. - Prevent decoding of post-
s_endpgmbytes that can incorrectly map operands and trigger false panics.
File summaries
| File | Description |
|---|---|
src/gpu-compute/fetch_unit.cc |
Stops decode from consuming buffered bytes past branch/return/end-of-kernel boundaries by honoring Wavefront::stopFetch() during decode. |
Review details
Suppressed comments (1)
src/gpu-compute/fetch_unit.cc:597
wavefront->stopFetch()is an O(n) scan overinstructionBuffer(seeWavefront::stopFetch()), and calling it in thewhilecondition can makedecodeInsts()potentially O(n^2) per buffer fill (one full scan per decoded instruction). You can keep the same correctness behavior while avoiding repeated scans by checkingstopFetch()once up-front, then breaking based on the newly-decoded instruction being a branch/return/end-of-kernel.
hasFetchDataToProcess() && !wavefront->stopFetch()) {
if (splitDecode()) {
decodeSplitInst();
} else {
TheGpuISA::MachInst mach_inst =
- Files reviewed: 1/1 changed files
- Comments generated: 0
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fetch has two steps: request/cache instruction bytes, then decode buffered bytes into the wavefront instruction buffer.
Wavefront::stopFetch()already treated branches, returns, ands_endpgmas boundaries, butdecodeInsts()checked itonly when scheduling future fetch requests. Its inner loop continued consuming bytes already buffered after a boundary.
When
s_endpgmand following bytes occupied the same fetched region, gem5 decoded those following bytes under the ending kernel's wavefront. Constructing the following instruction immediately mapped its operands against that kernel'svalid reservation and produced a false panic. The instruction would never have executed:
s_endpgmerases later buffered instructions and terminates the wave.Fix:
Honor
Wavefront::stopFetch()before decoding either a split instruction or additional instructions from the fetch buffer. This prevents the decoder from continuing to populate the instruction buffer after the wavefront has entered a state that blocks further fetch and decode work.