Skip to content

arch-amdgpu: support ACC-selected DS data operands - #7

Open
Basemism wants to merge 11 commits into
stagingfrom
staging-basem/cdna3-ds-acc
Open

arch-amdgpu: support ACC-selected DS data operands#7
Basemism wants to merge 11 commits into
stagingfrom
staging-basem/cdna3-ds-acc

Conversation

@Basemism

@Basemism Basemism commented Sep 2, 2026

Copy link
Copy Markdown

Add support for the CDNA DS ACC encoding, which selects accumulator registers for DS instruction data operands.

Before, DS instructions emitted with ACC=1 are either rejected by the decoder or incorrectly read and write ordinary VGPRs instead of the wavefront’s AGPR allocation.

The DS encoding now decodes the former reserved bit as ACC.

When ACC=1:

  • DS source data operands are mapped through the AGPR window of the unified vector register file.
  • DS destination data operands are mapped through the same AGPR window.
  • The DS address operand remains an ordinary VGPR.
  • wf->accumOffset is added to data-register indices during both physical-register mapping and instruction execution.

Applying the offset in both paths ensures that dependency tracking and the instruction data path refer to the same physical registers.

The change covers 32-, 64-, 96-, and 128-bit DS reads and writes, including multi-data-register instructions.

v-ramadas and others added 11 commits August 26, 2026 17:31
This commit fixes the instruction execution latency of LDS instructions
based on what microbenchmarks suggest
The LDS uses 4 byte wide banks in most GPUs. This requires each 4B
address to map to a bank. The previous model mapped each byte to a bank
and introduced several bank conflicts. This commit fixes the bank
mapping to map each 4B word to an LDS bank

Change-Id: Iaca0bcd3525a31c4ae55e379bf341a1de3ef8b31
Change-Id: I6dc1c9118027434a9e0bd26cc3843d3f56d955ad
Previously, the LDS model assumed each instruction has the same bus data
transfer costs. This commit updates that to use the instruction dword
length instead

Change-Id: Id3decd6bb6fba3c50c84377ac5c7550660113092
Previously, `GlobalMemPipeline::getNextReadyResp` only checked the
absolute oldest request in the entire ComputeUnit. If this single
request was incomplete, it blocked all subsequent completed requests
from being processed, degrading global memory pipeline performance.

This commit updates the response selection logic to pick the oldest
completed request on a per-wavefront basis.

This prevents a single stalled wavefront from starving the entire CU,
improving overall global memory pipeline throughput.

Change-Id: Ia9bfde40a1940f450aee181ab816cc47dc677644
Introduce separate fabric_clk and memory_clk for the GPUFS,
replacing the single ruby_clock domain.

Change-Id: I817ab2e5a215b76728fd31dcc067535fd6590126
Previously, gpu clock was not applied to the model

Change-Id: Ic3e39f984bc79fc0787282bb73d004a1f474e904
Decode the CDNA DS ACC bit and map DS data operands, but not address operands,
through the AGPR window of the unified vector register file. Apply the same
accumulator offset during operand mapping and instruction execution for
32-, 64-, 96-, and 128-bit DS reads and writes.

Also accept all DS primary-encoding variants used by the ACC form and check
that mapped AGPR operands remain inside the wavefront's reserved vector
register allocation.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants