Skip to content

configs: model the GPU scalar cache separately from the SQC - #3

Open
Basemism wants to merge 11 commits into
stagingfrom
staging-basem/scalar-cache
Open

configs: model the GPU scalar cache separately from the SQC#3
Basemism wants to merge 11 commits into
stagingfrom
staging-basem/scalar-cache

Conversation

@Basemism

@Basemism Basemism commented Sep 2, 2026

Copy link
Copy Markdown

Create a distinct scalar-cache object and controller instead of reusing the SQC cache configuration for scalar memory requests.

Add size and associativity options while continuing to share the SQC Ruby machine type and controller-version namespace.

Keep existing defaults and route scalar sequencers through the new cache without enabling vector-data requests.

v-ramadas and others added 11 commits August 26, 2026 17:31
This commit fixes the instruction execution latency of LDS instructions
based on what microbenchmarks suggest
The LDS uses 4 byte wide banks in most GPUs. This requires each 4B
address to map to a bank. The previous model mapped each byte to a bank
and introduced several bank conflicts. This commit fixes the bank
mapping to map each 4B word to an LDS bank

Change-Id: Iaca0bcd3525a31c4ae55e379bf341a1de3ef8b31
Change-Id: I6dc1c9118027434a9e0bd26cc3843d3f56d955ad
Previously, the LDS model assumed each instruction has the same bus data
transfer costs. This commit updates that to use the instruction dword
length instead

Change-Id: Id3decd6bb6fba3c50c84377ac5c7550660113092
Previously, `GlobalMemPipeline::getNextReadyResp` only checked the
absolute oldest request in the entire ComputeUnit. If this single
request was incomplete, it blocked all subsequent completed requests
from being processed, degrading global memory pipeline performance.

This commit updates the response selection logic to pick the oldest
completed request on a per-wavefront basis.

This prevents a single stalled wavefront from starving the entire CU,
improving overall global memory pipeline throughput.

Change-Id: Ia9bfde40a1940f450aee181ab816cc47dc677644
Introduce separate fabric_clk and memory_clk for the GPUFS,
replacing the single ruby_clock domain.

Change-Id: I817ab2e5a215b76728fd31dcc067535fd6590126
Previously, gpu clock was not applied to the model

Change-Id: Ic3e39f984bc79fc0787282bb73d004a1f474e904
Create a distinct scalar-cache object and controller instead of reusing the
SQC cache configuration for scalar memory requests. Add independent size and
associativity options while continuing to share the SQC Ruby machine type and
controller-version namespace required by GPU_VIPER.

Keep existing defaults and route scalar sequencers through the new cache
without enabling vector-data requests.
@Basemism
Basemism requested review from TomXia, mattsinc and v-ramadas and a lite review from Copilot September 2, 2026 12:13
@Basemism Basemism self-assigned this Sep 2, 2026

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

The changes are localized to configuration code and appear consistent with the stated goal while preserving defaults, with only minor maintainability feedback.

Pull request overview

This PR updates the GPU_VIPER Ruby configuration to model GPU scalar memory requests with a dedicated scalar L1 cache/controller instead of reusing the existing SQC cache configuration, while preserving existing default behavior.

Changes:

  • Adds ScalarCache and ScalarCntrl to provide a distinct cache object/controller for scalar requests.
  • Introduces --scalar-size and --scalar-assoc CLI options to tune the scalar cache independently.
  • Routes scalar sequencers through the new scalar cache/controller while keeping scalar requests on the existing SQC machine type/versioning namespace.
File summaries
File Description
configs/ruby/GPU_VIPER.py Introduces dedicated scalar cache/controller classes, adds scalar cache sizing options, and switches scalar controller construction to use the new controller.
Review details
  • Files reviewed: 1/1 changed files
  • Comments generated: 1
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread configs/ruby/GPU_VIPER.py
Comment on lines +278 to +289
class ScalarCache(RubyCache):
dataArrayBanks = 8
tagArrayBanks = 8
dataAccessLatency = 1
tagAccessLatency = 1

def create(self, options):
self.size = MemorySize(options.scalar_size)
self.assoc = options.scalar_assoc
if hasattr(options, "sqc_rp"):
self.replacement_policy = ObjectList.rp_list.get(options.sqc_rp)()

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants