Skip to content

About

Deterministic Go simulator for congestion-aware routing in GPU training fabrics

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

GPU Fabric Congestion Simulator

CI Go License

A deterministic discrete-event simulator for congestion and routing in multi-hop GPU training fabrics. The project models a two-stage Clos network, replays all-reduce, mixture-of-experts all-to-all, and incast workloads, and compares static ECMP against telemetry-driven adaptive flowlet routing.

Completed adaptive-routing run

What it models

  • Directional links with configurable bandwidth, propagation delay, and finite output buffers
  • Ring all-reduce phases, MoE all-to-all token exchange, many-to-one incast, and background traffic
  • Stable per-flow ECMP path selection
  • Adaptive routing using delayed queue, utilization, and drop telemetry
  • Flowlet boundaries, route cooldowns, and hysteresis
  • Reliable chunk retries after buffer drops
  • p50/p95/p99 flow completion, collective completion, queue depth, utilization, retry volume, route changes, and processed events
  • Deterministic replay for a fixed scenario and seed

The simulation clock advances through a priority queue of logical events. It does not use wall-clock sleeps or one goroutine per packet, so operating-system scheduling does not change a run's result.

Results

The checked-in congested.json fixture produces the following deterministic comparison:

Metric ECMP Adaptive Change
Collective completion 8.61 ms 5.17 ms -39.97%
Flow p99 286.9 us 172.2 us -39.97%
Retried bytes 3.36 GB 1.56 GB -53.42%
Route changes 0 480 -

These are simulator measurements for a synthetic fixture, not hardware benchmarks. The mixed-background scenario adds seed-stable start jitter for paired experiments with traffic variation.

At its checked-in seed (53), the moe-all-to-all.json fixture models two expert-token exchanges across eight GPUs. Each GPU divides a 64 KiB aggregate payload across every other GPU:

Metric ECMP Adaptive Change
Collective completion 648.7 us 624.2 us -3.78%
Flow p99 320.3 us 308.0 us -3.83%
Retried bytes 224.2 MB 189.9 MB -15.30%
Route changes 0 24 -
go run ./cmd/fabricsim \
  -config scenarios/moe-all-to-all.json \
  -router both

Run a paired benchmark and generate a Markdown report:

go run ./cmd/fabricbench \
  -config scenarios/mixed-background.json \
  -runs 30 \
  -output results/mixed.jsonl

python3 analysis/report.py results/mixed.jsonl \
  --output results/mixed.md

Quick start

Requirements: Go 1.26 or newer. Python 3.10 or newer is only needed for the analysis helpers.

git clone https://github.com/iahsanGill/gpu-fabric-sim.git
cd gpu-fabric-sim

make check
go run ./cmd/fabricsim \
  -config scenarios/congested.json \
  -router both

Dashboard

go run ./cmd/fabricviz \
  -config scenarios/mixed-background.json \
  -addr :8080

Open http://localhost:8080, select ECMP or adaptive routing, and run the scenario. The dashboard is embedded in the Go binary and has no external UI or CDN dependencies.

Peak-pressure fabric heatmap

Dashboard endpoints:

Method Path Purpose
GET /api/topology Nodes and directional links
POST /api/run?router=adaptive Start a simulation
GET /api/state Current frame and final metrics
GET /api/trace Complete telemetry and route-change trace
GET /api/events Server-sent telemetry stream

Export the same trace without the dashboard:

go run ./cmd/fabricsim \
  -config scenarios/congested.json \
  -router adaptive \
  -trace results/adaptive-trace.json

Architecture

flowchart LR
  C[Scenario JSON] --> T[Clos topology]
  C --> W[Workload generator]
  T --> E[Event engine]
  W --> E
  E --> D[Link queues]
  D --> M[Delayed telemetry]
  M --> R[ECMP or adaptive routing]
  R --> D
  E --> X[Metrics and trace observer]
  X --> B[Benchmark report]
  X --> V[Dashboard]
Loading

The adaptive controller receives sampled state after telemetry and processing delays. It scores equal-cost candidate paths by queue occupancy, utilization, latency, and path length. A route changes only at an eligible flowlet boundary after the cooldown has elapsed and the new score clears the hysteresis margin.

See docs/ARCHITECTURE.md for the event model, workload barriers, transport boundary, and observability design.

Scenarios

File Workload Purpose
allreduce-small.json Ring all-reduce Fast deterministic smoke test
congested.json Ring all-reduce ECMP/adaptive stress comparison
incast.json Many-to-one incast Receiver-side hotspot
mixed-background.json All-reduce plus background flows Contention and tail latency
moe-all-to-all.json MoE all-to-all Expert-token exchange across every GPU pair

All inputs are explicit JSON: topology dimensions, link properties, tensor and chunk sizes, telemetry/control delay, thresholds, flowlet size, background traffic, and random seed.

Commands

make check       # format, tests, race detector, vet, Python syntax
make smoke       # run every scenario with both routers
make report      # write results/latest.jsonl and results/latest.md
make dashboard   # start the dashboard using SCENARIO
make udp         # run the loopback UDP loss/retry demo

The UDP demo uses sequence numbers, acknowledgements, deterministic loss, and retries. It is a separate transport exercise and is not used to produce simulator results.

Repository layout

cmd/          simulator, benchmark, dashboard, and UDP entry points
internal/     topology, event engine, routing, metrics, dashboard, transport
scenarios/    reproducible workload and fabric configurations
analysis/     dependency-free benchmark summarization and report generation
docs/         architecture, validation protocol, roadmap, and resume notes

Validation and scope

CI runs formatting checks, unit/integration tests, the race detector, go vet, every checked-in scenario, and a benchmark-report smoke test. The full local protocol is documented in docs/VALIDATION.md.

This repository demonstrates network simulation, traffic engineering, distributed-training communication, observability, and Go systems engineering. It does not implement CUDA kernels, production switch software, or hardware RDMA measurements.

License

MIT

About

Deterministic Go simulator for congestion-aware routing in GPU training fabrics

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages