A deterministic discrete-event simulator for congestion and routing in multi-hop GPU training fabrics. The project models a two-stage Clos network, replays all-reduce, mixture-of-experts all-to-all, and incast workloads, and compares static ECMP against telemetry-driven adaptive flowlet routing.
- Directional links with configurable bandwidth, propagation delay, and finite output buffers
- Ring all-reduce phases, MoE all-to-all token exchange, many-to-one incast, and background traffic
- Stable per-flow ECMP path selection
- Adaptive routing using delayed queue, utilization, and drop telemetry
- Flowlet boundaries, route cooldowns, and hysteresis
- Reliable chunk retries after buffer drops
- p50/p95/p99 flow completion, collective completion, queue depth, utilization, retry volume, route changes, and processed events
- Deterministic replay for a fixed scenario and seed
The simulation clock advances through a priority queue of logical events. It does not use wall-clock sleeps or one goroutine per packet, so operating-system scheduling does not change a run's result.
The checked-in congested.json fixture produces the following deterministic
comparison:
| Metric | ECMP | Adaptive | Change |
|---|---|---|---|
| Collective completion | 8.61 ms | 5.17 ms | -39.97% |
| Flow p99 | 286.9 us | 172.2 us | -39.97% |
| Retried bytes | 3.36 GB | 1.56 GB | -53.42% |
| Route changes | 0 | 480 | - |
These are simulator measurements for a synthetic fixture, not hardware benchmarks. The mixed-background scenario adds seed-stable start jitter for paired experiments with traffic variation.
At its checked-in seed (53), the moe-all-to-all.json fixture models two
expert-token exchanges across eight GPUs. Each GPU divides a 64 KiB aggregate
payload across every other GPU:
| Metric | ECMP | Adaptive | Change |
|---|---|---|---|
| Collective completion | 648.7 us | 624.2 us | -3.78% |
| Flow p99 | 320.3 us | 308.0 us | -3.83% |
| Retried bytes | 224.2 MB | 189.9 MB | -15.30% |
| Route changes | 0 | 24 | - |
go run ./cmd/fabricsim \
-config scenarios/moe-all-to-all.json \
-router bothRun a paired benchmark and generate a Markdown report:
go run ./cmd/fabricbench \
-config scenarios/mixed-background.json \
-runs 30 \
-output results/mixed.jsonl
python3 analysis/report.py results/mixed.jsonl \
--output results/mixed.mdRequirements: Go 1.26 or newer. Python 3.10 or newer is only needed for the analysis helpers.
git clone https://github.com/iahsanGill/gpu-fabric-sim.git
cd gpu-fabric-sim
make check
go run ./cmd/fabricsim \
-config scenarios/congested.json \
-router bothgo run ./cmd/fabricviz \
-config scenarios/mixed-background.json \
-addr :8080Open http://localhost:8080, select ECMP or adaptive routing, and run the scenario. The dashboard is embedded in the Go binary and has no external UI or CDN dependencies.
Dashboard endpoints:
| Method | Path | Purpose |
|---|---|---|
GET |
/api/topology |
Nodes and directional links |
POST |
/api/run?router=adaptive |
Start a simulation |
GET |
/api/state |
Current frame and final metrics |
GET |
/api/trace |
Complete telemetry and route-change trace |
GET |
/api/events |
Server-sent telemetry stream |
Export the same trace without the dashboard:
go run ./cmd/fabricsim \
-config scenarios/congested.json \
-router adaptive \
-trace results/adaptive-trace.jsonflowchart LR
C[Scenario JSON] --> T[Clos topology]
C --> W[Workload generator]
T --> E[Event engine]
W --> E
E --> D[Link queues]
D --> M[Delayed telemetry]
M --> R[ECMP or adaptive routing]
R --> D
E --> X[Metrics and trace observer]
X --> B[Benchmark report]
X --> V[Dashboard]
The adaptive controller receives sampled state after telemetry and processing delays. It scores equal-cost candidate paths by queue occupancy, utilization, latency, and path length. A route changes only at an eligible flowlet boundary after the cooldown has elapsed and the new score clears the hysteresis margin.
See docs/ARCHITECTURE.md for the event model, workload barriers, transport boundary, and observability design.
| File | Workload | Purpose |
|---|---|---|
allreduce-small.json |
Ring all-reduce | Fast deterministic smoke test |
congested.json |
Ring all-reduce | ECMP/adaptive stress comparison |
incast.json |
Many-to-one incast | Receiver-side hotspot |
mixed-background.json |
All-reduce plus background flows | Contention and tail latency |
moe-all-to-all.json |
MoE all-to-all | Expert-token exchange across every GPU pair |
All inputs are explicit JSON: topology dimensions, link properties, tensor and chunk sizes, telemetry/control delay, thresholds, flowlet size, background traffic, and random seed.
make check # format, tests, race detector, vet, Python syntax
make smoke # run every scenario with both routers
make report # write results/latest.jsonl and results/latest.md
make dashboard # start the dashboard using SCENARIO
make udp # run the loopback UDP loss/retry demoThe UDP demo uses sequence numbers, acknowledgements, deterministic loss, and retries. It is a separate transport exercise and is not used to produce simulator results.
cmd/ simulator, benchmark, dashboard, and UDP entry points
internal/ topology, event engine, routing, metrics, dashboard, transport
scenarios/ reproducible workload and fabric configurations
analysis/ dependency-free benchmark summarization and report generation
docs/ architecture, validation protocol, roadmap, and resume notes
CI runs formatting checks, unit/integration tests, the race detector, go vet,
every checked-in scenario, and a benchmark-report smoke test. The full local
protocol is documented in docs/VALIDATION.md.
This repository demonstrates network simulation, traffic engineering, distributed-training communication, observability, and Go systems engineering. It does not implement CUDA kernels, production switch software, or hardware RDMA measurements.

