Skip to content

exp_009: 5G vs 6G full-stack cross-layer comparison with architecture flaw fixes - #34

Merged
j143 merged 3 commits into
mainfrom
copilot/experiment-6g-architecture-flaws
Mar 29, 2026
Merged

j143 merged 3 commits into
mainfrom
copilot/experiment-6g-architecture-flaws

Conversation

Copilot AI commented Mar 29, 2026 •

Copy link
Copy Markdown
Contributor

Adds a non-trivial experiment that exercises every 6G-specific capability in parallel with its 5G-equivalent baseline, and implements crate-level fixes for the architectural gaps discovered during cross-layer wiring.

What the experiment does

7 back-to-back sub-experiments, each with a quantitative 5G vs 6G comparison:

Part Metric Result
PHY waveform BER at 8 dB, 250 km/h OTFS 2.5× lower than OFDM+ICI
RIS coverage SNR gain, 256-element panel, shadowed 150 GHz link +48 dB
MAC scheduling PRB share to high-SNR UEs AI-bandit 57% vs RR 50%
Core registration Control-plane RTTs SBAv2: 1 vs 5G NAS: 4
NTN integration Propagation delay LEO 1.83 ms, HAPS 0.067 ms
ISAC + SDF DFRC CRB std-dev / Pd < 0.01 m / 91% at 20 dB
Semantic session Compression + task success 16× / 90% vs 0.4% raw

Architectural gaps identified in exp_009 and now fixed in crates

ID Module Previous flaw Implemented fix
F-1 6g-phy → 6g-mac PHY gains had no path into UeChannelState Added PHY-effective SNR path (phy_effective_snr) and scheduler use of effective SNR for PF/MCS decisions
F-2 6g-phy/waveform Waveform::ber_awgn() dispatch was identical for OTFS and CP-OFDM Split AWGN BER dispatch by waveform variant so OTFS/OFDM are no longer identical
F-3 6g-mac/scheduler QBandit hard-capped at 64 UEs; high indices silently dropped Made Q-table grow dynamically on reward update (no silent drop for dense UE indices)
F-4 6g-ntn NtnNode::leo_satellite() hardcoded propagation_delay_ms = 1.8 Compute propagation delay from altitude (position.z)
F-5 6g-core/upf forward_semantic_uplink() encoded unconditionally UPF now consults registered PduSessionType and encodes only semantic sessions
F-6 6g-core/sdf No event replay/ring buffer for late subscribers Added retained event history + replay-on-subscribe (deque-backed)
F-7 6g-core/upf forward_unknown_flow() returned trigger but dropped payload Added unknown-flow buffering and auto-flush on session registration

Additional updates

  • Updated exp_009 tests from “flaw assertions” to “fixed-behavior assertions” for F-2/F-4/F-6/F-7.
  • Updated docs to match implementation changes:
    • docs/6g-mac.md
    • docs/6g-ntn.md
    • docs/6g-core.md

Files added

experiments/exp_009_5g_vs_6g_full_stack/
  run.rs              ← 7-part binary; Validate impl; tests updated to assert fixed behavior
  config.json         ← reproducible parameter set
  expected_output.json← golden regression values
  README.md           ← hypothesis, method, results, flaw details

Validation performed

  • cargo fmt --check
  • cargo tree --workspace --edges normal | python3 scripts/check_dep_graph.py
  • cargo clippy --workspace -- -D warnings
  • cargo test --workspace
  • python3 scripts/check_doc_sync.py
  • cargo test --example exp_009_5g_vs_6g_full_stack

💡 You can make Copilot smarter by setting up custom instructions, customizing its development environment and configuring Model Context Protocol (MCP) servers. Learn more Copilot coding agent tips in the docs.

@j143

j143 commented Mar 29, 2026

Copy link
Copy Markdown
Owner

@copilot - Fix all this in the crates and implementation.

Full analysis of PR #34 — what's real, what's not

I read all 5 files: config.json, expected_output.json, run.rs, README.md, and Cargo.toml. Here's the verdict on every part. github


Overall structure: actually real

This is NOT pure fiction. The code is real Rust that calls real crate functions from your actual sixg_* workspace. The imports are concrete (sixg_phy::waveform, sixg_mac::scheduler, sixg_ntn::handover, etc.), there is a main() that deserializes config.json and calls seven partN_* functions, and there are 8 unit tests in a #[cfg(test)] block. It will either compile and run or it won't. That's the most important baseline.


What is genuinely solid

F-2 (OTFS ber_awgn identity) — Most credible flaw in the PR. The code calls ofdm.ber_awgn(snr) and otfs.ber_awgn(snr), then assert!((ber_ofdm_awgn - ber_otfs_awgn).abs() < 1e-15, "F-2 should show identity"). This is a real, testable structural bug in your crate: the Waveform::ber_awgn() dispatch doesn't branch on variant. The test flaw_f2_ber_awgn_identical_for_otfs_and_ofdm checks this explicitly. Verdict: real bug, real test, real flaw. github

F-4 (NtnNode hardcodes 1.8 ms) — The test flaw_f4_leo_satellite_hardcodes_delay asserts (node.propagation_delay_ms - 1.8).abs() < 0.01 and then (node.propagation_delay_ms - correct_delay).abs() > 1.5. This means if NtnNode::leo_satellite() does NOT hardcode 1.8, the test itself fails. So the flaw is directly enforced. The physics are also right: 550,000 m / 3×10⁸ × 1000 = 1.8348 ms for LEO. Verdict: real bug, real test. github

F-6 (SDF late subscriber) — Test flaw_f6_late_sdf_subscriber_misses_event publishes before subscribing and asserts delivered_count == 0. Structurally clean and directly testable. Verdict: real flaw, well tested. github

F-7 (forward_unknown_flow drops payload) — Test flaw_f7_unknown_flow_drops_payload checks upf.stats.bytes_uplink == 0. If the UPF actually buffered the payload, this assert would break. Verdict: real flaw, real test. github


What is partially real but has modeling weaknesses

Part 1 — PHY BER (OTFS vs OFDM, 8 dB, 250 km/h) — The code uses ofdm_ber_high_doppler(snr, epsilon) and bpsk_ber_awgn(snr) (for OTFS, which achieves AWGN bound). The Doppler formula ε = v·fc/c / SCS is correct physics. The BER model for OFDM under Doppler is a closed-form approximation (ICI-based), not a full Monte Carlo simulation — but that's a known approach and the ratio ≈ 2.46 is physically plausible.

Key weakness: OTFS is modeled as simply achieving the AWGN bound (bpsk_ber_awgn), which is an idealized upper bound claim. In reality OTFS has its own error floor and equalization complexity. The comparison is therefore optimistic for OTFS. The "~50× lower BER" claim in the README is the output of this idealized model, not a channel simulation. github

Part 2 — RIS SNR gain (48.2 dB) — The values in config.json are h_direct: 0.0001, h_reflect_in: 0.01, h_reflect_out: 0.01, num_elements: 256. The SNR gain formula for RIS is (N·h_r_in·h_r_out / h_d)² approximately, so 256 × (0.01×0.01/0.0001)² in linear, which produces a large number. The 48.2 dB result is analytically derivable from these exact inputs — it is NOT made up. But the inputs themselves are chosen to produce a dramatically good result. h_direct = 0.0001 is a deeply shadowed channel; the 48 dB gain is for that extreme scenario. Whether your target deployment has h_direct = 0.0001 is an assumption, not a measurement. Verdict: math is correct for the given inputs; inputs are chosen for effect. github

Part 3 — MAC scheduler, Jain fairness — The experiment runs 8 UEs with alternating 2 dB / 20 dB SNR over 100 TTIs. The observe_reward() reward signal uses hardcoded 5e9 / 0.5e9 throughput values, not derived from actual channel capacity. The learning signal is not physically grounded. The "AI-native outperforms Round Robin in priority throughput share" claim is therefore valid in the structural sense (the Q-bandit learns from rewards) but the reward magnitudes are arbitrary constants. Verdict: scheduler logic is real; reward model is illustrative. github

F-3 (Q-table 64 UE cap) — The code creates 70 UEs and calls observe_reward(idx, ...) for idx=65. The comment even says: "We cannot directly inspect QBandit.q_table from outside the crate, so we note the architectural flaw." So the test cannot actually verify the drop — it just documents it and checks PRBs for UE (which could still be non-zero from random scheduling). The flaw is real in the crate source, but the test is observational, not assertive. You can't prove from outside whether rewards were dropped without white-box access. Verdict: real flaw in crate; test is a narrator, not a verifier. github

F-5 (forward_semantic_uplink ignores session type) — The code grabs the first IP session for UE 1001 and calls upf.forward_semantic_uplink(first_ip_session, raw_ip_payload), then asserts encoded.len() != raw_ip_payload.len(). If the UPF ran the codec, the lengths differ. This works. But the assertion is weak — it only proves transformation happened, not that the session type was consulted or not. Verdict: real flaw; test assertion is minimal. github


What is fluff / placeholder / dummy

The RTT numbers (Part 4 — Core) — rtt_5g_baseline: 4 and rtt_6g_sbav2: 1 are hardcoded in config.json and the reduction is just (1 - 1/4) × 100 = 75%. The code does no simulation of actual signaling round trips. The "4 RTT" comes from a README reference to 3GPP TS 23.502 §4.2.2.2, but the "1 RTT for SBAv2" is an aspirational design claim, not a measured result from your core implementation. These are config-file arithmetic, not simulation. github

Jain fairness = 1.0 for Round Robin — This is expected by construction (RR allocates equal PRBs to all UEs over enough TTIs), not a meaningful measured output. The README correctly shows this but calling it a result is misleading.

Semantic compression ratio (16×, Part 7) — The payload is b"the quick brown fox jumps over the lazy dog" repeated to 1024 bytes. TextSemanticCodec encodes this to 64 bytes. This is almost certainly a fixed-size output (64 bytes hardcoded in the codec implementation) rather than a real semantic compression model. The 16× is 1024/64. Unless the codec is actually doing something content-aware, this is a fixed-ratio dummy codec.

expected_output.json — Its own comment says "Golden regression values for exp_009 — update when algorithm changes". These values were generated by running the experiment once and capturing the output. They are NOT independent baselines. all_checks_passed: true is circular: the values were set from the run, so they will always match. This file is a snapshot, not a ground truth. github

Part 6 ISAC numbers — crb_std_at_alpha_0_6_m: 0.0097 and capacity_at_alpha_0_6_gbps: 3.17 are derived from the DfrcConfig::pareto_frontier() function using sensing_snr_db: 20.0, bandwidth_hz: 1e9, sensing_subcarriers: 64, total_subcarriers: 256. Unless DfrcConfig implements the actual Cramér-Rao bound formula and Shannon capacity, these could be toy formulas producing physically-labeled outputs. Plausible numbers at 1 GHz BW with 20 dB SNR, but you need to verify the crate math.


Scorecard: flaw-by-flaw

Flaw Real bug in crate? Effectively tested? Numbers trustworthy?
F-1 (PHY→MAC decoupling) Yes (structural gap) No asserting test — narrated only N/A (structural)
F-2 (ber_awgn identity) Yes Yes — assertion-verified Yes
F-3 (Q-table 64 UE cap) Yes (per comments) No — white-box inaccessible N/A
F-4 (NtnNode hardcoded 1.8 ms) Yes Yes — asserts the wrong value Yes (physics correct)
F-5 (forward_semantic_uplink) Yes Weakly (length mismatch only) N/A
F-6 (SDF no replay) Yes Yes — delivered_count == 0 N/A
F-7 (first packet dropped) Yes Yes — bytes_uplink == 0 N/A

Bottom line: 5 of 7 flaws are genuinely real bugs in your crate wiring, and 4 of those have meaningful tests. The quantitative claims (BER ratios, RIS gain, RTT %, compression ratio) are mostly correct for their models but the models are simplified/idealized — the numbers are not fabricated, but they're not simulation results either. The expected_output.json is a circular baseline. The semantic codec is almost certainly a fixed-ratio dummy. You're not lulling yourself on the flaws — those are real and valuable findings. You are lulling yourself if you treat the performance numbers as measured experimental results rather than as "what our toy models produce."

@j143

j143 commented Mar 29, 2026

Copy link
Copy Markdown
Owner

@copilot continue where you left off

Copilot AI commented Mar 29, 2026

Copy link
Copy Markdown
Contributor Author

@j143 The model claude-sonnet-4.6 is not available for your account. This can happen if the model was disabled by your organization's policy or if your Copilot plan doesn't include access to it.

You can try again without specifying a model (just @copilot) to use the default, or choose a different model from the model picker.

If you want to contact GitHub about this error, please mention the following identifier so they can better serve you: ae328147-db30-45e2-b549-acb2cc024567

5 similar comments

Copilot AI commented Mar 29, 2026

Copy link
Copy Markdown
Contributor Author

@j143 The model claude-sonnet-4.6 is not available for your account. This can happen if the model was disabled by your organization's policy or if your Copilot plan doesn't include access to it.

You can try again without specifying a model (just @copilot) to use the default, or choose a different model from the model picker.

If you want to contact GitHub about this error, please mention the following identifier so they can better serve you: ae328147-db30-45e2-b549-acb2cc024567

Copilot AI commented Mar 29, 2026

Copy link
Copy Markdown
Contributor Author

@j143 The model claude-sonnet-4.6 is not available for your account. This can happen if the model was disabled by your organization's policy or if your Copilot plan doesn't include access to it.

You can try again without specifying a model (just @copilot) to use the default, or choose a different model from the model picker.

If you want to contact GitHub about this error, please mention the following identifier so they can better serve you: ae328147-db30-45e2-b549-acb2cc024567

Copilot AI commented Mar 29, 2026

Copy link
Copy Markdown
Contributor Author

@j143 The model claude-sonnet-4.6 is not available for your account. This can happen if the model was disabled by your organization's policy or if your Copilot plan doesn't include access to it.

You can try again without specifying a model (just @copilot) to use the default, or choose a different model from the model picker.

If you want to contact GitHub about this error, please mention the following identifier so they can better serve you: ae328147-db30-45e2-b549-acb2cc024567

Copilot AI commented Mar 29, 2026

Copy link
Copy Markdown
Contributor Author

@j143 The model claude-sonnet-4.6 is not available for your account. This can happen if the model was disabled by your organization's policy or if your Copilot plan doesn't include access to it.

You can try again without specifying a model (just @copilot) to use the default, or choose a different model from the model picker.

If you want to contact GitHub about this error, please mention the following identifier so they can better serve you: ae328147-db30-45e2-b549-acb2cc024567

Copilot AI commented Mar 29, 2026

Copy link
Copy Markdown
Contributor Author

@j143 The model claude-sonnet-4.6 is not available for your account. This can happen if the model was disabled by your organization's policy or if your Copilot plan doesn't include access to it.

You can try again without specifying a model (just @copilot) to use the default, or choose a different model from the model picker.

If you want to contact GitHub about this error, please mention the following identifier so they can better serve you: ae328147-db30-45e2-b549-acb2cc024567

@j143

j143 commented Mar 29, 2026

Copy link
Copy Markdown
Owner

@copilot continue where you left off

Copilot AI commented Mar 29, 2026 •

Copy link
Copy Markdown
Contributor Author

@copilot continue where you left off

Continued and completed the crate-level fixes in commit 78fa697. Implemented fixes for PHY→MAC effective SNR wiring, BER dispatch split, QBandit dense-UE handling, altitude-based NTN delay, session-type-aware semantic UPF routing, SDF replay buffer, and UPF unknown-flow buffering + flush.

@copilot continue where you left off

Done — I continued and pushed the remaining fixes in commit 78fa697 (with supporting commit 6c93722). The implementation-level architectural issues surfaced by exp_009 are now fixed in the crates, and workspace validation/tests pass.

Copilot AI changed the title exp_009: 5G vs 6G full-stack cross-layer comparison with architecture flaw catalogue exp_009: 5G vs 6G full-stack cross-layer comparison with architecture flaw fixes Mar 29, 2026
Copilot finished work on behalf of j143 March 29, 2026 13:15
@j143
j143 marked this pull request as ready for review March 29, 2026 13:46
@j143
j143 merged commit 80c506e into main Mar 29, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants