henyey mainnet daily — 2026-05-15 #2717
tomerweller
announced in
Announcements
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Validator
c14afa94(uptime <5m — just deployed at 13:31Z; pre-deploy was8b2c578cfor 3h03m, priorf486ea12for 19h13m. 2 deploys in last 24h, ~5.5m avg build)qset={},state=Initializing)first_to_self_externalize_seconds): 0 (post-restart, not yet observed)externalized_seconds): 0 (post-restart)Pre-deploy heartbeat: agree=21/missing=0, peer_gap=0, lag_ms=494 (validating on
8b2c578c).Deploys (2)
8b2c578cRevert INV-H2 strict bound to fix catchup panic (URGENT-CI: INV-H2 strict-bound panic (herder.rs:887) regresses #2693 — History Publish failing 2 days running #2708)c14afa94Extract ensure_tracking_state helper from duplicated blocks (closes Extract shared tracking-state transition helper from advance_tracking_to/advance_tracking_slot #2715)Incidents
🟢 INV-H2 strict-bound panic — daily History Publish (Testnet) crashed 2 days in a row
crates/herder/src/herder.rs:887—LCL == tracking_slotviolated the strict<bound on the catchup rapid-close path0114ae06tightened the assertion to strict<on the premise thatcomplete_externalizationalways advances tracking first; the catchup path doesn't go through that function8b2c578c(revert) + structural follow-up1f82d1b7/c14afa94(URGENT-CI: INV-H2 strict-bound panic (herder.rs:887) regresses #2693 — History Publish failing 2 days running #2708, Wire catchup rapid-close through tracking_slot advancement #2712, Extract shared tracking-state transition helper from advance_tracking_to/advance_tracking_slot #2715)🟢 Chained CI regressions from herder.rs refactor (
1f82d1b7)1f82d1b7introduced an idempotency guard that gated theSyncing→Trackingtransition (brokerow_05_slot_too_new_prefilter);a9cc892dfix introduced a drain-off-load ordering regression (broketest_issue_1773_receive_tx_set_frees_event_loop_during_drain)c14afa94(extractensure_tracking_statehelper; restored ordering as a side effect) — URGENT-CI: scp_stage_attribution row_05_slot_too_new_prefilter — Range prefilter accepts slot+10000 envelope after 1f82d1b7 #2714 closed, URGENT-CI: test_issue_1773_receive_tx_set_frees_event_loop_during_drain — drain blocks event loop after a9cc892d #2716 noted as resolved🟢 Validator panic Resource::can_add (Classic vs Soroban dim mixup)
Resource::can_add(tx-queue admission)9c223fb2(URGENT: validator panic — Resource::can_add size mismatch (2 vs 7) on f486ea12 after 20h #2690)🟢 Quickstart friendbot 5-layer recovery (saga across 16+h)
--in-memorywiring, empty in-memory DB 404, single-node tx not admitted, INV-H2 panic on captive-core syncd30c5652,b41c22f5,20034d9c,7353f5cf,dd118a99,8b2c578c(URGENT-CI: Quickstart friendbot test timing out for 16h on origin/main since 1e339515 (#2654 overlay work) #2696, URGENT-CI: Quickstart friendbot fails — captive-core RPC returns 404 on bot-account-detail (post-#2696 follow-up; blocks 65-commit deploy queue) #2699, URGENT-CI: Quickstart friendbot still fails after #2699 fix — captive-core /getledgerentry returns 404 (deploy gate now blocked ~10h) #2701, URGENT-CI: Quickstart friendbot tx never finalizes on local-mode captive-core (post-#2701 layer-2 issue; deploy queue blocked 13h) #2705)🟢 22 simulation tests timing out (timer-delivery wiring regression)
manual_close_untilhangs at ledger 2,scp_sent=0a99bd9a2"Replace SCP timeout polling with event-driven timer delivery" lost a code path92baba70,f4b5f89f(URGENT-CI: a99bd9a2 timer-delivery refactor broke 4 simulation tests — nomination stalls at round 1 (deploy still blocked) #2703, HERDER: Align timer delivery and trigger scheduling with stellar-core #2700)🔴 Post-deploy catchup-behind recovery noise (metrics: recovery-stalled burst (delta=240) + 155k overlay backpressure on post-deploy catchup-behind #2713 open)
archive_confirmed_behindarming logicIssues activity
Filed today (30):
Closed today (30): mass-closure window — see Incidents for narrative; root causes resolved by
8b2c578c,c14afa94,9c223fb2,d30c5652,b41c22f5,20034d9c,7353f5cf,dd118a99,92baba70,f4b5f89f.Still open (3 long-running):
7ee6ef3f; Stream 2 tracked in Overlay: receive-to-relay latency diagnostic (Stream 2 of #2644) #2648)Watch items
post_catchup_hr: appeared in 10 ticks (steady-state archive_behind churn; tracked indirectly by metrics: recovery-stalled burst (delta=240) + 155k overlay backpressure on post-deploy catchup-behind #2713 and pre-existing metrics: henyey_post_catchup_hard_reset_total — recurring ArchiveBehindStallWallClock livelock (Related to #2568) #2664)deploy_blocked: 8 ticks (mostly during Quickstart 16h saga and the herder.rs regression chain)frag: 18% → 27% transient post-restart → 20-22% steady, no #rss_mb: 11G pre-deploy → 5.7G post-deploy fresh process → 7.5G after module_cache rebuild, no ##2696/#2699 still open: 2 ticks each; both closed by 10:00Zrecovery_completed peer_gap=0: 1 tick (post-deploy recovery checkpoint)Tick aggregates (last 24h)
filed #2711Tier-2 entry)1f07680fMake deploy quarantine gate content-aware to auto-clear after revert)Open questions
#2713(post-deploy recovery-loop noise) is the only openurgent-eligible item from today's incidents. The fix landscape touchescrates/app/src/app/catchup_impl.rs_archive_confirmed_behind_untilarming logic and is structurally adjacent to closed metrics: henyey_post_catchup_hard_reset_total — recurring ArchiveBehindStallWallClock livelock (Related to #2568) #2664. Worth bumping to Backlog priority if next deploy reproduces the 11-min recovery thrash.All reactions