Repository navigation
fix: 取地址公网优先,NAT 时经 IPv4 连 hub - #5
Merged
Merged
Conversation
- 每族一个地址,公网优先于内网、CGNAT、ULA 与 198.18/15;IPv6 避开临时、 废弃、tentative 与 DAD 失败的地址(读 /proc/net/if_inet6 的 flags) - 网卡 IPv4 为内网地址时先连 hub 的 IPv4,hub 由此看到 NAT 出口 - 逐个拨号,除最后一个地址外每个限 5 秒。tokio 的 connect 单个地址不限时, 黑洞地址会占满内核 127 秒的 SYN 重试,超过 120 秒的连接期限,后面的地址 不会被尝试 Refs monitor-probe/monitor#19
This was referenced Sep 18, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
问题
monitor-probe/monitor#19 场景一:LXC NAT 小鸡的 eth0 上是
10.10.x/22和 ULAfd42:…/64,eth1 上是公网2401:…/128。agent 每族取第一个地址,报上去的是 ULA。节点经 v6 连 hub,NAT 出口的公网 v4 hub 看不到,网卡上也没有,于是谁都不知道。同一处还有个隐患:hub 同时有 A 和 AAAA 记录、而节点的 v6 路由不通时,
TcpStream::connect会按解析顺序逐个尝试,单个地址不设超时。黑洞的 v6 要等内核 127 秒的 SYN 重试,超过了 120 秒的CONNECT_DEADLINE,v4 地址永远轮不到,节点一直连不上。修改
src/collect.rspick():每族一个地址,公网地址优先,同一档内保留内核的顺序。is_public()在 v4 一侧排除 RFC 1918、CGNAT、回环、链路本地、0/8、192.0.0/24、198.18/15(TUN 模式代理的 fake-ip 段)、组播和保留段;v6 只认 2000::/3transient_v6():读/proc/net/if_inet6的 flags,IPv6 避开临时地址(隐私扩展)、已废弃地址(换前缀后留下的旧地址)、tentative 和 DAD 失败的地址。只剩这类地址时照样上报src/main.rsdial():自己解析 hub 的地址。网卡 IPv4 是内网地址时先连 IPv4,hub 由此看到 NAT 出口;其他情况保持解析器的顺序connect_first():逐个拨号,除最后一个地址外每个限时DIAL_FALLBACK(5 秒,覆盖首个 SYN 和两次重传)。最后一个地址不单独限时,所以只有一个慢地址的 hub 照样能连上。全部失败时,日志逐个列出每个地址的失败原因connected改为connected to <addr>,能看出走的是哪个地址族Facts改为在拨号前采集,同一份数据既用来决定先连哪个地址族,也用于 hello协议字段不变。
测试
cargo fmt --check、cargo clippy --all-targets -- -D warnings、cargo test(23 个)全部通过。the_reported_address_is_the_public_one_whatever_the_kernel_lists_first:覆盖 issue 里的 LXC 网卡布局;TUN 代理、CGNAT、局域网地址排在公网地址前面的情况;SLAAC 的临时地址和已废弃地址only_globally_routable_addresses_count_as_publica_black_holed_address_gives_way_to_the_next_within_the_fallback:accept 队列已满的 listener 会丢弃 SYN 而不是拒绝,用它在回环上模拟黑洞;测试使用暂停时钟a_host_behind_nat_dials_the_hubs_v4_first:依赖localhost解析出的::1排在前面。没有 IPv6 回环、或解析器把127.0.0.1排在前面时,这个测试跳过并打印原因(已在关闭 IPv6 的 network namespace 里验证)--server http://localhost:9911连只监听127.0.0.1的 hub,先尝试::1被拒,随即输出connected to 127.0.0.1:9911兼容
新 agent 配旧 hub、旧 agent 配新 hub 都能正常工作,发版没有先后要求。
行为变化:网卡 IPv4 是内网地址的节点改为经 IPv4 连 hub,面板上的连接来源随之从 v6 变成 NAT 出口的 v4。
Refs monitor-probe/monitor#19