Skip to content

dns: reach nameservers over IPv6 - #313

Closed
david-yu wants to merge 1 commit into
redpanda-data:v26.3.xfrom
david-yu:dyu/dns-ipv6-nameservers
Closed

david-yu wants to merge 1 commit into
redpanda-data:v26.3.xfrom
david-yu:dyu/dns-ipv6-nameservers

Conversation

@david-yu

@david-yu david-yu commented Sep 13, 2026 •

Copy link
Copy Markdown

The c-ares integration could only talk to IPv4 nameservers: do_socket created every UDP channel as AF_INET regardless of the family c-ares asked for, and sock_addr threw "No ipv6 yet" on an AF_INET6 sockaddr. On a host whose resolv.conf lists an IPv6 nameserver — an IPv6-only Kubernetes pod with CoreDNS — every lookup fails with ARES_ECONNREFUSED, which the error category renders as "Connection refused" although no TCP connect was attempted.

  • do_socket: datagram channel with the requested af.
  • sock_addr: accept AF_INET6, check the sockaddr length.
  • do_recvfrom: copy the source by its real length (sockaddr truncates sockaddr_in6) and report it in *from_len.
  • options.servers: installed via ares_set_servers_csv after ares_init_options; ARES_OPT_SERVERS is in_addr-only.
  • ARES_ECONNREFUSED text → c-ares' own "Could not contact DNS servers".

IPv4 nameservers take the unchanged AF_INET branches.

Tests

Unit (this PR) — tests/unit/dns_test.cc

  • test_resolve_udp_ipv6_nameserver: a mock UDP nameserver bound to [::1] answers an A query; the resolver is configured with servers = {"::1"}. Before this change construction throws Servers must be ipv4 addresses; after it, the answer (127.0.0.42) comes back over the IPv6 transport. Skips when the host has no IPv6. The existing TCP split-response test's response builder was refactored to share the message body.

Downstream (redpanda, with the pin bumped to this commit)

  • //src/v/rpc/test:rpc_gen_cycling_test echo_round_trip_ipv6_hostname: localhost resolved through c-ares with family INET6 → ::1, RPC round-trip succeeds.
  • End-to-end on EKS 1.34 with IPv6-only pods (resolv.conf → fd1d:cca8:a7aa::a): an unpatched v26.2.2 cluster loops forever on C-Ares:11; with this change the 3-broker cluster forms, and survives rolling upgrade/downgrade, pod kill, all-at-once restart, scale 3→4→3, tiered storage and cloud topics to S3, config changes — 29/29 cases. Producer throughput is within 1% of an identical IPv4 cluster. Also verified on dual-stack AKS and GKE.

Context: redpanda-data/redpanda-operator#1855. Companion core change: redpanda-data/streaming-enterprise#330. Will also be proposed upstream to scylladb/seastar.

The c-ares integration created every UDP socket as AF_INET and rejected
any AF_INET6 sockaddr with "No ipv6 yet", so a host whose resolv.conf
lists an IPv6 nameserver could not resolve anything: c-ares reported
ARES_ECONNREFUSED for every server. That is what an IPv6-only
Kubernetes pod looks like (CoreDNS over IPv6).

- do_socket creates the datagram channel with the family c-ares asked
  for instead of hardcoding AF_INET.
- sock_addr accepts AF_INET6 and checks the sockaddr length.
- do_recvfrom copies the source address by its real length; sockaddr is
  only wide enough for AF_INET, so an IPv6 source was being truncated
  and its length misreported.
- options.servers is installed with ares_set_servers_csv after channel
  creation; ARES_OPT_SERVERS carries in_addr only.
- ARES_ECONNREFUSED is rendered with c-ares' own text, "Could not
  contact DNS servers", so it stops reading like a refused TCP connect.

IPv4 nameservers take the same AF_INET branches as before.

Adds a unit test that answers a query from a mock nameserver bound to
[::1].
@travisdowns

Copy link
Copy Markdown
Member

Please upstream to sesatar first, then downstream the change here.

@david-yu

Copy link
Copy Markdown
Author

No problem will close for now.

@david-yu david-yu closed this Sep 15, 2026
@baryluk

baryluk commented Sep 18, 2026

Copy link
Copy Markdown

Hi @david-yu , i wonder if you have a PR ready for seastar to upstream the support?

@david-yu

Copy link
Copy Markdown
Author

@baryluk I'll see what I can do, will try to perhaps upstream a PR next week.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants