Skip to content

Add ESP32-P4 support; fix voice sequence numbers and VAD cut-offs - #2

Open
haarts wants to merge 3 commits into
dchote:mainfrom
haarts:p4-support
Open

haarts wants to merge 3 commits into
dchote:mainfrom
haarts:p4-support

Conversation

@haarts

@haarts haarts commented Oct 3, 2026

Copy link
Copy Markdown

Adds support for the Waveshare ESP32-P4-WIFI6-POE-ETH and fixes two audio problems that affect every board. Three commits that can be reviewed separately:

1. Add ESP32-P4 support

  • Opus: OPUS_XTENSA_LX7 is now set only for the ESP32-S3 instead of in library.json, so the P4 (RISC-V) uses the generic C code. lib/micro-opus/CMakeLists.txt registers the vendored Opus as an IDF component for ESPHome's native esp-idf toolchain.
  • Toolchain: the P4 config uses the native esp-idf toolchain, because PlatformIO's can't link rev3 P4 builds at the moment (pioarduino 55.03.39 expects sections.rev3.ld.in, which its ESP-IDF 5.5.5 doesn't ship). The S3 configs keep using PlatformIO, unchanged.
  • Wi-Fi-only calls are guarded; the default username uses the Ethernet MAC when there's no Wi-Fi.
  • New mic_warmup option: settle time after acquiring the I2S bus. It defaults to the previous 200 ms; the P4 config uses 50 ms.
  • Board config esphome/waveshare-esp32-p4-poe.yaml: IP101 Ethernet, ES8311 with the board's analog mic. ESPHome's use_microphone: true selects a PDM mic, so it stays off. Mic preamp and digital gain are runtime settings (the defaults of 72 dB in total overdrove the input).

2. Fix voice packet sequence numbers (all boards)

The voice sequence counts 10 ms frames: desktop Mumble advances it by 2 per 20 ms Opus packet and keeps it running across utterances. We advanced it by 1 and reset it to 0 at every utterance. Mumla and desktop Mumble time their jitter buffer by it and dropped most of the audio, so speech came through in fragments of a few hundred ms. pymumble and the server don't care, which is probably why it went unnoticed.

3. Improve voice detection and push-to-talk (all boards)

  • The adaptive noise floor crept up under continuous speech and cut TX after about a second. It now only learns the ambient level between utterances.
  • Hangover goes from 300 to 800 ms, so pauses between words don't end TX.
  • Push-to-talk sends everything while the button is held, with no VAD gating and no calibration window swallowing the first word.
  • Always-on sends a 200 ms pre-roll when VAD triggers.

Testing

  • On hardware: Waveshare ESP32-P4-WIFI6-POE-ETH (rev v3.2), against Murmur 1.5 (official Docker image), listening in Mumla on Android.
  • End to end: speech was played next to the board, recorded from the server with a pymumble client, and transcribed with Whisper. A Dutch test sentence went from about 60% word error rate to 0% in both always-on and push-to-talk.
  • Builds: pre-commit run --all-files passes on each commit. The P4 config compiles at each commit, and esp32-s3-box.yaml compiles with all three.
  • Not tested on S3 hardware: I don't have an S3 board. Commits 2 and 3 change shared behavior, so a quick check on a Box or Voice PE would be welcome.

I haven't added the P4 board to the CI matrix or the release/OTA manifests, because it needs the native esp-idf toolchain instead of PlatformIO. Happy to add that if you want it.

🤖 Generated with Claude Code

https://claude.ai/code/session_01NmT8zZQWZTiXj38MZhU3tT

haarts and others added 3 commits October 3, 2026 15:30
The ESP32-P4 is RISC-V, has no built-in Wi-Fi and uses Ethernet:

- micro-opus: enable the Xtensa LX7 code paths only on the ESP32-S3 instead
  of hard-coding OPUS_XTENSA_LX7, and register the vendored Opus as an IDF
  component (lib/micro-opus/CMakeLists.txt) so it builds with ESPHome's
  native esp-idf toolchain. The PlatformIO toolchain can't link rev3 P4
  builds at the moment (pioarduino 55.03.39 expects sections.rev3.ld.in,
  which its ESP-IDF 5.5.5 doesn't ship).
- Guard the Wi-Fi-only calls (power save, Wi-Fi MAC) and use the Ethernet
  MAC for the default username when there is no Wi-Fi.
- New mic_warmup option: settle time after acquiring the I2S bus. Defaults
  to the previous 200 ms (the ES7210 on the S3 boxes needs it); the
  ES8311 on the P4 board is fine with 50 ms.
- Board config esphome/waveshare-esp32-p4-poe.yaml: IP101 Ethernet, ES8311
  with the board's analog mic (ESPHome's use_microphone selects a PDM mic,
  so it stays off), mic preamp/digital gain as runtime settings, PTT on a
  header GPIO and a "Talk" switch for push-to-talk from Home Assistant.

Tested on the board (rev v3.2 silicon) against Murmur 1.5.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NmT8zZQWZTiXj38MZhU3tT
The Mumble voice sequence counts 10 ms frames: desktop Mumble advances it by
2 for a 20 ms Opus packet and keeps it running across utterances. We
advanced it by 1 per 20 ms packet and reset it to 0 at the start of every
utterance. Receivers time their jitter buffer by the sequence, so Mumla and
desktop Mumble dropped most of our audio: speech came through in fragments
of a few hundred milliseconds. Affects every board.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NmT8zZQWZTiXj38MZhU3tT
- The adaptive noise floor no longer rises while speech is being sent. It
  crept up under continuous speech (0.01 per frame, plus fast adaptation on
  quiet syllables) and cut transmission off after about a second; now it
  only learns the ambient level between utterances, with a much slower
  recovery creep (0.0005).
- Hangover 300 -> 800 ms, so normal pauses between words don't end TX.
- Push-to-talk sends everything while the button is held: no VAD gating
  and no calibration window that swallowed the first word.
- Always-on keeps a 200 ms pre-roll while waiting for VAD and sends it when
  TX starts, so the start of the first word isn't lost to the attack time.

Measured end to end (speech played near the board, recorded from the
server, transcribed with Whisper): a Dutch test sentence went from about
60% word error rate to 0% in both always-on and push-to-talk.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NmT8zZQWZTiXj38MZhU3tT

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant