Global push-to-talk voice input for Linux. Hold a key, speak, release — transcribed text (with punctuation) is pasted at your cursor, in any app. Fully offline.
按住快捷键说话,松开即把带标点的识别结果粘贴到光标处。完全离线。中文说明
hold F9 ●─────────────── speak 🎙 ─────────────── release ○
record (pw-record) → transcribe (sherpa-onnx, ~0.15 s)
→ clipboard → Ctrl+V injected at cursor → notification 🔔
| Engines | X-ASR zh+en transducer (built-in punctuation, int8, ~136 MB) · SenseVoice int8 (zh/en/ja/ko/yue) + CT-Transformer punct |
| Latency | ~0.1–0.3 s transcription for a sentence (CPU, 8 threads; no GPU needed) |
| Display server | Wayland (GNOME/KDE/Hyprland/Sway) and X11 |
| Audio | PipeWire (pw-record/pw-play) |
| Injection | ydotool ≥ 1.0 via persistent ydotoold (uinput) — no per-app plugins |
| Privacy | 100% local. No cloud, no network after model download. |
| Trigger | evdev hotkey listener (hold = record) + optional Ctrl+Alt+V toggle fallback |
On Ubuntu 22.04 + GNOME Wayland + fcitx5 there was no working open-source voice input:
- fcitx5-vinput — great project, but its capture silently returns nothing on PipeWire 0.3.48 (the 22.04 default). It works on 24.04+; use it there.
- SenseVoice GUI tools — X11 only. IBus-based tools — broken with fcitx5. Cloud-based tools (unhush/whisrs/Bolo) — need API keys.
voice-ptyd fills that gap with a boring, auditable ~300-line Python daemon:
evdev (all keyboards, hotplug rescan) ← hold-to-talk trigger
↓
pw-record 16 kHz mono → energy trim → sherpa-onnx offline ASR
↓
wl-copy (+ xclip fallback) → ydotool key combo (Ctrl+V / Shift+Ins)
↓
notify-send + ding/dong + history.jsonl
The udev rule ships a uaccess tag so the active seat user can read keyboard
events (same mechanism as YubiKeys) — no root daemon, no input group, no re-login.
- Linux with systemd user sessions; apt-based distro for
install.sh(tested: Ubuntu 22.04) - PipeWire audio (installer offers to switch from PulseAudio — that switch is required on 22.04)
- A microphone 🙂
git clone https://github.com/changer-changer/voice-ptyd
cd voice-ptyd
./install.sh # sudo for apt/udev steps; everything else per-userThen hold F9 and speak. Fallback toggle: Ctrl+Alt+V.
Uninstall: ./uninstall.sh (add --full to also remove ydotoold + udev rule).
After editing: systemctl --user restart voice-ptyd.
| Symptom | Fix |
|---|---|
PermissionError on /dev/input/event* in logs |
Re-run installer or sudo udevadm trigger -c add -s input; make sure your session is the active seat (loginctl list-sessions) |
| Recognition always empty | pactl info must say “on PipeWire”; test capture: pw-record --rate 16000 /tmp/t.wav (speak, Ctrl-C, check the file isn't silent). On 22.04 the installer's PipeWire switch is mandatory. |
| Text pastes into the wrong window | By design it goes to the focused window; click the target first, then hold the key. |
Nothing in terminals with ctrlv |
Terminals usually want Ctrl+Shift+V; set "paste_mode": "shiftins" instead. |
| F9 does nothing on laptop keyboards | Try Fn+F9, or bind another key_code, or use the Ctrl+Alt+V toggle. |
| ydotool “latency” warnings | You're running 0.1.x without ydotoold; the installer builds 1.0.4 (scripts/build-ydotool.sh). |
Logs: journalctl --user -u voice-ptyd -f · History: ~/.local/share/voice-ptyd/history.jsonl
| voice-ptyd | fcitx5-vinput | voiceio | nerd-dictation | Speech Note | |
|---|---|---|---|---|---|
| Offline Chinese w/ punctuation | ✅ | ✅ | ✅ (whisper) | ❌ (vosk) | ✅ |
| GNOME Wayland + fcitx5 (no IME switch) | ✅ | ✅ (PipeWire ≥1.0) | ❌ (IBus) | ❌ (xdotool) | GUI-app only |
| Push-to-talk (hold) | ✅ evdev | ✅ | ✅ evdev | partial | portal (GNOME 48+) |
| No cloud / no API key | ✅ | ✅ (optional cloud) | ✅ local | ✅ | ✅ |
| Ubuntu 22.04 | ✅ | ❌ capture broken | untested | X11 only | partial |
| Root daemon | ❌ (uaccess) | ❌ | user group input |
varies | ❌ |
- The daemon reads keyboard events only to watch your trigger key (filter on
key_code); the code is ~300 lines and auditable in one sitting. - uaccess grants the active local seat user (you) access — same as fingerprint readers.
- ydotoold runs as your user, not root. Audio never leaves the machine.
Models are downloaded from the official k2-fsa/sherpa-onnx releases (see their model cards for terms; SenseVoice is from FunAudioLLM/FunASR). This repo ships no model weights.
- Pre-roll ring buffer (zero first-syllable clipping)
- Hotword/lexicon boosting for the X-ASR engine
- Streaming (partial results while holding)
- Packaging (deb / AUR), install support beyond apt
voice-ptyd 解决的是 Ubuntu 22.04 + GNOME Wayland + fcitx5 环境下“没有能用的开源语音输入”这个具体问题: fcitx5-vinput 被 22.04 的旧 PipeWire(0.3.48)卡死、SenseVoice 系工具仅支持 X11、IBus 系与 fcitx5 冲突、云 API 系要 Key。
git clone https://github.com/changer-changer/voice-ptyd && cd voice-ptyd
./install.sh # 会询问是否把音频栈切到 PipeWire(22.04 必须,可回滚)- 按住 F9 说话,松开即粘贴(中英文混说、自动标点、提示音、通知、历史记录)
- 完全离线:模型 ~136 MB(X-ASR 中英标点 int8),另可选 SenseVoice 五语种 + 标点模型
- 不需要 root 守护进程:udev uaccess 规则让当前登录用户读取键盘事件(与 YubiKey 同机制)
- 注入走持久
ydotoold(uinput),Wayland 下任意应用可用;终端建议"paste_mode": "shiftins" - 配置
config.json(换键、换引擎、关提示音),改完systemctl --user restart voice-ptyd
升级到 Ubuntu 24.04 后,也可以启用 fcitx5-vinput 获得更完整的体验(悬浮条、命令模式、豆包云 ASR)——两者共存不冲突。
MIT — see LICENSE. Models keep their upstream licenses.
{ "key_code": 67, // KEY_F9; KEY_F10=68, KEY_RIGHTALT=100, … (linux/input-event-codes.h) "key_name": "F9", "engine": "xasr", // "xasr" (zh-en, built-in punct) | "sensevoice" (5-lang) "paste_mode": "ctrlv", // "shiftins" for terminals | "clipboard" (copy only) "max_seconds": 90, "min_audio_seconds": 0.35, "notify": true, "ding": true, // start/stop chime "language": "auto" // sensevoice engine }