Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

voice-ptyd

Global push-to-talk voice input for Linux. Hold a key, speak, release — transcribed text (with punctuation) is pasted at your cursor, in any app. Fully offline.

按住快捷键说话,松开即把带标点的识别结果粘贴到光标处。完全离线。中文说明

hold F9 ●─────────────── speak 🎙 ─────────────── release ○
   record (pw-record) → transcribe (sherpa-onnx, ~0.15 s)
   → clipboard → Ctrl+V injected at cursor → notification 🔔
Engines X-ASR zh+en transducer (built-in punctuation, int8, ~136 MB) · SenseVoice int8 (zh/en/ja/ko/yue) + CT-Transformer punct
Latency ~0.1–0.3 s transcription for a sentence (CPU, 8 threads; no GPU needed)
Display server Wayland (GNOME/KDE/Hyprland/Sway) and X11
Audio PipeWire (pw-record/pw-play)
Injection ydotool ≥ 1.0 via persistent ydotoold (uinput) — no per-app plugins
Privacy 100% local. No cloud, no network after model download.
Trigger evdev hotkey listener (hold = record) + optional Ctrl+Alt+V toggle fallback

Why this exists

On Ubuntu 22.04 + GNOME Wayland + fcitx5 there was no working open-source voice input:

  • fcitx5-vinput — great project, but its capture silently returns nothing on PipeWire 0.3.48 (the 22.04 default). It works on 24.04+; use it there.
  • SenseVoice GUI tools — X11 only. IBus-based tools — broken with fcitx5. Cloud-based tools (unhush/whisrs/Bolo) — need API keys.

voice-ptyd fills that gap with a boring, auditable ~300-line Python daemon:

evdev (all keyboards, hotplug rescan)          ← hold-to-talk trigger
   ↓
pw-record 16 kHz mono  →  energy trim  →  sherpa-onnx offline ASR
   ↓
wl-copy (+ xclip fallback)  →  ydotool key combo (Ctrl+V / Shift+Ins)
   ↓
notify-send + ding/dong + history.jsonl

The udev rule ships a uaccess tag so the active seat user can read keyboard events (same mechanism as YubiKeys) — no root daemon, no input group, no re-login.

Requirements

  • Linux with systemd user sessions; apt-based distro for install.sh (tested: Ubuntu 22.04)
  • PipeWire audio (installer offers to switch from PulseAudio — that switch is required on 22.04)
  • A microphone 🙂

Quick start

git clone https://github.com/changer-changer/voice-ptyd
cd voice-ptyd
./install.sh          # sudo for apt/udev steps; everything else per-user

Then hold F9 and speak. Fallback toggle: Ctrl+Alt+V.

Uninstall: ./uninstall.sh (add --full to also remove ydotoold + udev rule).

Configuration — config.json

{
  "key_code": 67,          // KEY_F9; KEY_F10=68, KEY_RIGHTALT=100, … (linux/input-event-codes.h)
  "key_name": "F9",
  "engine": "xasr",        // "xasr" (zh-en, built-in punct) | "sensevoice" (5-lang)
  "paste_mode": "ctrlv",   // "shiftins" for terminals | "clipboard" (copy only)
  "max_seconds": 90,
  "min_audio_seconds": 0.35,
  "notify": true,
  "ding": true,            // start/stop chime
  "language": "auto"       // sensevoice engine
}

After editing: systemctl --user restart voice-ptyd.

Troubleshooting

Symptom Fix
PermissionError on /dev/input/event* in logs Re-run installer or sudo udevadm trigger -c add -s input; make sure your session is the active seat (loginctl list-sessions)
Recognition always empty pactl info must say “on PipeWire”; test capture: pw-record --rate 16000 /tmp/t.wav (speak, Ctrl-C, check the file isn't silent). On 22.04 the installer's PipeWire switch is mandatory.
Text pastes into the wrong window By design it goes to the focused window; click the target first, then hold the key.
Nothing in terminals with ctrlv Terminals usually want Ctrl+Shift+V; set "paste_mode": "shiftins" instead.
F9 does nothing on laptop keyboards Try Fn+F9, or bind another key_code, or use the Ctrl+Alt+V toggle.
ydotool “latency” warnings You're running 0.1.x without ydotoold; the installer builds 1.0.4 (scripts/build-ydotool.sh).

Logs: journalctl --user -u voice-ptyd -f · History: ~/.local/share/voice-ptyd/history.jsonl

How it compares

voice-ptyd fcitx5-vinput voiceio nerd-dictation Speech Note
Offline Chinese w/ punctuation ✅ ✅ ✅ (whisper) ❌ (vosk) ✅
GNOME Wayland + fcitx5 (no IME switch) ✅ ✅ (PipeWire ≥1.0) ❌ (IBus) ❌ (xdotool) GUI-app only
Push-to-talk (hold) ✅ evdev ✅ ✅ evdev partial portal (GNOME 48+)
No cloud / no API key ✅ ✅ (optional cloud) ✅ local ✅ ✅
Ubuntu 22.04 ✅ ❌ capture broken untested X11 only partial
Root daemon ❌ (uaccess) ❌ user group input varies ❌

Security notes

  • The daemon reads keyboard events only to watch your trigger key (filter on key_code); the code is ~300 lines and auditable in one sitting.
  • uaccess grants the active local seat user (you) access — same as fingerprint readers.
  • ydotoold runs as your user, not root. Audio never leaves the machine.

Model licenses

Models are downloaded from the official k2-fsa/sherpa-onnx releases (see their model cards for terms; SenseVoice is from FunAudioLLM/FunASR). This repo ships no model weights.

Roadmap

  • Pre-roll ring buffer (zero first-syllable clipping)
  • Hotword/lexicon boosting for the X-ASR engine
  • Streaming (partial results while holding)
  • Packaging (deb / AUR), install support beyond apt

中文说明

voice-ptyd 解决的是 Ubuntu 22.04 + GNOME Wayland + fcitx5 环境下“没有能用的开源语音输入”这个具体问题: fcitx5-vinput 被 22.04 的旧 PipeWire(0.3.48)卡死、SenseVoice 系工具仅支持 X11、IBus 系与 fcitx5 冲突、云 API 系要 Key。

git clone https://github.com/changer-changer/voice-ptyd && cd voice-ptyd
./install.sh     # 会询问是否把音频栈切到 PipeWire(22.04 必须,可回滚)
  • 按住 F9 说话,松开即粘贴(中英文混说、自动标点、提示音、通知、历史记录)
  • 完全离线:模型 ~136 MB(X-ASR 中英标点 int8),另可选 SenseVoice 五语种 + 标点模型
  • 不需要 root 守护进程:udev uaccess 规则让当前登录用户读取键盘事件(与 YubiKey 同机制)
  • 注入走持久 ydotoold(uinput),Wayland 下任意应用可用;终端建议 "paste_mode": "shiftins"
  • 配置 config.json(换键、换引擎、关提示音),改完 systemctl --user restart voice-ptyd

升级到 Ubuntu 24.04 后,也可以启用 fcitx5-vinput 获得更完整的体验(悬浮条、命令模式、豆包云 ASR)——两者共存不冲突。

License

MIT — see LICENSE. Models keep their upstream licenses.

About

Global push-to-talk voice input for Linux — hold a key, speak, punctuated zh/en text pasted at cursor. 100% offline (sherpa-onnx). Works on Wayland & X11.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages