Local Voice Agent combines speech recognition, a language model, and speech generation in one supervised application. It selects a model stack that fits the available hardware and memory.
Tip
The Jetson profile supports real-time voice conversations on a Jetson Orin Nano.
The application includes:
- A browser voice interface.
- Local streaming speech recognition with Nemotron Q8.
- Local language models through llama.cpp.
- Local speech generation with Kokoro.
- Automatic setup for CPU, NVIDIA, Apple Silicon, and Jetson.
- A remote-client mode for devices that run without a local browser.
Clone this repository before you start:
git clone https://github.com/ShayneP/local-voice-ai.git
cd local-voice-aiThe setup launcher needs Python 3.10 or later.
Install the additional tools for your platform:
| Platform | Requirement |
|---|---|
| Linux CPU | Docker Engine with Docker Compose |
| Desktop NVIDIA | Docker Engine, Docker Compose, and NVIDIA Container Toolkit |
| Jetson Orin | JetPack 6.2, L4T 36.4, and the NVIDIA Docker runtime |
| Apple Silicon | Python 3.11–3.13, uv, livekit-server, and llama-server |
On Apple Silicon, install the native server tools with Homebrew:
brew install livekit llama.cpp
uv sync --extra ml --extra devThe first start needs an internet connection. Later starts reuse downloaded model files and native components. Docker also reuses its image layers.
Start the setup launcher:
python3 run.pyThe launcher shows the detected hardware, memory budget, and recommended models. Accept the recommendation or select a different profile.
When the application is ready, open http://localhost:8080. When the browser requests microphone access, permit it.
For a non-interactive start, run:
python3 run.py start --profile auto --yesAutomatic selection uses the device type to select a runtime. It then uses the memory budget to select a model profile.
| Profile | Memory target | Language model | Context | Speech recognition | Voice |
|---|---|---|---|---|---|
lean |
About 4.7 GB | Qwen3 1.7B | 4K | Nemotron Q8 | Kokoro ONNX |
jetson-realtime |
About 4.7 GB | Qwen3 1.7B | 4K | Nemotron Q8 | Kokoro ONNX |
compact |
About 5.5 GB | Gemma 4 E2B | 4K | Nemotron Q8 | Kokoro |
balanced |
About 6.5 GB | Gemma 4 E2B | 16K | Nemotron Q8 | Kokoro |
The memory values are planning targets, not hard limits. The automatic mode keeps memory available for the operating system and active conversations.
All profiles use the native streaming Nemotron Q8 runtime. The launcher selects the CPU, CUDA, or Metal runtime for the device.
The default language is English. For English, the application uses the English-specific Nemotron model. For another supported language, it uses Nemotron 3.5.
Set the language in .env.local:
STT_LANGUAGE=fr-FRIf the speaker language can change, use STT_LANGUAGE=auto. This value selects
the multilingual model. Whisper remains available as a manual fallback:
STT_PROVIDER=whisperWhisper waits for a complete utterance before transcription. Nemotron sends partial transcripts while the user speaks, so Nemotron has lower voice latency.
To set a memory budget, use --memory-gb:
python3 run.py start --profile auto --memory-gb 5.5 --yesThe recommended Jetson setup runs the voice stack on the Jetson and the browser
interface on a laptop. This gives the browser a localhost address for
microphone access.
The Jetson setup needs approximately 29 GB of free disk space. The first build compiles native components, so it takes longer than later builds.
On the Jetson, find its LAN address:
ip -4 -brief addressCreate .env.local in the repository root. Replace the example address with
the Jetson address:
LIVEKIT_URL=ws://192.168.1.40:7880
LIVEKIT_NODE_IP=192.168.1.40
MANAGE_LIVEKIT=1The laptop needs these ports on the Jetson:
| Port | Protocol | Use |
|---|---|---|
8080 |
TCP | Connection details and status |
7880 |
TCP | LiveKit connection |
7881 |
TCP | WebRTC fallback media |
7882 |
UDP | WebRTC media |
If UFW is active, permit only the local subnet. Replace the example subnet with your local subnet:
sudo ufw status
sudo ufw allow proto tcp from 192.168.1.0/24 to any port 7880,7881,8080 comment 'local voice ai'
sudo ufw allow proto udp from 192.168.1.0/24 to any port 7882 comment 'local voice ai media'CAUTION: Do not expose these ports to the public internet. The default service uses development credentials.
python3 run.py start --profile auto --memory-gb 5.5 --yesWait until the launcher reports that all services are ready.
Install Node.js 20 on the laptop. Then run:
git clone https://github.com/ShayneP/local-voice-ai.git
cd local-voice-ai
corepack enable
python3 run.py client --server 192.168.1.40Open http://localhost:3000. The client installs its frontend packages on the first start.
| Command | Purpose |
|---|---|
python3 run.py |
Configure and start the application |
python3 run.py configure |
Select a different profile |
python3 run.py plan |
Show the selected runtime and models |
python3 run.py status |
Show service readiness |
python3 run.py logs |
Follow the application logs |
python3 run.py down |
Stop the Docker application |
python3 run.py client --server <host> |
Run the interface for a remote server |
The launcher saves the selected profile in .local-voice-ai.toml. This file is
local to the device and is not committed to Git.
Put device-specific configuration in .env.local. This file overrides the
selected profile and the defaults in .env.
Common values include:
| Value | Purpose |
|---|---|
LIVEKIT_URL |
LiveKit server address |
LIVEKIT_NODE_IP |
LAN address advertised by a managed LiveKit server |
LLAMA_MODEL |
Model name used by the agent |
LLAMA_HF_REPO |
GGUF model repository and quantization |
STT_PROVIDER |
Speech engine. The default is nemotron-cpp |
STT_LANGUAGE |
Speech language. The default is en |
TTS_VOICE |
Kokoro voice name |
WAKE_WORD=1 |
Require “Hey LiveKit” before the agent listens |
WEB_PORT |
Browser interface port. The default is 8080 |
See .env for the complete list.
Set a remote base URL to replace one local service. The supervisor does not start the matching local process.
| Service | Configuration |
|---|---|
| LiveKit Cloud | LIVEKIT_URL, LIVEKIT_API_KEY, LIVEKIT_API_SECRET |
| Language model | LLAMA_BASE_URL, LLAMA_MODEL, LLAMA_API_KEY |
| Speech recognition | STT_BASE_URL, STT_MODEL, STT_API_KEY |
| Speech generation | TTS_BASE_URL, TTS_API_KEY |
Store API keys in .env.local. Do not commit this file.
The startup value is the model cache size on disk. It is not the memory used by the process.
From the laptop, request the Jetson status:
curl -fsS http://192.168.1.40:8080/api/status | python3 -m json.toolIf this command times out, make sure that the firewall permits the laptop subnet.
Make sure that UDP port 7882 is open between the laptop and the Jetson.
Show the current status and logs:
python3 run.py status
python3 run.py logsLocal development needs Python 3.11–3.13, uv, Node.js 20, pnpm,
livekit-server, and llama-server.
Install the Python environment:
uv sync --extra ml --extra dev
.venv/bin/python -m local_voice_ai.agent download-filesStart the application:
.venv/bin/python -m local_voice_ai serveThis reads .env.local, then the saved profile in .local-voice-ai.toml, then
.env, so it starts with the same settings python3 run.py would use.
serve binds the web port to every interface, but it tells browsers to connect
to LiveKit on loopback, which no other machine can reach. Name the address they
should use instead:
# .env.local
LIVEKIT_PUBLIC_URL=ws://192.168.1.40:7880That is the only variable needed: the ICE address follows it, and LiveKit is
still started here because LIVEKIT_URL remains on loopback. Set LIVEKIT_URL
itself only to use a LiveKit you run elsewhere, such as LiveKit Cloud.
serve does not host the web interface, so connect from the other computer
with python3 run.py client --server 192.168.1.40, which needs Node.js and
pnpm there but not Docker.
If you change the frontend, start its development server in another terminal:
corepack enable
pnpm --dir frontend install --frozen-lockfile
pnpm --dir frontend devRun the automated tests:
.venv/bin/python -m pytest -q
pnpm --dir frontend buildThe default configuration is for local development and trusted private networks. It does not provide authentication for local model endpoints.
- Keep
.env.localout of Git. - Limit firewall rules to the local subnet.
- Do not publish the LiveKit or model ports directly to the internet.
- Use authentication and TLS before you expose the application through a public service.
- LiveKit
- LiveKit Agents
- NVIDIA Nemotron Speech
- NVIDIA Nemotron 3.5 ASR
- NVIDIA NeMo-Speech.cpp
- llama.cpp
- Gemma 4
- Kokoro
- Kokoro ONNX
- faster-whisper
Questions and feature requests are welcome through GitHub Issues.
