Local Desk AI Assistant is a local-first touchscreen assistant for an ESP32-S3 7-inch 800x480 RGB panel. It is created and maintained by lepczynski-cloud.
Repository:
https://github.com/lepczynski-cloud/crowpanel_AI_assistant
The panel sends prompts only to services on the local network:
- Ollama generates text answers.
- The optional Voice server performs local speech recognition with faster-whisper.
- Optional Piper TTS turns the answer into a local WAV file for panel playback.
No cloud AI API, vendor account, telemetry service, or proprietary factory application is required.
Version 0.5.7 keeps the hardware-confirmed stable renderer from 0.5.5 and fixes the remaining Windows Voice launcher failure.
The 0.5.6 launcher correctly prepared a PCM-only faster-whisper runtime, but then ran an unnecessary pip uninstall av command. When PyAV was already absent, pip returned success while printing a warning to stderr. Windows PowerShell 5.1 could promote that harmless warning to NativeCommandError because the script used strict error handling.
The 0.5.7 launcher:
- no longer runs the unnecessary PyAV uninstall command,
- keeps fresh environments PyAV-free by installing faster-whisper with
--no-deps, - updates a partial 0.5.6 environment in place instead of downloading everything again,
- evaluates Python and pip commands by their real process exit code,
- preserves strict error handling for PowerShell operations,
- verifies the complete STT/TTS runtime before starting Uvicorn,
- writes the environment schema marker only after verification succeeds,
- keeps the one-double-click startup workflow.
No Windows security-policy exception is required for the PyAV issue.
- modern dark/cyan touch interface with rounded controls,
- stable low-cost renderer verified on the target panel,
- animated startup diagnostics without full-screen interaction flicker,
- camera-safe Home screen with version, Wi-Fi, Ollama, and optional Voice status,
- on-screen keyboard for typed prompts and settings,
- Quick prompts,
- editable Wi-Fi name and password stored in ESP32 NVS,
- editable local server host, Ollama port, Voice port, and Ollama model,
- English, Polish, and automatic reply-language modes,
- English as the default recognition and answer language,
- direct local Ollama
/api/tagsand/api/generateintegration, - six-second microphone recording with local STT retries and diagnostics,
- optional spoken replies through Piper,
- one-click Windows Voice launcher and optional desktop shortcut,
- CLion/PlatformIO project layout,
- no IP addresses on Home or the boot screen; detailed addresses remain in Settings.
ESP32-S3 touch panel
|
+-- local HTTP --> Ollama :11434
| /api/tags
| /api/generate
|
+-- optional ----> Voice server :8080
/health
/api/stt -> PCM decoder -> faster-whisper
/api/tts -> Piper -> WAV
Typed and Quick prompts go directly to Ollama. Voice mode records PCM WAV on the panel, sends it to the optional Voice server, forwards the recognized text to Ollama, and optionally plays a Piper-generated WAV through the panel speaker.
Use the normal build with the rear switches set to:
S1 = 0
S0 = 0
Mode = MIC & SPK
PlatformIO environment = crowpanel_v12_local_ai
The working display baseline is preserved:
Resolution: 800 x 480
Pixel clock: 16 MHz
Horizontal timing: 40 / 48 / 13
Vertical timing: 1 / 31 / 13
Microphone: standard I2S
The PDM targets are diagnostic alternatives for other hardware revisions. Do not use them when the normal microphone works in the 0/0 switch position.
- compatible ESP32-S3 7-inch 800x480 RGB touch panel using this project's pinout,
- 16 MB flash and PSRAM,
- CLion with PlatformIO or PlatformIO Core,
- Git available in
PATH, - USB serial driver when required by Windows.
- Windows 10/11 for the included one-click Voice launcher,
- Python 3.10-3.13 for optional Voice,
- Ollama running on the same LAN,
- at least one local Ollama model.
Example:
ollama pull smollm2:latest
ollama listOllama must accept LAN connections on port 11434. From another device on the same network, this address should return JSON:
http://YOUR-PC-IP:11434/api/tags
Copy:
include/secrets.example.h
to:
include/secrets.h
Set first-boot defaults:
#define WIFI_SSID "YOUR_WIFI_NAME"
#define WIFI_PASSWORD "YOUR_WIFI_PASSWORD"
#define DEFAULT_AI_HOST "192.168.1.100"
#define DEFAULT_OLLAMA_PORT 11434
#define DEFAULT_VOICE_PORT 8080
#define DEFAULT_OLLAMA_MODEL "smollm2:latest"include/secrets.h is ignored by Git. Real Wi-Fi values are copied to ESP32 NVS only when no saved values exist. Wi-Fi, server host, ports, model, language, and spoken-reply state can later be changed on the touchscreen.
Open platformio.ini in CLion and upload:
pio run -e crowpanel_v12_local_ai -t upload
pio device monitor -b 115200Do not erase NVS during a normal update. An erase removes all saved settings.
Home shows only camera-safe information:
- application version,
- Wi-Fi state,
- Ollama state,
- optional Voice state,
- readiness message.
It provides Type, Voice, Quick, Settings, and spoken-reply controls.
Settings uses three tabs:
- Wi-Fi — network name, masked password, reconnect action, and panel network details.
- Servers — local server host, Ollama port, Voice port, and current service state.
- Assistant — Ollama model, reply language, spoken reply, and microphone profile.
The bottom bar contains only Health, About, and Back.
The default is English. The selector offers:
English | Polish | Auto
- English recognizes and answers in English.
- Polish recognizes and answers in Polish.
- Auto detects the spoken language and asks Ollama to answer in the request language.
Double-click:
START_VOICE_SERVER.cmd
On first use, the launcher creates server/.venv, installs dependencies, downloads the local Whisper model and English Piper voice, validates the runtime, and starts port 8080.
When updating directly over a 0.5.6 project folder, the launcher repairs the partial private environment in place. When using a new folder, copy only server/models from the previous folder to avoid downloading the models again; do not copy server/.venv unless you intend to migrate it in place.
Wait for:
[Local Desk AI] faster-whisper is ready.
[Local Desk AI] Piper TTS is ready.
INFO: Application startup complete.
The Voice server can be started before or after the panel. Every tap on Voice performs a fresh health check, so a panel reboot is not required.
To create a desktop shortcut, run once:
CREATE_VOICE_SERVER_SHORTCUT.cmd
The server is not installed as a Windows service. Close its window or press Ctrl+C to stop it and release memory.
The panel sends one known audio format: 16-bit mono PCM WAV. Version 0.5.7 intentionally omits PyAV in fresh environments and uses a small in-memory compatibility placeholder only so faster-whisper can import. The generic PyAV media decoder is never executed.
This specifically avoids errors such as:
ImportError: DLL load failed while importing filter: application control policy blocked this file
Do not disable Windows App Control for this project. If another required native module is blocked, the launcher reports the exact failed import instead of repeatedly reinstalling everything.
The most recent microphone request is stored locally as:
server/debug/last_stt.wav
This file is ignored by Git. Play it when recognition is unreliable:
- clear speech with empty text usually indicates language/model tuning,
- very quiet audio suggests moving closer to the microphone,
- silence suggests the wrong hardware switch or microphone environment,
- distorted audio suggests the wrong I2S/PDM profile.
Health endpoints:
http://127.0.0.1:8080/health
http://YOUR-PC-IP:8080/health
platformio.ini PlatformIO environments
src/main.cpp main firmware
src/display_smoke_test.cpp display-only diagnostic target
include/BoardConfig.h board pins and audio profile
include/CrowPanelDisplay.h working RGB/touch driver configuration
include/DefaultConfig.h first-boot defaults
include/secrets.example.h safe template; copy to secrets.h locally
START_VOICE_SERVER.cmd one-click Windows Voice launcher
CREATE_VOICE_SERVER_SHORTCUT.cmd optional desktop shortcut creator
server/ optional local STT/TTS companion
docs/ focused setup and troubleshooting guides
CHANGELOG.md release history
THIRD_PARTY_NOTICES.md dependency and model notices
The repository intentionally excludes:
include/secrets.h
.pio/
server/.venv/
server/models/
server/debug/
*.wav
- Upload with CLion and PlatformIO
- Voice setup and troubleshooting
- Local network checklist
- Wi-Fi configuration
- Hardware audio notes
- Black-screen recovery
- Release notes
Firmware, documentation, and root helper scripts are released under the MIT License; see LICENSE.
The optional server/ companion is released under GPL-3.0-or-later because it integrates the current GPL-licensed Piper package; see server/LICENSE.
Downloaded libraries and model files remain under their upstream licenses. See THIRD_PARTY_NOTICES.md.