Local-first realtime Russian voice avatar with STT, Qwen LLM, TTS, Audio2Face, LiveKit, and VRM rendering.
-
Updated
Sep 11, 2026 - Python
Local-first realtime Russian voice avatar with STT, Qwen LLM, TTS, Audio2Face, LiveKit, and VRM rendering.
Agentic framework for unified full-stack observability and cloud LLMs
Drop-in long-term memory for LLM apps — application-layer Memory Caching (Behrouz et al. 2026, arXiv:2602.24281). 14.6x fewer tokens at higher accuracy. 3-line API.
On-prem GPU inference infrastructure — infra-as-code, hardening baseline, topology, and ADRs for private model serving.
Synthetic industrial-manual RAG evaluation pack for revisions, wrong-model traps, access boundaries, abstention, and citations
Health-aware, cost-optimized tiered LLM inference routing — reference architecture, ADRs, and a runnable implementation.
To associate your repository with the on-prem-ai topic, visit your repo's landing page and select "manage topics."