Suggestion Description
I am requesting support for lightweight Text-to-Speech models (e.g., Kokoro-82M or Piper) compiled for the NPU, alongside an exposed OpenAI-compatible /v1/audio/speech endpoint.
I already use the fastflowLM whisper model for speech to text in openwebui. So an NPU accellerated text to speech model for reading out chat responses would be really useful.
Requested:
TTS Model Suppor - Introduce quantized/compiled versions of highly efficient TTS models (like Kokoro-82M or Piper) optimized for the XDNA 2 architecture.
API Integration - Expose the standard OpenAI-compatible audio generation endpoint (/v1/audio/speech) in the FastFlowLM server mode.
I'd be happy to help test versions as I have an AMD Ryzen AI 9 HX 370 (XDNA 2 / 50 TOPS).
Operating System
Windows 11, Ubuntu 24.04
GPU
NPU: AMD Ryzen AI 9 HX 370 (XDNA 2 / 50 TOPS)
ROCm Component
FastflowLM
Suggestion Description
I am requesting support for lightweight Text-to-Speech models (e.g., Kokoro-82M or Piper) compiled for the NPU, alongside an exposed OpenAI-compatible /v1/audio/speech endpoint.
I already use the fastflowLM whisper model for speech to text in openwebui. So an NPU accellerated text to speech model for reading out chat responses would be really useful.
Requested:
TTS Model Suppor - Introduce quantized/compiled versions of highly efficient TTS models (like Kokoro-82M or Piper) optimized for the XDNA 2 architecture.
API Integration - Expose the standard OpenAI-compatible audio generation endpoint (/v1/audio/speech) in the FastFlowLM server mode.
I'd be happy to help test versions as I have an AMD Ryzen AI 9 HX 370 (XDNA 2 / 50 TOPS).
Operating System
Windows 11, Ubuntu 24.04
GPU
NPU: AMD Ryzen AI 9 HX 370 (XDNA 2 / 50 TOPS)
ROCm Component
FastflowLM