Skip to content

[Feature/Model Request]: Support for NPU-accelerated Text-to-Speech (TTS) models #731

Description

@Seyij

Suggestion Description

I am requesting support for lightweight Text-to-Speech models (e.g., Kokoro-82M or Piper) compiled for the NPU, alongside an exposed OpenAI-compatible /v1/audio/speech endpoint.
I already use the fastflowLM whisper model for speech to text in openwebui. So an NPU accellerated text to speech model for reading out chat responses would be really useful.

Requested:
TTS Model Suppor - Introduce quantized/compiled versions of highly efficient TTS models (like Kokoro-82M or Piper) optimized for the XDNA 2 architecture.
API Integration - Expose the standard OpenAI-compatible audio generation endpoint (/v1/audio/speech) in the FastFlowLM server mode.

I'd be happy to help test versions as I have an AMD Ryzen AI 9 HX 370 (XDNA 2 / 50 TOPS).

Operating System

Windows 11, Ubuntu 24.04

GPU

NPU: AMD Ryzen AI 9 HX 370 (XDNA 2 / 50 TOPS)

ROCm Component

FastflowLM

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions