Skip to content

Latest commit

 

History

35 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

douchi

A privacy-first CLI tool for generating music playlists from a natural language description of what you want to hear.

douchi playlist "90s NYC hip-hop" -o tonight.m3u
douchi playlist "mellow late-night jazz"
douchi playlist "albums similar to OK Computer"
douchi playlist "melancholy British albums similar to Kind of Blue"

No audio is ever processed by an LLM — the library is represented as a structured database and queries are translated to SQL. Supports local inference via Ollama or cloud inference via the Claude API.


How it works

Standard audio file tags (genre, artist, year) are insufficient for context-aware queries like "Scandinavian black metal" or "upbeat 80s synth-pop". douchi solves this with a three-step design:

1. Ingest — Scan your music library, read audio tags, and populate a local SQLite database with your albums and tracks.

2. Enrich — Perform a one-time pass against the Discogs API to augment each album with richer metadata: granular sub-genres and styles, geographic origin, label, and catalog number. Also generates embedding vectors for each album (used for similarity queries).

3. Query — An LLM translates your natural language query into a SQL statement, which is executed against the local database and written to an .m3u playlist file. The LLM can run locally via Ollama or in the cloud via the Claude API.

The LLM only ever sees your query and the database schema — never your audio files or personal data.


Architecture

Text-to-SQL pipeline

User query
    │
    ▼
 Router ──────────────────────────────────────────────┐
    │                                                  │
    │ filtering query                    similarity / hybrid query
    │ ("90s jazz")                       ("albums similar to X")
    ▼                                                  │
 Genre/style resolver (small LLM)                      │
 identifies relevant DB values                         │
    │                                                  │
    ▼                                                  ▼
 SQL generator (large LLM)              Vector search (sqlite-vec)
 produces SELECT statement              finds nearest neighbours
    │                                                  │
    └──────────────────────┬───────────────────────────┘
                           ▼
                    SQLite database
                           │
                           ▼
                      .m3u playlist

Filtering queries go through two LLM steps. First, a lightweight resolver model (qwen2.5:3b by default) identifies which genre and style names from the database are relevant to the query — using the model's music domain knowledge to handle semantic matches (e.g. "afrobeats" resolves to styles like Amapiano and Gqom) and distinguish near-homonyms (Afrobeat vs Afrobeats). Then, the larger SQL generation model receives the resolved names along with the database schema, few-shot examples, and the user's query, and outputs a SQL SELECT statement which is validated and executed locally.

Similarity queries bypass the LLM entirely: the reference album's embedding vector is looked up, and the nearest neighbours in vector space are returned using sqlite-vec.

Hybrid queries ("British jazz similar to Kind of Blue") combine both: similarity search narrows the candidate set, then the LLM filters within it.

Database

A local SQLite database stores:

Table Contents
artists, albums, tracks Core library data from audio tags
genres, styles Discogs taxonomy (many-to-many via junction tables)
album_tags Extensible EAV table for mood/vibe/theme attributes
album_embeddings 768-dimensional vectors for similarity search
discogs_cache Raw Discogs API responses (for re-processing without re-fetching)
enrichment_log Audit trail of enrichment attempts per album

Requirements

  • Python 3.11+
  • Ollama running locally (for local inference and embeddings)
  • A Discogs account with a free API token (for the one-time enrichment pass)
  • Optional: An Anthropic API key (for Claude-based cloud inference)

LLM backend

douchi supports two LLM backends:

  • Ollama (default) — fully offline, local inference. Best for privacy; requires sufficient RAM for the chosen model.
  • Claude — cloud inference via the Anthropic API. Much faster on constrained hardware (seconds vs. minutes). Requires an API key and internet connection. Album metadata is sent to the API.

Even when using Claude for chat-based tasks (SQL generation, genre resolution, tagging), Ollama is still required for embedding generation since Claude does not offer an embedding API.

Ollama models

Pull these before running (all required even if using Claude backend):

ollama pull nomic-embed-text     # Embedding generation (required, all hardware)

If using Ollama as your LLM backend (the default), also pull:

ollama pull qwen2.5-coder:32b    # SQL generation (M4 / 36GB+)
ollama pull qwen2.5:3b           # Genre/style resolution (all hardware)

On Intel hardware or with less than ~20GB of available RAM, use a smaller SQL generation model:

ollama pull qwen2.5-coder:7b

See Hardware notes for details.


Installation

pip install douchi
# or, for isolated installs:
pipx install douchi

To use the Claude backend:

pip install "douchi[claude]"

To install from source:

git clone https://github.com/nathancrtr/douchi
cd douchi
pip install -e ".[dev]"
# For Claude support:
pip install -e ".[dev,claude]"

Configuration

Run the interactive setup wizard on first use — you do not need to edit any files manually:

douchi init

The wizard prompts for your music library path, Discogs API token, and Ollama model, then writes ~/.config/douchi/config.toml for you. Press Enter to accept the displayed default for any field.

Verifying your environment

douchi check

Checks that all dependencies are reachable before you start a long ingest or enrichment run:

[ok] config          ~/.config/douchi/config.toml
[ok] library         /Volumes/Music (12 847 audio files)
[ok] database        ~/.local/share/douchi/library.db
[ok] ollama          http://localhost:11434 — qwen2.5-coder:32b loaded
[ok] resolver model  qwen2.5:3b loaded
[ok] embed model     nomic-embed-text loaded
[ok] discogs         token present
[ok] sqlite-vec      loaded (vector search enabled)

Any item that cannot be reached is printed as [fail] with a short explanation and a suggested fix.

Manual configuration (reference)

douchi init writes a TOML file at ~/.config/douchi/config.toml. You can edit it directly if you need to change a value after initial setup:

[library]
path = "/path/to/your/music"

[database]
path = "~/.local/share/douchi/library.db"

[ollama]
host = "http://localhost:11434"
embedding_model = "nomic-embed-text"

[llm]
backend = "ollama"                           # "ollama" or "claude"
model = "qwen2.5-coder:32b"                 # Ollama model for SQL generation
resolver_model = "qwen2.5:3b"               # Ollama model for genre resolution
claude_model = "claude-sonnet-4-20250514"          # Claude model for SQL generation
claude_resolver_model = "claude-haiku-4-5-20251001"  # Claude model for resolution/tagging

[discogs]
token = ""   # get a free token at discogs.com/settings/developers

All values can also be overridden with environment variables:

Variable Overrides
DOUCHI_LIBRARY_PATH library.path
DOUCHI_DB_PATH database.path
DOUCHI_OLLAMA_HOST ollama.host
DOUCHI_OLLAMA_MODEL llm.model
DOUCHI_LLM_BACKEND llm.backend
ANTHROPIC_API_KEY Claude API authentication
DISCOGS_TOKEN discogs.token

Usage

First-time setup

# 1. Run the setup wizard (creates config, initialises the database)
douchi init

# 2. Verify everything is reachable
douchi check

# 3. Scan your library and read audio tags (uses path from config)
douchi ingest

# Re-scan all files, not just new ones
douchi ingest --force

# 4. Enrich with Discogs metadata, embeddings, and mood/vibe tags
#    (needs internet for Discogs; Ollama for tags; one-time only)
#    Token is read from config; --discogs-token overrides it
douchi enrich

Enrichment runs three passes: Discogs metadata, embedding generation, and LLM-powered mood/vibe/theme tagging. After enrichment, all commands run fully offline.

Generating playlists

# Write to a file
douchi playlist "90s NYC hip-hop" -o hip-hop.m3u

# Print to stdout (pipe to a player or script)
douchi playlist "mellow late-night jazz"

# Similarity search
douchi playlist "albums similar to OK Computer" -o similar.m3u

# Hybrid: similarity + filtering
douchi playlist "melancholy British albums similar to Kind of Blue"

# Inspect the generated SQL before executing
douchi playlist "Scandinavian black metal" --show-sql

# Show which genres/styles the resolver identified
douchi playlist "2020s afrobeats" --show-resolved

# Generate SQL without executing (for debugging)
douchi playlist "upbeat 80s synth-pop" --dry-run

# Override the model for a single query
douchi playlist "jazz" --model qwen2.5-coder:7b

Browsing the library

# Find similar albums (embedding-based)
douchi similar "OK Computer"
douchi similar "Kind of Blue" -n 20

# List available genres and styles (useful for refining queries)
douchi genres
douchi styles

# Database statistics and enrichment coverage
douchi stats

# Check that all dependencies are reachable
douchi check

Normalizing genres and styles

Genre and style names are automatically normalized during ingest and Discogs enrichment (case, punctuation, and alias mapping). To clean up existing data that was ingested before normalization was added:

# Preview what would change
douchi normalize --dry-run

# Apply normalization
douchi normalize

Enrichment options

The enrichment step is resumable. If it's interrupted, re-run with --skip-matched to process only albums not yet enriched:

douchi enrich --skip-matched

To re-enrich a single album by its database ID:

douchi enrich --album-id 42

To re-fetch Discogs data for already-matched albums (e.g. after a bug fix that improved data extraction):

douchi enrich --refresh

To skip the LLM tagging pass (useful if Ollama is not running):

douchi enrich --skip-tags

To regenerate embeddings without re-fetching Discogs data or re-tagging (e.g. after changing the embedding model or text format):

douchi enrich --embeddings-only

Query tips

douchi routes queries automatically:

  • Filtering: genre, decade, geography, mood/vibe, label, style

    90s UK post-punk
    electronic albums on Warp Records
    mellow late-night ambient music
    warm summery vibes
    dark cinematic albums
    Chicago blues from the 1960s
    
  • Similarity: reference an album or artist in your library

    albums similar to OK Computer
    music that sounds like Boards of Canada
    something in the style of Kind of Blue
    
  • Hybrid: combine both

    melancholy British albums similar to OK Computer
    jazz albums from the 1960s similar to Kind of Blue
    

If a query returns zero results:

  • Check what genres and styles are in your database: douchi genres, douchi styles
  • Use --show-resolved to see which genres/styles the resolver identified
  • Use --show-sql to inspect the generated SQL
  • Rephrase using exact genre/style names from douchi genres

Hardware notes

The SQL generation model is the primary hardware constraint. Model files must fit in RAM (or VRAM on discrete GPU systems) to run at usable speeds.

Apple Silicon (M-series)

The recommended platform. Ollama uses Metal GPU acceleration and Apple Silicon's unified memory architecture means the full RAM pool is available to the model — a 36GB system can run a 32B parameter model at Q4 quantization (~18–20GB). Expect 2–5 seconds per query with a 32B model.

Systems with a discrete GPU (NVIDIA/AMD)

Ollama will use CUDA or ROCm automatically if available. The model must fit in VRAM for full GPU acceleration; if it overflows into system RAM, inference will be slower. A GPU with 16GB+ VRAM can run a 14–32B model comfortably.

CPU-only (no GPU, or integrated graphics)

Inference runs on CPU, which is significantly slower — expect 30–90 seconds per query depending on the model size and CPU. Use the smallest model that gives acceptable quality for your use case.

Choosing a model by available RAM

Available RAM Recommended model Notes
32GB+ qwen2.5-coder:32b Best quality; ~18–20GB at Q4
16–24GB qwen2.5-coder:14b Good quality; ~8–9GB at Q4
8–16GB qwen2.5-coder:7b Adequate for most queries
< 8GB qwen2.5-coder:3b Fast but reduced accuracy

qwen2.5-coder is recommended over general-purpose models because it is fine-tuned for code and SQL tasks, meaning smaller variants perform better on this workload than their parameter count would suggest.

Regardless of SQL model size, the resolver model (qwen2.5:3b, ~2GB) and nomic-embed-text (~274MB, used for similarity search) are small and run quickly on any hardware.

On lower-end hardware:

  • Speed degrades significantly; CPU-only inference can make queries feel slow
  • Quality is somewhat reduced on complex multi-table queries (style + country
    • decade combined), though simple queries remain reliable

Development

# Install with dev dependencies
pip install -e ".[dev]"

# Run the test suite
python -m pytest

# Run with coverage
python -m pytest --cov=douchi

# Lint
ruff check src/

Tests run entirely offline — no Ollama or Discogs connection required. Integration tests (requiring a live Ollama instance) are not yet implemented and should be marked @pytest.mark.integration when added.


Project structure

src/douchi/
├── cli.py              Entry point; all CLI commands
├── config.py           Config loading (TOML + env vars)
├── llm.py              LLM backend abstraction (Ollama + Claude)
├── db/
│   ├── connection.py   SQLite connection factory; loads sqlite-vec
│   └── schema.py       DDL; apply_schema()
├── ingest/
│   ├── scanner.py      Filesystem walker; discovers album folders
│   └── tagger.py       Audio tag reader (FLAC, MP3, M4A, ALAC)
├── enrich/
│   ├── data/           JSON data files (aliases, vocabulary, junk lists)
│   ├── discogs.py      Discogs API matching and enrichment
│   ├── embeddings.py   Embedding generation and storage
│   ├── normalize.py    Genre/style name normalization and alias mapping
│   └── tagger.py       LLM-powered mood/vibe/theme tagging
├── query/
│   ├── genre_resolver.py  LLM-based genre/style resolution for queries
│   ├── router.py          Query classifier (SQL / similarity / hybrid)
│   ├── similarity.py      Vector search with sqlite-vec and Python fallback
│   └── text_to_sql.py     Prompt assembly, SQL extraction, validation
├── prompts/
│   ├── system.py       System prompt template
│   └── examples.py     Few-shot SQL examples
└── playlist/
    └── m3u.py          Extended M3U writer

About

Create music playlists with natural language

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages