Skip to content

Incremental embedding: embed new messages at insert time instead of batch-only #41

Description

@splaice

Problem

The embed phase is strictly a batch operation. Every invocation of `run_pipeline` goes: split → parse (inserts rows with `embedding IS NULL`) → index → embed (Ollama, in parallel workers, fills in embeddings). That's fine for the one-shot mbox import case, but becomes awkward when:

  1. A user ingests a small incremental mbox update and wants search to work on those messages immediately.
  2. Gmail sync (Phase 7) arrives — each sync pulls a handful of new messages at a time; running a 4-worker batch embed for 20 messages is silly.

Current state

  • `src/maildb/ingest/parse.py::_process_single_chunk` inserts rows with no `embedding` column set → it's `NULL`.
  • `src/maildb/ingest/embed.py::embed_worker` picks up `embedding IS NULL` rows via `SKIP LOCKED` in batches of 50.
  • The `embed` phase runs after `parse` completes in `run_pipeline`.
  • `count_unembedded` gates whether the embed phase runs at all.

Proposed solution

Add an "inline" mode where parse workers embed each row as they insert it, and the batch embed phase becomes opportunistic (only runs if there are leftovers from skipped work or older NULL rows).

Two possible shapes:

Option A — parse worker calls EmbeddingClient directly.

Each parse worker holds its own `EmbeddingClient` and calls `embed(build_embedding_text(...))` before `INSERT`. Simpler but couples parse throughput to Ollama throughput.

Option B — parse worker queues, embed worker drains continuously.

Parse workers insert with `embedding IS NULL` (current behavior); embed workers start concurrently and drain via `SKIP LOCKED` as rows land. No pipeline ordering change needed. This is already most of how it works — the gap is that `run_pipeline` currently waits for parse to finish before starting embed.

Recommendation: Option B. Minimal structural change, keeps parse CPU-bound and embed GPU-bound, retains batch efficiency. Implementation sketch:

  • Launch embed workers inside the parse phase's `ProcessPoolExecutor`, not after it.
  • Parse workers and embed workers share the queue implicitly via the `embedding IS NULL` predicate + `SKIP LOCKED`.
  • After parse workers finish, wait for embed workers to drain, then create the HNSW index.

Add a `--inline-embed / --batch-embed` flag on `maildb ingest run` so operators can choose based on corpus size — batch is better throughput for a cold 800K import; inline is better latency for a 50-message sync.

Acceptance criteria

  • `run_pipeline` supports concurrent parse + embed (via option B or equivalent)
  • CLI flag `--inline-embed` / `--batch-embed` (default picks one based on expected corpus size or falls back to env var)
  • New integration test: ingest a small mbox with `--inline-embed` and assert the emails have non-null embeddings before `run_pipeline` returns
  • Existing batch-mode tests still pass
  • `skip_embed=True` still skips both modes
  • DESIGN.md §7.1 updated to describe both modes

Tradeoffs

  • Pro: Much better UX for incremental imports. Sets us up for Gmail sync without a separate pathway.
  • Con: More concurrency surface. Parse workers now need an Ollama handle (per-worker connection, or shared pool).
  • Risk: If Ollama is down during ingestion, inline mode means parse itself appears to stall. Mitigation: fall back to leaving `embedding=NULL` after N failed attempts and let the batch path pick it up later.

Defer trigger

This issue is low priority until Gmail sync (Phase 7) starts. For a one-shot mbox import the batch mode is fine. Open to pull forward if someone finds themselves doing many small mbox imports in a row.

References

  • `src/maildb/ingest/orchestrator.py::run_pipeline` — phase ordering
  • `src/maildb/ingest/parse.py` — where inline embedding would fit
  • `src/maildb/ingest/embed.py` — existing batch worker
  • `src/maildb/embeddings.py::build_embedding_text` — already token-aware

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions