Skip to content

fix: give each query lexeme its own BM25 candidate budget - #279

Open
nuemaan wants to merge 1 commit into
Ontos-AI:mainfrom
nuemaan:fix/nuemaan/bm25-per-lexeme-candidate-budget
Open

fix: give each query lexeme its own BM25 candidate budget#279
nuemaan wants to merge 1 commit into
Ontos-AI:mainfrom
nuemaan:fix/nuemaan/bm25-per-lexeme-candidate-budget

Conversation

@nuemaan

@nuemaan nuemaan commented Aug 13, 2026

Copy link
Copy Markdown

Summary

Measured on 5001 chunks where 5000 densely repeat a common term and one holds a rare term, querying data zebra:

ts_rank_cd position of the rare chunk:   5001 of 5001
rank_rows_by_bm25 position:                 1 of 5001

candidate pool, global ordering, limit 2000:  2000 rows, rare chunk absent
candidate pool, per-lexeme budget:            1001 rows, rare chunk present

The per-lexeme pool is smaller and still keeps the match BM25 wants. Cost is one GIN probe per lexeme rather than one overall, and _MAX_FTS_QUERY_TOKENS already caps that at 50.

Verification

  • uv run pytest apps/api/tests/unit/test_bm25_channel_tsquery.py apps/api/tests/contract/test_bm25_fts_prefilter_contract.py — 16 passed against real Postgres 16
  • uv run pytest packages/shared-python/shared/tests/test_retrieval_search_channels.py — 4 passed
  • uvx ruff check packages/shared-python apps/api/tests — clean
  • Reverted the channel change and re-ran the new regression test to confirm it fails without the fix: AssertionError: assert 'common-2' == 'rare-zebra'

test_content_channel_uses_bounded_or_fts_after_scope_filters asserted on the literal LIMIT :fts_candidate_limit as a position marker for the candidate bound. That clause is now LIMIT lb.per_lexeme_limit, so the assertions were repointed. The intent from #244, that scope filters land ahead of any candidate bound so exclusions cannot consume the budget, is unchanged and still asserted.

Not tested: production-scale corpora, and lateral-probe cost at the 50 lexeme ceiling on a large namespace. Both need a dataset I do not have locally.

Deployment Notes

  • No new environment variables. RETRIEVAL_POSTGRES_FTS_CANDIDATE_LIMIT keeps its meaning as the total budget, now spent per lexeme.
  • No database migrations.
  • Backwards compatible. The full-scan fallback is untouched, so rolling back is a code-only revert.

Checklist

  • Tests were added or updated when behavior changed
  • Public docs, examples, or OpenAPI contracts were updated when needed
  • Database migrations are idempotent and safe to deploy
  • Logs, errors, and validation paths avoid leaking secrets or user data
  • The pull request description explains any breaking or user-visible change

The FTS prefilter bounded candidates with a single global ts_rank_cd
ordering. ts_rank_cd scores term density inside a chunk and ignores how
rare a term is across the corpus, while BM25 weights rare terms heavily.
A short chunk holding the one rare term in a query therefore sorts near
the bottom of that ordering and is truncated first, even though BM25
ranks it top.

Measured on 5001 chunks where 5000 densely repeat a common term and one
holds a rare term: the rare chunk ranks 5001 of 5001 under ts_rank_cd
and 1 of 5001 under rank_rows_by_bm25, so a 2000 candidate limit dropped
the best match before BM25 ran.

Each lexeme now draws from its own share of the budget through a lateral
join, so a lexeme matching few chunks always contributes them. A floor
keeps many-lexeme queries from dividing the budget into slivers. The
same corpus now yields 1001 candidates including the rare chunk, fewer
rows than the old path loaded while keeping the match that matters.

Also logs a warning when the pool saturates. The debug line reported
candidates == limit whether the corpus held exactly that many or far
more, so silent truncation looked identical to a healthy query.

The bounded-prefilter test asserted on the literal LIMIT clause as a
position marker. Its intent, that scope filters land ahead of any
candidate bound, is unchanged and now asserts against the per-lexeme
clause.

Closes Ontos-AI#278
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Classic BM25 FTS prefilter can truncate the top BM25 match in large namespaces

1 participant