Skip to content

docs/planner: scoped bm25 through a traversal is ~15x slower than global bm25 on a 2.26M-row type #758

Description

@ragnorc

Summary

On a local store with 2,264,574 Passage rows (9,288 revisions, 266 matters), after optimize:

Query Time
bm25($p.text, $q) over all passages, limit 10 0.04 s
same, scoped to one matter via $p passageOfRevision $r $r: SourceRevision { matter_number: $m } (ranked variable first) 0.65 s
bm25 over 125,931 Chunk rows, limit 10 0.02 s

This matches the documented behavior (the traversal leaves the bounded top-k window unfilled, so the engine rescans without the bound), so it is not a defect. Two suggestions: (1) a sentence in docs/user/search/index.md telling catalog authors to scope on the ranked type's own indexed property where possible, or to rank a coarser type first; (2) if cheap, let the bounded scan use a selective traversal prefilter for bm25 the way rrf does, so scoped ranking does not degrade to a full-type rescan.

Version / environment

omnigraph 0.11.0, macOS arm64, file:// store, Homebrew binary.

Related: #750 (ranking on a traversal-introduced target; the query above declares the ranked variable first, so it is the supported shape), #563 (ranked read plus join materialization, closed).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

P-mediumMedium priorityacceptedTriaged and validated; open for a PRfeatureFeature proposalperformanceCorrect result, wrong cost: work scales with the table, not with the delta or the query

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions