Skip to content

search: pick snippet window that covers the most distinct query terms - #16

Open
ashishbhateja wants to merge 1 commit into
mainfrom
search/snippet-best-coverage
Open

ashishbhateja wants to merge 1 commit into
mainfrom
search/snippet-best-coverage

Conversation

@ashishbhateja

Copy link
Copy Markdown
Owner

What and why

Advances issue #3 (search result quality — snippets and highlighting).

snippet() in src/search.js previously always centred the context window on the earliest occurrence of any matched term. For a multi-term query this produces a poor result when the terms cluster together only later in the text: the reader sees a snippet that shows just one of their search words.

Before (search('integral yoga') where "integral" and "yoga" only co-appear near line 6 of the body, but "yoga" also appears alone in line 1):

Yoga is a path of transformation. …

After:

… In integral yoga, silence and steadiness unite.

How

snippet() now collects every occurrence of every matched term in the text, then iterates over those positions and picks the one whose ±half window contains the highest count of distinct terms. Single-term behaviour is mathematically identical to before; the change only kicks in when two or more terms are matched.

The algorithm is O(occurrences²) which is negligible at magazine scale (a handful of terms, each appearing a few times in a few hundred characters of body text).

Changes

File Change
src/search.js snippet() — best-coverage window selection
scripts/smoke.mjs New assertion covering the multi-term case

Tests

20 checks passed.   ← was 19

The new check (snippet centres on the window covering the most distinct terms) would have failed against the previous implementation, confirming the regression guard is real.


Generated by Claude Code

The previous snippet() always centred the context window on the earliest
occurrence of any matched term.  For a multi-term query where the terms
cluster together only later in the text (e.g. one isolated early hit and
both terms appearing together further on), this produced a snippet showing
only one of the matched terms — a poor result for the reader.

The new implementation collects every occurrence of every matched term,
then picks the candidate position whose window (±half of contextChars)
contains the highest count of distinct terms.  Single-term behaviour is
unchanged; for multi-term queries the displayed snippet is now far more
likely to show the words the reader actually searched for, advancing the
search-result-quality goal tracked in issue #3.

A new smoke test captures the exact failing case: 'silence' appears early
(isolated) and again near 'integral' much later; the snippet must now
include both.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants