Skip to content

Performance knowledge: keep low-selectivity fields out of key columns - #226

Open
Miljan Milosavljević (miljance) wants to merge 1 commit into
microsoft:mainfrom
miljance:perf/low-selectivity-key-columns
Open

Miljan Milosavljević (miljance) wants to merge 1 commit into
microsoft:mainfrom
miljance:perf/low-selectivity-key-columns

Conversation

@miljance

Copy link
Copy Markdown

Fixes #224

Mistake prevented

design-covering-keys-from-read-pattern.md says which predicates a key should support, but no article says which predicates should stay out of a key. Agents applying the guidance turn every filtered field into a key column, including Boolean, small Option/Enum and status fields (Processed, Blocked, Status), sometimes as the leading column. Such a column:

  • widens every index row;
  • adds write cost, most of all when the process the key serves rewrites the field (filter Processed = false, then ModifyAll(Processed, true));
  • as a leading column, gives no seek to any read that does not filter it with equality.

Change

  • New article microsoft/knowledge/performance/keep-low-selectivity-fields-out-of-key-columns.md, with a .bad.al/.good.al pair. The samples use the same reader code, and only the key differs.
  • The article lists the legitimate exceptions so reviewers do not report false positives:
    • SIFT keys, which must contain every filtered field (cross-referenced to flowfield-source-key-needs-sumindexfields.md);
    • read-heavy covering through IncludedFields, when the process does not update the field;
    • skewed distributions filtered on the rare value;
    • unique keys.
  • al-performance-review.md gains one worklist sentence that loads the article when an added or changed key, or its IncludedFields, contains a Boolean, Option, Enum or status field. The existing rule that a changed key alone is not a finding still applies.

Evidence

  • The article links Microsoft Learn pages on table keys and their costs, on keys and performance, on SIFT, and the SQL Server index design guide on column order and selectivity.
  • The claim that SQL Server cannot skip-scan past an unfiltered leading column is general SQL Server behaviour. The article states it in those terms, not as a BC platform guarantee.

Versions and domain

  • bc-version: [all]. IncludedFields requires runtime 8.0 (BC 19).
  • The performance domain is Microsoft-owned, and this is general key behaviour rather than an organization policy, so the article belongs in the Microsoft layer.

Checks

  • validate_frontmatter.py: 0 errors. Its 2 warnings come from pre-existing files this PR does not touch.
  • Test-KnowledgeIndex.ps1, Test-ReviewFixtures.ps1 and Test-ReviewContract.ps1 all pass.
  • No model-based evaluation was run.

🤖 Generated with Claude Code

The covering-key guidance says which predicates a key should support but
not which ones to leave out, so agents turn every filtered Boolean,
Option/Enum or status field into a key column. Add an article that keeps
such fields out of key columns and out of IncludedFields when the serving
process rewrites them, with explicit exceptions for SIFT, read-heavy
covering, skewed data and unique keys, plus a good/bad sample pair. The
performance review worklists it when a changed key contains such a field.

Refs microsoft#224

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Performance] Covering-key guidance needs a counterweight: keep low-selectivity columns out of key columns

1 participant