Repository navigation
Fix lake cold startup, restart cache persistence, and pgwire deadlines - #1019
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Cold remote lake searches were opening segment readers serially and scanning every dictionary block to populate diagnostic statistics. Disk-cache writes were also rejected on macOS because capacity freshness compared native monotonic timestamps with Io awake timestamps. Pgwire passed the same mismatched absolute clock through native catalog/storage boundaries, causing six lake SQL tests to return QueryCanceled.
Defer range-backed dictionary diagnostics, admit readers in bounded parallel jobs and transfer them to the writer without reopening, overlap bounded required header reads, and fetch complete immutable artifacts with one bounded GET plus exact length/SHA-256 verification. Normalize cache capacity and pgwire deadlines into the native clock domain. Catalog reader protocols 24–33 remain parseable for reconciliation; serving still requires the current protocol.
Raw BigQuery HN export in antfly-dev-01, local Debug binary, 10,000 rows:
The new runs copied catalog state into directories with no cache files before startup; the harness SQL count precedes text search. Empty-cache search still pays required remote metadata/posting/hydration I/O. These are smoke measurements, not archive-scale or deployment-region benchmarks. Update the Hacker News example and qualification record with the results.
Validation: