Problem
CacheKey holds nothing but the relative path. The key carries
no storage identity, while the cache state behind it is shared:
StorageCache (foyer, the bytes) — one per process, shared by the Iceberg path
and the object-store path;
CacheLayer (stat cache, KeyLocks) — one per OperatorRegistry, cloned into
every operator it builds;
PrefetchLayer: seen: DashSet<Arc<str>> and
InFlightTracker.pending: DashMap<Arc<str>, _>— keyed by the
bare path as well.
With two or more buckets (warehouse + catalog is the standard deployment) this
means:
- a read of
tables/logs/metadata/v1.metadata.json in the second bucket can be
served the first bucket's bytes;
- a stat (HEAD) against the second bucket returns the first bucket's metadata;
- a delete or write in one bucket invalidates the other bucket's stat entry;
- prefetch treats an object as already seen because of a different bucket and
skips the read-ahead.
Risk: wrong data in query results, and GC deleting objects it did not mean to.
Fix
- Introduce
StorageNamespace — the address of the data, credentials excluded:
endpoint, bucket, root. Compare it structurally; if a compact key is
wanted, derive a collision-resistant digest, not a 64-bit DefaultHasher.
- Put the namespace into
CacheKey and into the keys of stat_cache,
KeyLocks, the prefetch seen set and InFlightTracker.
- Thread the namespace from
registry::OperatorKey to the point where
CacheLayer and PrefetchLayer are applied. StorageLayers::wrap_operator
currently does not know which operator it wraps, so its signature changes.
Alternatives rejected:
- a separate
CacheLayer per operator — does not fix it, StorageCache stays
shared;
- a separate foyer cache per bucket — splits the shared memory budget and
duplicates the disk cache.
Compatibility
Changing CacheKey changes the serialized key, i.e. foyer's on-disk format.
Existing disk-cache entries become misses. It warms up after the deploy; no
migration needed.
Definition of done
- a cache hit never crosses endpoint/bucket/root;
- the same relative path in two buckets returns the expected, different bytes and
different metadata;
- a delete or write in one bucket does not invalidate the other's stat entry;
- rotating credentials under an unchanged namespace keeps the cache shared —
otherwise the operator-leak fix loses its point for a credential-vending
catalog;
- the test proves an actual cache hit (via
CacheMetrics counters), not merely a
successful read from the backend;
- coverage: a component test for the path collision plus an S3 integration test
over two buckets.
Problem
CacheKeyholds nothing but the relativepath. The key carriesno storage identity, while the cache state behind it is shared:
StorageCache(foyer, the bytes) — one per process, shared by the Iceberg pathand the object-store path;
CacheLayer(stat cache,KeyLocks) — one perOperatorRegistry, cloned intoevery operator it builds;
PrefetchLayer:seen: DashSet<Arc<str>>andInFlightTracker.pending: DashMap<Arc<str>, _>— keyed by thebare path as well.
With two or more buckets (warehouse + catalog is the standard deployment) this
means:
tables/logs/metadata/v1.metadata.jsonin the second bucket can beserved the first bucket's bytes;
skips the read-ahead.
Risk: wrong data in query results, and GC deleting objects it did not mean to.
Fix
StorageNamespace— the address of the data, credentials excluded:endpoint,bucket,root. Compare it structurally; if a compact key iswanted, derive a collision-resistant digest, not a 64-bit
DefaultHasher.CacheKeyand into the keys ofstat_cache,KeyLocks, the prefetchseenset andInFlightTracker.registry::OperatorKeyto the point whereCacheLayerandPrefetchLayerare applied.StorageLayers::wrap_operatorcurrently does not know which operator it wraps, so its signature changes.
Alternatives rejected:
CacheLayerper operator — does not fix it,StorageCachestaysshared;
duplicates the disk cache.
Compatibility
Changing
CacheKeychanges the serialized key, i.e. foyer's on-disk format.Existing disk-cache entries become misses. It warms up after the deploy; no
migration needed.
Definition of done
different metadata;
otherwise the operator-leak fix loses its point for a credential-vending
catalog;
CacheMetricscounters), not merely asuccessful read from the backend;
over two buckets.