Skip to content

Isolate the storage cache by storage namespace #173

Description

@s-prosvirnin

Problem

CacheKey holds nothing but the relative path. The key carries
no storage identity, while the cache state behind it is shared:

  • StorageCache (foyer, the bytes) — one per process, shared by the Iceberg path
    and the object-store path;
  • CacheLayer (stat cache, KeyLocks) — one per OperatorRegistry, cloned into
    every operator it builds;
  • PrefetchLayer: seen: DashSet<Arc<str>> and
    InFlightTracker.pending: DashMap<Arc<str>, _>— keyed by the
    bare path as well.

With two or more buckets (warehouse + catalog is the standard deployment) this
means:

  • a read of tables/logs/metadata/v1.metadata.json in the second bucket can be
    served the first bucket's bytes;
  • a stat (HEAD) against the second bucket returns the first bucket's metadata;
  • a delete or write in one bucket invalidates the other bucket's stat entry;
  • prefetch treats an object as already seen because of a different bucket and
    skips the read-ahead.

Risk: wrong data in query results, and GC deleting objects it did not mean to.

Fix

  1. Introduce StorageNamespace — the address of the data, credentials excluded:
    endpoint, bucket, root. Compare it structurally; if a compact key is
    wanted, derive a collision-resistant digest, not a 64-bit DefaultHasher.
  2. Put the namespace into CacheKey and into the keys of stat_cache,
    KeyLocks, the prefetch seen set and InFlightTracker.
  3. Thread the namespace from registry::OperatorKey to the point where
    CacheLayer and PrefetchLayer are applied. StorageLayers::wrap_operator
    currently does not know which operator it wraps, so its signature changes.

Alternatives rejected:

  • a separate CacheLayer per operator — does not fix it, StorageCache stays
    shared;
  • a separate foyer cache per bucket — splits the shared memory budget and
    duplicates the disk cache.

Compatibility

Changing CacheKey changes the serialized key, i.e. foyer's on-disk format.
Existing disk-cache entries become misses. It warms up after the deploy; no
migration needed.

Definition of done

  • a cache hit never crosses endpoint/bucket/root;
  • the same relative path in two buckets returns the expected, different bytes and
    different metadata;
  • a delete or write in one bucket does not invalidate the other's stat entry;
  • rotating credentials under an unchanged namespace keeps the cache shared —
    otherwise the operator-leak fix loses its point for a credential-vending
    catalog;
  • the test proves an actual cache hit (via CacheMetrics counters), not merely a
    successful read from the backend;
  • coverage: a component test for the path collision plus an S3 integration test
    over two buckets.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions