Skip to content

[Query] Split the foyer cache: isolate Iceberg metadata from Parquet data files #170

Description

@s-prosvirnin

Problem

We build a single foyer HybridCache<CacheKey, CacheValue> in IoHandle::from_config and share that one instance between two consumers:

  • the Iceberg FileIO, via IceGateStorageFactory — metadata.json, manifest lists, manifests (Avro), and Parquet data files;
  • the WAL object store, via create_object_store.

The cache key is the object path only. There is no per-object-class partitioning, no separate budgets, and no priorities — the weighter is the entry's byte size and everything competes in one LRU under one byte budget.

Consequence: a large table scan evicts Iceberg metadata. Every subsequent query then re-reads metadata.json and the manifests from S3 before it can even start planning. The prefetch layer makes this worse — it sits outside the cache layer and streams Parquet column chunks into the same LRU.

max_write_cache_size does not help. It only applies to the write path in CacheWriter; the read path merges and inserts any fetched range unconditionally.

Note what is already isolated and out of scope: the stat/HEAD cache is a separate DashMap with a TTL, the query engine's catalog provider has its own watch channel, and ParquetQueueReader has its own metadata-entries cache. Only the cached bytes share a budget.

Goal

Metadata reads must not be evicted by data-file reads. Two foyer instances with independent memory and disk budgets:

  • metadata_cache — .metadata.json, manifest lists, manifest files (.avro), catalog root.json;
  • data_cache — .parquet and WAL segments.

CacheLayer routes by path on read, write, and delete. Classification lives in one function so the rule is testable and stated once.

Scope

  • CacheConfig gains separate metadata/data sizing. Keep the existing fields working: if only the flat sizes are given, split them by a documented default ratio rather than failing. Validate the new fields alongside the existing ones.
  • IoHandle holds and closes both caches; close() must drain both.
  • CacheLayer takes both caches and picks one per key. Foyer metrics registration reports each tier separately (add a tier attribute or distinct instrument names — decide and document which).
  • IceGateStorageFactory and create_s3_store thread both caches through.

Rejected alternative

A read-path size cap symmetric to max_write_cache_size is cheaper but does not fix this: prefetch produces many small column chunks that individually pass any threshold and collectively evict metadata anyway. It provides no isolation guarantee.

Acceptance

  • Filling data_cache past its capacity leaves metadata entries resident — cover this with a test that asserts metadata is still served from cache after the data tier is thrashed.
  • Existing configs that set only memory_size_mb / disk_size_mb still start, with the resulting split logged at startup.
  • Metrics distinguish the two tiers.
  • Update the cache section of the affected crate docs — this changes the documented contract of max_write_cache_size (it stops being the metadata-protection mechanism).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions