Skip to content

CatalogProvider Limitations: Query Scoped Resolution #24

Description

@alexanderbianchi

Motivation

The current catalog integration combines long-lived catalog discovery, cached table providers, and fresh metadata loading during physical planning. These lifetimes do not align well with query-scoped authentication or consistent schema resolution.

@DerGut's apache/iceberg-rust#3000 addresses some of these issues in connecting the SessionCatalog work in iceberg-rust apache/iceberg-rust#2920. Discussions there are relevant for this issue.

This issue is intended to discuss the current limitations and general direction for the catalog integration. I’ll create an EPIC with separate follow-up issues for each component.

1. Catalog contents and providers are cached for the provider’s lifetime

The catalog provider eagerly stores every schema:

  pub struct IcebergCatalogProvider {
      schemas: HashMap<String, Arc<dyn SchemaProvider>>,
  }

Each schema provider eagerly stores every table provider:

  pub(crate) struct IcebergSchemaProvider {
      catalog: Arc<dyn Catalog>,
      namespace: NamespaceIdent,
      tables: Arc<DashMap<String, Arc<IcebergTableProvider>>>,
  }

Construction lists namespaces, then lists tables in each discovered namespace, effectively enumerating the whole catalog.

  client.list_namespaces(None).await?
   ...
  client.list_tables(&namespace).await?

The same Arc is returned to every query:

  Ok(self
      .tables
      .get(name)
      .map(|entry| entry.value().clone() as Arc<dyn TableProvider>))

Consequences:

  • tables created or dropped outside this provider are not reflected
  • startup requires catalog-wide listing permissions
  • initialization cost scales with the entire catalog rather than the query
  • providers cannot naturally be scoped to a tenant or authenticated request
  • querying one table requires initializing providers for unrelated tables.

2. Table resolution happens too late, task planning (might) happen too late

The table provider caches its schema during construction, but reloads the table during scan and insert planning. If the table changes between those steps, planning can combine the cached schema with newer metadata. This also repeats catalog requests for a table already loaded during resolution.

Task Planning happens during execution

File discovery and scan planning also happen during execution, we could use the information from the planned file tasks to feed more accurate statistics to the physical planner (in the case of file pruning, for example).

   let tasks: Vec<FileScanTask> =  table_scan.plan_files().await?.try_collect().await?;

3. Remote catalog operations do not fit the synchronous provider lifecycle

Remote Iceberg operations are asynchronous, while parts of DataFusion’s catalog interface are synchronous. The current integration handles discovery by eagerly loading and caching catalog contents during construction. For mutations, SchemaProvider::register_table() and deregister_table() instead block on remote create/drop operations. This occupies a blocking-pool thread while awaiting remote I/O and makes cancellation and runtime behavior harder to reason about

These are two consequences of the same mismatch: remote catalog access needs an explicit asynchronous boundary, with resolved providers serving subsequent planning lookups.

What a proper integration should provide

Asynchronous, query-scoped resolution

  • Load only referenced tables, resolving repeated references once.
  • Avoid requiring catalog-wide listing permissions to query a known table.
  • Reuse the catalog client across queries and resolved providers throughout each query.
  • Keep catalog resolution, scan planning, and execution as separate stages.

AsyncCatalogProvider helps, but context must be bound explicitly

DataFusion’s asynchronous catalog traits support resolving references into query-local providers before planning. However, resolve() receives only SessionConfig, and the lower-level schema/table lookups receive no session context. The integration therefore needs an explicit way to bind trusted request context before resolution.

Request context bound before catalog access

  • Let the embedding application supply trusted identity and credentials before resolution.
  • Preserve that context through catalog operations, including commit refresh and retries.
  • Isolate concurrent queries, define a stable session identity, and fail early when required context is missing.
  • Keep catalog credentials out of SQL-settable options and avoid automatically forwarding them to execution workers.

Fresh table resolution before planning

  • Load referenced tables through the schema provider for each query.
  • Build query-local table providers so schema resolution and scan planning use the same loaded metadata.
  • Let subsequent queries observe updates without rebuilding the shared catalog client.

Reusable providers with replaceable data sources

DataFusion’s DataSourceExec is a standard physical scan node that delegates reading, partitioning, statistics, and metrics to a DataSource implementation (DFD example). Keeping scan construction replaceable would let other usecases (like dataFusion-distributed) reuse the upstream catalog, schema, and table providers while supplying its own data source; the execution details belong in a separate issue.

Proposed architecture

Embedding application
  └── authenticates request and supplies trusted context
          |
          v
Query-bound catalog/schema resolution
  ├── uses the shared catalog client with that context
  ├── loads referenced tables for this query
  └── creates query-local table providers
          |
          v
Iceberg TableProvider
  ├── schema(): exposes the loaded table's schema
  └── scan():
        ├── selects snapshot
        ├── converts predicates
        ├── invokes Iceberg's scan planner
        └── passes planned FileScanTasks to the factory
          |
          v
Configurable data-source factory
  └── organizes planned tasks for local or distributed execution
          |
          v
DataSourceExec
  ├── local IcebergDataSource
  └── downstream/distributed IcebergDataSource
          |
          v
Execution
  └── uses Iceberg's Arrow reader to read assigned tasks

Discussion

  1. Should referenced tables be resolved asynchronously per query, before logical planning?
  2. How should trusted request context be bound before the first catalog call?
  3. Should scan construction be replaceable so downstream engines can distribute those tasks?
  4. Should Iceberg file scan tasks be planned during physical planning, before execution?
  5. How should remote create/drop operations work without blocking synchronous catalog methods?

The most relevant / important issues here have to do with the literal catalog operations + authentication. Some of the questions raised fall into a broader category of structuring the planning flow.

I can go more into detail about how I see the implementation in a subsequent issue, or edit and append here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions