Skip to content

docs(validation): give the context walker a true justification - #150

Merged
simontaurus merged 3 commits into
mainfrom
docs/context-walker-justification
Sep 11, 2026
Merged

simontaurus merged 3 commits into
mainfrom
docs/context-walker-justification

Conversation

@LukasGold

Copy link
Copy Markdown
Contributor

Closes #118 by taking its second branch: "if it does not hold, record which specific capability is missing so the module has a justification that survives review."

The walker stays. context_resolution.py is unchanged apart from its docstring, which was arguing from a premise that had already been disproven.

What was wrong

Two claims:

  • "What a processor does not expose is the flattened active context itself as a value a caller can inspect." pyld does expose it, through the public JsonLdProcessor.process_context.
  • "this module exists because none does today, not because a processor could not in principle resolve the chain itself." The first half was false, so the sentence justified the module on the one ground that did not hold.

Diffed across all 13 top-level fixtures, pyld reproduces the walker's top-level flattening faithfully and in places better: it reports pre-expanded IRIs where the walker reports compact CURIEs, and on Organization.schema.json's employees, a @reverse term, the walker's .get("@id") yields None where pyld surfaces the reverse target.

What is actually missing

process_context never resolves a term's scoped @context. It leaves the value exactly as authored and defers it to expansion, so {"@id": ..., "@context": "Pet.schema.json"} comes back as an unresolved string. Two committed fixtures use that form: Organization.schema.json (address) and PersonWithPet.schema.json (pets).

That would be harmless if expansion could resolve it later, but the callers here expand without a document loader: pipeline.py:487 calls check_predicates with no options, and check_predicates defaults options to None. Demonstrated on Organization.schema.json's real context:

raw @context, no document loader     -> ValueError: Found invalid relative IRI 'Thing.schema.json'
walker's resolved context, no loader -> expand OK, identical predicates
raw @context + a document loader     -> expand OK, identical predicates

So something has to embed the scoped content eagerly, which is most of what _resolve_inline does today.

Successive calls share no memory of prior hops, so nothing stops a caller re-entering the processor around a reference cycle. The stack in _walk_reference and max_scoped_depth in _resolve_inline are that bound, and delegating the per-hop merge to pyld does not eliminate them.

Verification

Docstring only, no code path touched.

  • full suite: 585 passed, 9 skipped, unchanged from main
  • make check: exit 0
  • parity: 6 passed, 588 deselected, so it ran rather than skipping silently
  • CRLF preserved, no mixed line endings

- pyld does expose the flattened context via public process_context
- what it does not do is resolve a term's scoped @context
- callers expand with no document loader, so it cannot resolve later
- successive calls share no memory, so the cycle bound stays here
@github-actions

github-actions Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Release preview

No version bump from the current commits (stays at v0.19.0). Use conventional commit types (feat, fix, ...) to trigger a release.

Changelog preview (truncated)

Preview via python-semantic-release and conventional commits.

@github-actions

github-actions Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

📊 Benchmark Results

Click to see benchmark comparison
📊 Benchmark Comparison (threshold: 1.3x)
============================================================

➖ Unchanged (within threshold):
  ➖ test_simple_dict_document_store: 0.0015s → 0.0015s (+1.0%)
  ➖ test_sqlite_document_store: 0.0016s → 0.0016s (-1.9%)
  ➖ test_local_sparql_store: 0.0324s → 0.0323s (-0.2%)
  ➖ test_oneof_subschema: 0.0523s → 0.0538s (+2.8%)
  ➖ test_enum_docstrings: 0.0445s → 0.0442s (-0.8%)
  ➖ test_subclass_inheritance: 0.0474s → 0.0485s (+2.4%)
  ➖ test_class_hierarchy: 0.0441s → 0.0450s (+2.2%)
  ➖ test_core[v1]: 0.0338s → 0.0330s (-2.4%)
  ➖ test_core[v2]: 0.0385s → 0.0387s (+0.5%)
  ➖ test_schema_generation[v1]: 0.0016s → 0.0015s (-0.3%)
  ➖ test_schema_generation[v2]: 0.0025s → 0.0025s (-0.1%)
  ➖ test_simple_json: 0.0006s → 0.0006s (-0.3%)
  ➖ test_complex_graph: 0.0013s → 0.0013s (+0.1%)

============================================================
Summary: 0 regressions, 0 improvements, 13 unchanged
============================================================

✅ No significant performance regressions

Threshold: 1.3x (30% slower triggers a regression warning)

Note: Benchmarks are informational only and won't fail the build.

💡 Tip: Download the benchmark-results artifact for detailed JSON data

@codecov

codecov Bot commented Sep 2, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@simontaurus

Copy link
Copy Markdown
Contributor

Conclusion holds, both mechanisms are stated wrong.

Scoped context. process_context does not "leave the value exactly as authored and defer it to expansion". It resolves and validates the scoped context eagerly, and without a loader raises before returning:

JsonLdError: Invalid JSON-LD syntax; invalid scoped context.
Details: {'context': 'Pet.schema.json', 'term': 'pets'}

So "by the time that reference is reached there is nothing left to resolve it against" is not the failure. What is true is that pyld does not keep the result: with a loader supplied, mappings["pets"]["@context"] is still the string "Pet.schema.json". That is the justification for _resolve_inline - the value is resolved and then discarded, not deferred.

Cycles. pyld detects cyclic @context URLs within a single call:

JsonLdError: Cyclical @context URLs detected.  (code: context overflow)

What has no memory is the hop-by-hop driving that #118 needs for provenance. Worth saying that, since as written it reads as "the processor is unguarded", and the real point is that recovering provenance is what costs you its guard.

Both conclusions held, both mechanisms were wrong. process_context does
not defer a scoped @context to expansion; it resolves eagerly and keeps
nothing, raising invalid scoped context when no loader is available. And
pyld does detect cyclic @context URLs within one call - what defeats the
guard is driving the chain a hop at a time, which is what provenance
requires.
@github-actions

Copy link
Copy Markdown
Contributor

📊 Benchmark Results

Click to see benchmark comparison
📊 Benchmark Comparison (threshold: 1.3x)
============================================================

✅ Performance Improvements:
  ✅ test_enum_docstrings: 0.0554s → 0.0405s (-26.8%, ratio: 0.73x)
  ✅ test_subclass_inheritance: 0.0744s → 0.0430s (-42.2%, ratio: 0.58x)
  ✅ test_core[v2]: 0.0495s → 0.0381s (-23.1%, ratio: 0.77x)

➖ Unchanged (within threshold):
  ➖ test_simple_dict_document_store: 0.0012s → 0.0012s (-2.0%)
  ➖ test_sqlite_document_store: 0.0013s → 0.0013s (-1.4%)
  ➖ test_local_sparql_store: 0.0290s → 0.0290s (-0.2%)
  ➖ test_oneof_subschema: 0.0517s → 0.0489s (-5.3%)
  ➖ test_class_hierarchy: 0.0535s → 0.0470s (-12.2%)
  ➖ test_core[v1]: 0.0396s → 0.0324s (-18.2%)
  ➖ test_schema_generation[v1]: 0.0012s → 0.0012s (-0.8%)
  ➖ test_schema_generation[v2]: 0.0021s → 0.0020s (-1.2%)
  ➖ test_simple_json: 0.0004s → 0.0004s (+0.5%)
  ➖ test_complex_graph: 0.0011s → 0.0012s (+4.2%)

============================================================
Summary: 0 regressions, 3 improvements, 10 unchanged
============================================================

✅ No significant performance regressions

Threshold: 1.3x (30% slower triggers a regression warning)

Note: Benchmarks are informational only and won't fail the build.

💡 Tip: Download the benchmark-results artifact for detailed JSON data

The justification names what pyld does not return; that only holds while
the two agree on everything else, and nothing checked it. Both are facts
now: the effective term set matches across all 13 committed schemas, and
a scoped @context comes back from process_context exactly as authored.

Where a caller needs only the effective mapping, process_context is the
better source - it applies JSON-LD's override semantics rather than
terms()'s later-key-wins flatten.
@github-actions

Copy link
Copy Markdown
Contributor

📊 Benchmark Results

Click to see benchmark comparison
📊 Benchmark Comparison (threshold: 1.3x)
============================================================

➖ Unchanged (within threshold):
  ➖ test_simple_dict_document_store: 0.0017s → 0.0017s (+0.2%)
  ➖ test_sqlite_document_store: 0.0019s → 0.0019s (+0.1%)
  ➖ test_local_sparql_store: 0.0405s → 0.0397s (-2.0%)
  ➖ test_oneof_subschema: 0.0644s → 0.0652s (+1.3%)
  ➖ test_enum_docstrings: 0.0528s → 0.0535s (+1.3%)
  ➖ test_subclass_inheritance: 0.0565s → 0.0574s (+1.6%)
  ➖ test_class_hierarchy: 0.0546s → 0.0651s (+19.1%)
  ➖ test_core[v1]: 0.0402s → 0.0406s (+1.1%)
  ➖ test_core[v2]: 0.0461s → 0.0465s (+0.9%)
  ➖ test_schema_generation[v1]: 0.0017s → 0.0017s (-2.3%)
  ➖ test_schema_generation[v2]: 0.0028s → 0.0028s (+1.3%)
  ➖ test_simple_json: 0.0007s → 0.0007s (+1.2%)
  ➖ test_complex_graph: 0.0016s → 0.0017s (+1.3%)

============================================================
Summary: 0 regressions, 0 improvements, 13 unchanged
============================================================

✅ No significant performance regressions

Threshold: 1.3x (30% slower triggers a regression warning)

Note: Benchmarks are informational only and won't fail the build.

💡 Tip: Download the benchmark-results artifact for detailed JSON data

@simontaurus
simontaurus merged commit 4556612 into main Sep 11, 2026
21 checks passed
@simontaurus
simontaurus deleted the docs/context-walker-justification branch September 11, 2026 10:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Evaluate replacing the hand-written context walker with pyld

3 participants