Skip to content

Repository files navigation

oold-llm-bench

Hypothesis-driven benchmark for schema-constrained LLM extraction.

It measures where structure has to be enforced for an LLM to produce data that validates. The candidates are the prompt, decode time, commit time, and nowhere at all. Arms range from free text with no schema to a fully grounded OO-LD schema, and one deterministic grader scores all of them.

Scope

Question Does schema enforcement change extraction quality, and where does it have to happen
Arms A0-prose, A0-json, A1 (schema in prompt), A2 (decode-time constraint), A3 (ontology grounding)
Metric Triple-level F1 over (entity, property, value), with class, unit and provenance reported separately
Corpora schema.org from real and synthetic documents; QUDT quantity kinds with closed unit enums

The benchmark depends on oold-python as one consumer among several, and must be able to run an arm that depends on nothing at all. That is why it does not live inside the library it evaluates.

Development

One command mirrors CI. Run it before every push:

make ci

It runs these in order, and each is also usable on its own:

Command Does
make check Lock file consistency, pre-commit on tracked and untracked files, ty, deptry
make test pytest with coverage
make docs-test Builds the documentation strictly, failing on any warning

Everything else, or make help for the full list:

Command Does
make install Creates the environment and installs the git hooks
make docs Serves the documentation locally
make build Builds a wheel into dist/
make check lints untracked files on purpose: pre-commit run -a skips them, which is how a new file passes locally and then fails in CI the moment it is committed.

Commits follow Conventional Commits; releases, the changelog and versioned documentation are automated on merge to main. Until a release App is configured, the release workflow skips itself instead of failing.

Updates from the template

This repository was generated from coregraft and records which version in .copier-answers.yml. A weekly workflow checks whether the template has moved on and, if so, replays its changes here. Your own edits are preserved: copier re-applies them on top of the new template version rather than over it, and anything it cannot merge arrives as <<<<<<< conflict markers for a human to decide.

By default the update arrives as an issue naming the new version and the command to run:

uvx copier update --skip-answered --trust --conflict inline

To get it as a pull request instead, this repository needs a GitHub App, because GitHub refuses to let the built-in token create or update anything under .github/workflows/, and template updates routinely do. One-time setup:

  1. Use or create a GitHub App with Contents: read and write, Pull requests: read and write and Workflows: read and write. The last one is the whole point; without it the push is rejected. Organisations often already have a release App, which may need the Workflows permission added and the updated permission accepted on each installation.
  2. Install the App on this repository.
  3. Provide RELEASE_APP_ID and RELEASE_APP_PRIVATE_KEY as repository secrets, or share the organisation secrets with this repository. Installing the App and sharing the secret are two separate settings.

The workflow detects the secret at runtime, so it opens a pull request once the App is available and falls back to an issue when it is not. Either way the merge is the same; only the automation differs.


Generated from coregraft.

About

Hypothesis-driven benchmark for schema-constrained LLM extraction: free-text, JSON Schema and OO-LD arms

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

Generated from OO-LD/coregraft