Skip to content

design: durable preflight results with manifest and SQLite evolution #2040

Description

@FabioLeitao

Context

The CPU preflight currently persists one row per measurement in TSV. This is sufficient for the current local workflow: it is append-friendly, inspectable with standard CLI tools, and supports the current --resume key.

As the HV/Vicuna evaluation pipeline grows across models, context tiers, hosts, repeated runs, and external harnesses, persistence and querying will become a design concern rather than an implementation detail.

Proposed design

Keep the current TSV format while the measurement schema is still stabilizing, and add a per-run JSON manifest for metadata such as:

  • wrapper/script version or git revision;
  • host and CPU details;
  • Ollama and harness versions;
  • complete command-line arguments and environment;
  • prompt/protocol version;
  • model inventory, capabilities, context limits, and selected variants;
  • run start/end status.

When the number of runs and comparison queries justifies it, make SQLite the canonical local store and retain TSV/CSV export for review and analysis.

The SQLite design should provide:

  • a uniqueness key for model, variant, context, threads, probe, repetition, and thinking level;
  • transactional/resumable writes;
  • multiple runs and hosts without overwriting history;
  • explicit status values for OK, timeout, unsupported, and errors;
  • normalized run/config metadata;
  • easy exports to TSV/CSV and downstream analysis.

Scope and non-goals

This is a design/follow-up issue, not a request to migrate the current wrapper immediately. The immediate priority is to stabilize the measurement schema and methodology. Any SQLite migration should preserve existing TSV results or provide an importer.

Suggested acceptance criteria

  • Decide whether TSV remains a supported export format.
  • Define the run/measurement schema and uniqueness semantics.
  • Define the JSON manifest fields and versioning.
  • Define SQLite tables, indexes, and migration/import strategy.
  • Ensure --resume behavior is deterministic across interrupted runs.
  • Document how preflight, coder-challenge, llmfit throughput, and llmfit quality results relate without collapsing unlike metrics into one score.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions