Skip to content

v0.3: model-group aggregation and compatibility hardening #5

Description

@AviBackToBlack

Goal

Build v0.3 around safe model-group semantics and the remaining compatibility gaps discovered during v0.1/v0.2 validation.

Proposed order

1. Multi-deployment / model-group aggregation

LiteLLM legitimately supports multiple deployments with the same user-facing model_name. v0.2 rejects duplicates, so the generator cannot currently represent a routed model group.

Research and implement a deterministic aggregation contract instead of selecting an arbitrary deployment.

Questions to settle before implementation:

  • how LiteLLM routes capability-constrained requests across deployments in a group
  • whether group capability values should be intersection/guarantee semantics, union/routable semantics, or field-specific
  • how null vs explicit false propagates across deployments
  • context/output token policy when deployments disagree
  • canonical/base-model identity when deployments route to different underlying models/providers
  • exact Codex template eligibility when only some deployments resolve to the same exact Codex model
  • reasoning-effort aggregation and explicit denial precedence
  • supported OpenAI parameter aggregation

Relevant upstream evidence:

Acceptance criteria:

  • duplicate model_name values no longer fail merely because the group has multiple deployments
  • aggregation is deterministic and order-independent
  • no capability is advertised more strongly than the selected aggregation contract permits
  • ambiguous/conflicting exact-template identity fails closed or falls back conservatively; never arbitrarily picks a donor
  • explain exposes group size, disagreement notes, and field provenance
  • tests cover same-model duplicate deployments, heterogeneous providers, conflicting true/false/null capabilities, reasoning levels, context windows, and order independence

2. Offline bundle version identity

Replace the current caller-trust boundary between --catalog-file, --codex-prompt-file, and --codex-schema-file with a verifiable bundle/manifest contract.

Acceptance criteria:

  • one explicit Codex version/ref identity for all offline resources
  • digest verification for bundled files
  • mixed-version resources fail closed
  • exact-only offline generation remains lightweight

3. Per-model overrides and globs

Extend the exact string allowlist without making selection ambiguous.

Acceptance criteria:

  • backwards-compatible exact strings
  • deterministic glob expansion
  • explicit per-model overrides with provenance
  • no silent inclusion of non-chat/non-responses models

4. Foreign web-search compatibility

Research before enabling. Do not simply map supports_web_search=true to Codex search support until the wire/tool semantics are proven compatible.

Current reason for caution:

Acceptance criteria:

  • identify exact Codex search-tool wire expectations for the matched Codex version
  • distinguish provider-native web search from OpenAI-compatible search parameter support
  • only enable when the LiteLLM route and Codex protocol are both evidenced compatible
  • otherwise retain conservative false/disabled behavior

Non-goals

  • no model-specific donor cloning for foreign models
  • no weakening of v0.2 schema-drift fail-closed behavior
  • no hidden inference from LiteLLM null
  • no real/private LiteLLM payloads or infrastructure data in fixtures

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions