diff --git a/README.md b/README.md index 3a068d5..6fe80a7 100644 --- a/README.md +++ b/README.md @@ -102,7 +102,7 @@ Existing `models = ["..."]` entries are always exact strings; wildcard character `model_overrides` is keyed by an exact user-facing `model_name`; override keys are never globs. The target must be covered by either `models` or `model_globs`. Overrides operate on aggregated compatibility evidence, not on individual deployments and not on arbitrary Codex catalog fields. Supported fields are the six capability booleans (`supports_vision`, `supports_audio_input`, `supports_function_calling`, `supports_parallel_function_calling`, `supports_web_search`, `supports_reasoning`), `max_input_tokens`, `max_output_tokens`, `supported_openai_params`, and `reasoning_effort_levels`. `max_output_tokens` is currently evidence/audit-only because the generated Codex model schema has no independent output-token-limit field that consumes it; it is retained for explainability and future fields that may depend on that evidence. -An override is an explicit trusted assertion and can replace aggregate `guaranteed`, `denied`, or `unknown` evidence. `explain` keeps the original aggregate state/value beside the configured value and effective state/value. Dependency closure still applies: disabling reasoning disables reasoning transport/efforts, and disabling function calling prevents parallel tool calls. Overrides cannot force exact-template identity or enable foreign-model Codex web search. +An override is an explicit trusted assertion and can replace aggregate `guaranteed`, `denied`, or `unknown` evidence. `explain` keeps the original aggregate state/value beside the configured value and effective state/value. Dependency closure still applies: disabling reasoning disables reasoning transport/efforts, and disabling function calling prevents parallel tool calls. `supports_web_search` overrides remain evidence-only for hosted-search compatibility: they do not control Codex `supports_search_tool`, and the generated model catalog alone does not guarantee hosted web-search enablement or suppression. See [docs/foreign-web-search.md](docs/foreign-web-search.md). `version = "auto"` runs `codex --version` and fetches the catalog from the corresponding `rust-v` tag in `openai/codex`. This avoids using a `main` catalog whose schema may not match the installed Codex binary. @@ -175,11 +175,11 @@ Rows with the same exact raw `model_name` are treated as one LiteLLM routing gro For an **exact Codex template group**, every deployment must independently resolve unambiguously to the same version-matched Codex template. The Codex entry remains authoritative, while explicit deployment denials can conservatively downgrade capabilities. Unknown evidence alone does not narrow exact-template behavior. A configured override may replace the aggregate compatibility evidence used for those downgrade decisions, but cannot change the proven exact identity or invent Codex-owned fields. -For a **foreign group**, the generated catalog entry describes what is safe for an arbitrary routed request: boolean capabilities require a group guarantee, supported parameter and reasoning sets are intersected, and context/output limits use the safe minimum only when every deployment provides a valid value. Explicit overrides may replace those aggregate evidence values. Foreign web search remains disabled even when the effective web-search evidence is true. +For a **foreign group**, the generated catalog entry describes what is safe for an arbitrary routed request: boolean capabilities require a group guarantee, supported parameter and reasoning sets are intersected, and context/output limits use the safe minimum only when every deployment provides a valid value. Explicit overrides may replace those aggregate evidence values. Web-search evidence is retained for audit/explain but is not mapped to Codex `supports_search_tool` or treated as sufficient proof for hosted-search compatibility; hosted web search also depends on Codex provider/runtime behavior. ## Context-window policy -For an **exact Codex template match**, `context_window` and `max_context_window` remain the Codex values. LiteLLM `max_input_tokens` is treated as validation evidence because the two fields do not have identical semantics. A `max_input_tokens` override likewise changes validation evidence only; it does not replace exact-template Codex context fields. +For an **exact Codex template match**, `context_window` and `max_context_window` remain the Codex values. LiteLLM `max_input_tokens` is treated as validation evidence because the two fields do not have identical semantics. A `max_input_tokens` override likewise changes the LiteLLM validation evidence only; it does not replace the exact template's Codex context fields. For a **foreign model group**, the generator uses the minimum known LiteLLM `max_input_tokens` across every deployment as the best safe approximation for both context fields. A multi-deployment foreign group with missing or invalid context evidence fails closed rather than advertising a guessed window. An explicit `max_input_tokens` override can supply the trusted effective evidence needed to synthesize that group. A single foreign deployment preserves the v0.2 fallback behavior. @@ -196,7 +196,7 @@ For a **foreign model group**, the generator uses the minimum known LiteLLM `max ## Current limitations -- Foreign-model web search remains disabled even when LiteLLM or a configured override advertises web search; Codex search-tool wire semantics need an explicit compatibility rule. +- Foreign hosted web-search compatibility is not inferred from LiteLLM `supports_web_search`. Current Codex hosted search is provider/runtime controlled, while `supports_search_tool` has separate tool-discovery semantics; see [docs/foreign-web-search.md](docs/foreign-web-search.md). The generated model catalog alone does not guarantee hosted-search enablement or suppression. - Foreign-model context-window mapping is an approximation, as described above. - `model_overrides` are trusted user assertions; they can deliberately replace conservative aggregate evidence, so incorrect overrides can over-advertise gateway compatibility even though Codex-owned template identity/fields remain protected. - Hand-authored offline bundle manifests provide digest integrity and one declared Codex identity, but are not signed provenance attestations that the files came from the named Git ref. diff --git a/docs/foreign-web-search-sources.md b/docs/foreign-web-search-sources.md new file mode 100644 index 0000000..badf68a --- /dev/null +++ b/docs/foreign-web-search-sources.md @@ -0,0 +1,60 @@ +# Foreign web-search research source index + +This companion page records the upstream evidence used by `foreign-web-search.md` so the contract can be revalidated against future Codex/LiteLLM versions without redoing discovery from scratch. + +## Evidence scope + +The Codex behavior used for the v0.3 decision was verified against **`rust-v0.154.0`**, the latest stable Codex release on 2026-09-14, and then cross-checked against upstream `main` for drift. The contract does not assume that every historical or future Codex ref has identical semantics. + +For any generated catalog, the version-matched Codex ref remains authoritative. If a future matched ref changes these semantics, the generator must re-evaluate the mapping rather than inherit this snapshot blindly. + +## Codex — version-matched stable evidence (`rust-v0.154.0`) + +- Hosted web-search construction and `WebSearchToolType` consumption: + - https://github.com/openai/codex/blob/rust-v0.154.0/codex-rs/core/src/tools/hosted_spec.rs +- Hosted web-search provider/config gating and separate `search_tool_enabled()` path: + - https://github.com/openai/codex/blob/rust-v0.154.0/codex-rs/core/src/tools/spec_plan.rs +- `WebSearchToolType` schema (`text`, `text_and_image`; no disabled variant): + - https://github.com/openai/codex/blob/rust-v0.154.0/codex-rs/protocol/src/openai_models.rs +- Configured-provider capability defaults (`web_search=true` via `ProviderCapabilities::default()`): + - https://github.com/openai/codex/blob/rust-v0.154.0/codex-rs/model-provider/src/provider.rs + +The same semantics were still present on upstream `main` when researched on 2026-09-14: + +- https://github.com/openai/codex/blob/main/codex-rs/core/src/tools/hosted_spec.rs +- https://github.com/openai/codex/blob/main/codex-rs/core/src/tools/spec_plan.rs +- https://github.com/openai/codex/blob/main/codex-rs/protocol/src/openai_models.rs +- https://github.com/openai/codex/blob/main/codex-rs/model-provider/src/provider.rs +- https://github.com/openai/codex/blob/main/codex-rs/core/tests/suite/web_search.rs + +Additional behavioral evidence: + +- Custom catalog/tool-search failure showing `supports_search_tool` is about deferred tool discovery, while hosted web search is independent: + - https://github.com/openai/codex/issues/36382 +- Provider capability/config limitations relevant to suppression/opt-in: + - https://github.com/openai/codex/issues/21952 + - https://github.com/openai/codex/issues/24465 +- LiteLLM/custom-provider model discovery context: + - https://github.com/openai/codex/issues/30760 + - https://github.com/openai/codex/issues/37122 + +## LiteLLM + +- Mixed-provider routing bug demonstrating that `supports_web_search` metadata does not by itself prove deterministic compatibility for every native search request shape: + - https://github.com/BerriAI/litellm/issues/38982 +- Responses/web-search integration gaps and response-shape translation context: + - https://github.com/BerriAI/litellm/issues/26073 +- Provider-specific Responses bridge interaction for search: + - https://github.com/BerriAI/litellm/issues/37127 + +## Revalidation rule + +Before changing or applying the contract to a new Codex version, re-check the **version-matched** source used by the generator, not only `main`. In particular verify: + +1. whether `WebSearchToolType` gained a disabled/off variant; +2. whether provider capabilities became user-configurable; +3. whether `supports_search_tool` changed meaning; +4. whether hosted search is still gated by provider capability + runtime mode; +5. whether LiteLLM routing guarantees the exact OpenAI Responses hosted-search wire shape across every deployment in the selected model group. + +The v0.3 rule is intentionally conservative: without a version-matched proof of semantic equivalence, LiteLLM `supports_web_search` must not be mapped to Codex `supports_search_tool` or treated as sufficient proof for hosted-search enablement. diff --git a/docs/foreign-web-search.md b/docs/foreign-web-search.md new file mode 100644 index 0000000..47858a2 --- /dev/null +++ b/docs/foreign-web-search.md @@ -0,0 +1,172 @@ +# Foreign web-search compatibility contract + +Status: **research complete; hosted-search enablement is blocked/deferred**. + +This document closes the research/design gate for v0.3 foreign-model web search. It deliberately does **not** enable web search for foreign models. + +## Evidence scope + +The concrete Codex behavior below was verified against **`rust-v0.154.0`**, the latest stable Codex release on 2026-09-14, and cross-checked against upstream `main` for drift. It is not a claim that every historical or future Codex version has identical semantics. + +The generator remains version-aware: for a particular generated catalog, the matching `rust-v` source is authoritative. The safe repository-wide rule is therefore narrower: **never map similarly named LiteLLM and Codex search fields unless semantic equivalence is proven for the version-matched Codex ref and the routed LiteLLM wire path.** + +## Why this needs a separate contract + +LiteLLM and Codex expose similarly named search capability fields, but they are not the same contract. + +The generator must not infer Codex hosted web-search compatibility from LiteLLM `supports_web_search` alone. A generated catalog entry describes model metadata, while Codex decides whether to emit a hosted `web_search` tool from a combination of model metadata, provider capabilities, and runtime configuration. + +## Codex findings + +The current stable Codex ref (`rust-v0.154.0`) has two independent search concepts. + +### `supports_search_tool` is tool discovery, not hosted web search + +Codex `search_tool_enabled()` gates deferred/namespace tool discovery with: + +- `model_info.supports_search_tool`; and +- provider `namespace_tools` capability. + +It is used for `tool_search` and deferred tool exposure. It is **not** the hosted web-search enable bit. + +Version-matched source: + +- `codex-rs/core/src/tools/spec_plan.rs` +- https://github.com/openai/codex/blob/rust-v0.154.0/codex-rs/core/src/tools/spec_plan.rs + +This distinction is externally observable. OpenAI Codex issue #36382 documents custom models where `supports_search_tool=true` hides ordinary MCP tools behind deferred discovery while hosted web search remains independent: + +- https://github.com/openai/codex/issues/36382 + +Therefore v0.3 must not map LiteLLM `supports_web_search` to Codex `supports_search_tool`. A future Codex ref may be re-evaluated only from its version-matched semantics. + +### Hosted web search is provider/config controlled + +At `rust-v0.154.0`, Codex builds a hosted `web_search` tool when all relevant runtime conditions allow it: + +1. the active provider advertises `ProviderCapabilities.web_search=true`; +2. runtime `web_search_mode` is not disabled; and +3. the model's `web_search_tool_type` controls the **shape** of the emitted hosted tool. + +`WebSearchToolType` at that ref has only: + +- `text`; and +- `text_and_image`. + +There is no model-catalog `disabled` variant. + +Version-matched source: + +- https://github.com/openai/codex/blob/rust-v0.154.0/codex-rs/core/src/tools/hosted_spec.rs +- https://github.com/openai/codex/blob/rust-v0.154.0/codex-rs/core/src/tools/spec_plan.rs +- https://github.com/openai/codex/blob/rust-v0.154.0/codex-rs/protocol/src/openai_models.rs + +Configured/custom providers at that ref inherit `ProviderCapabilities::default()`, where `web_search=true`. The model catalog itself does not provide a per-model hosted-search off switch. + +Version-matched source: + +- https://github.com/openai/codex/blob/rust-v0.154.0/codex-rs/model-provider/src/provider.rs + +The same architecture was still present on upstream `main` on 2026-09-14. The companion source index records both the stable-ref evidence and the `main` drift cross-check. + +## LiteLLM findings + +LiteLLM `supports_web_search` is useful capability evidence, but it does not by itself prove that an arbitrary LiteLLM deployment accepts the exact OpenAI Responses `tools: [{"type":"web_search", ...}]` wire shape emitted by Codex. + +Providers can expose native search through different request shapes and transformations. Recent upstream evidence includes: + +- Anthropic-native `web_search_*` tool types versus OpenAI `web_search` / `web_search_preview` routing detection; +- provider-specific completion-to-Responses bridges for search; +- model/provider metadata gaps where `supports_web_search` is missing or wrong. + +In particular, LiteLLM issue #38982 documents a mixed-provider group where explicit per-deployment `supports_web_search=false` is not sufficient for Anthropic-native search routing because the router's request detector recognizes only OpenAI-style tool types in that path: + +- https://github.com/BerriAI/litellm/issues/38982 + +This reinforces the generator's existing group principle: metadata must describe what is safe for an arbitrary routed request, and similarly named provider-native capabilities are not interchangeable with a proven Codex/OpenAI wire contract. + +## Contract + +### 1. Never map LiteLLM web-search evidence to Codex `supports_search_tool` without version-matched proof + +For the researched stable ref, `supports_search_tool` belongs to Codex tool discovery/deferred exposure semantics while LiteLLM `supports_web_search` belongs to search capability evidence. They are independent. + +For exact Codex templates under this contract: + +- preserve the version-matched Codex template's `supports_search_tool` value; +- do not downgrade it because LiteLLM says `supports_web_search=false`; +- do not upgrade it because LiteLLM says `supports_web_search=true` or because a configured override says so. + +For foreign models: + +- keep the existing conservative `supports_search_tool=false` choice as a **tool-discovery** decision only; +- never describe that field as disabling hosted web search. + +If a future matched Codex ref changes `supports_search_tool` semantics, that ref must be researched before changing this rule. + +### 2. `web_search_tool_type` is not an enable bit in the researched stable ref + +At `rust-v0.154.0`, `text` versus `text_and_image` only chooses the hosted tool payload shape. A foreign entry may need a schema-valid `web_search_tool_type`, but the generator must not claim that `web_search_tool_type="text"` disables search. + +A future matched ref must be re-checked for new enum values or changed semantics. + +### 3. `supports_web_search` remains evidence-only for foreign hosted search + +Until a stronger compatibility proof exists, effective foreign `supports_web_search=true` may be recorded in aggregation/override evidence and `explain`, but it must not cause the generator to mutate Codex `supports_search_tool` or claim hosted-search compatibility. + +Likewise, `supports_web_search=false` is evidence about the LiteLLM route, but at the researched Codex ref the model catalog alone cannot enforce provider-level hosted-search suppression. + +### 4. Model catalog alone cannot guarantee hosted-search disablement for the researched stable ref + +At `rust-v0.154.0`: + +- `WebSearchToolType` has no disabled variant; +- configured providers default to `ProviderCapabilities.web_search=true`; +- hosted-search emission also depends on runtime `web_search_mode`. + +Therefore this project must not promise that foreign hosted web search is disabled solely because the generated model has `supports_search_tool=false`. + +Users who need guaranteed suppression must currently disable web search in the applicable Codex runtime/provider configuration. A future Codex provider capability override or model-level hosted-search disable field could provide a stronger enforceable boundary and should be evaluated from the matched ref. + +Related upstream requests/bugs: + +- https://github.com/openai/codex/issues/21952 +- https://github.com/openai/codex/issues/24465 +- https://github.com/openai/codex/issues/30760 +- https://github.com/openai/codex/issues/37122 + +### 5. Future enablement requires an explicit wire proof + +Foreign hosted web search may be enabled by this generator only when all of the following are proven for the **matched Codex version and selected LiteLLM route**: + +1. **Codex wire shape** — exact request tool type and fields emitted by the matched Codex version are known. +2. **Provider boundary** — the active Codex provider is allowed to emit hosted web search. +3. **LiteLLM route guarantee** — every deployment that can receive the public `model_name` accepts that exact wire shape, or LiteLLM deterministically filters to compatible deployments for that request shape. +4. **Response compatibility** — LiteLLM/provider responses preserve the response item/event shapes Codex expects, including `web_search_call` history when applicable. +5. **No semantic collision** — enabling hosted search does not repurpose `supports_search_tool`, tool-search discovery, MCP exposure, or another unrelated Codex capability. + +A plain LiteLLM `supports_web_search=true` is insufficient because it does not prove the full five-part contract. + +## Immediate implementation follow-up + +The next production PR should be intentionally small: + +1. remove the existing exact-template mapping from LiteLLM `supports_web_search=false` to Codex `supports_search_tool=false`; +2. remove override provenance that treats `supports_web_search` as ownership/validation of `supports_search_tool`; +3. keep foreign `supports_search_tool=false`, but describe it explicitly as conservative tool-search/deferred-discovery behavior; +4. keep README/current-limitations wording aligned with this contract so the project does not claim that model catalog metadata guarantees hosted web-search disablement; +5. add regressions proving exact `supports_search_tool` is Codex-owned regardless of LiteLLM web-search evidence. + +No v0.3 implementation should set foreign hosted-search behavior from `supports_web_search` until the five-part proof above is available for the matched Codex/LiteLLM path. + +## v0.3 disposition + +The research acceptance criteria are satisfied for the current v0.3 decision: + +- the current stable Codex hosted-search wire construction and separate `supports_search_tool` meaning are identified and source-backed at `rust-v0.154.0`; +- provider-native/LiteLLM search evidence is distinguished from Codex/OpenAI hosted-search wire compatibility; +- the current model-catalog boundary is documented as insufficient for safe foreign enablement or guaranteed provider-level disablement; +- future mappings require version-matched revalidation rather than extrapolation from `main`; +- the implementation remains fail-closed with respect to **automatic capability mapping**: no foreign hosted-search behavior is enabled from LiteLLM evidence. + +Actual hosted-search enablement is **deferred** pending a stronger Codex provider/model capability boundary plus route-level wire compatibility evidence. diff --git a/docs/model-selection-overrides.md b/docs/model-selection-overrides.md index 960c335..26393fc 100644 --- a/docs/model-selection-overrides.md +++ b/docs/model-selection-overrides.md @@ -14,7 +14,7 @@ This work does not: - allow arbitrary Codex catalog fields to be overwritten; - allow overrides to force exact-template identity; - allow `mode`, provider, deployment model, or base-model identity to be overridden; -- enable foreign-model web search; +- map LiteLLM `supports_web_search` to Codex hosted-search behavior or `supports_search_tool`; - infer capabilities from model names or provider names; - add exclusion patterns or regular-expression syntax in v0.3. @@ -129,7 +129,7 @@ In particular: - effective `supports_reasoning = false` disables reasoning efforts and reasoning-summary transport on generated output; - effective `supports_function_calling = false` prevents parallel tool calls; - parallel tool calls for foreign models still require all existing parameter/function/parallel guarantees; -- an override cannot turn on foreign Codex web search. `supports_web_search = true` may alter the recorded LiteLLM evidence or prevent an exact-template downgrade, but foreign `supports_search_tool` remains disabled until the separate web-search compatibility work is complete. +- `supports_web_search` overrides change recorded LiteLLM compatibility evidence only. They do not mutate Codex `supports_search_tool` and are not sufficient proof for hosted-search enablement. Hosted web search also depends on version-matched Codex provider/runtime semantics; see `foreign-web-search.md`. Directly contradictory settings inside one override table fail configuration validation. At minimum: