Skip to content

feat(models): discover context windows from Kiro catalog - #121

Open
leecoder wants to merge 4 commits into
tickernelz:masterfrom
leecoder:fix/overflow-response-for-compaction
Open

leecoder wants to merge 4 commits into
tickernelz:masterfrom
leecoder:fix/overflow-response-for-compaction

Conversation

@leecoder

@leecoder leecoder commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

Summary

Discover each account's model context window from Kiro's live model catalog instead of maintaining a hardcoded context-size table.

Changes

  • src/plugin/models.ts: Query GET https://q.{region}.amazonaws.com/ListAvailableModels?origin=AI_EDITOR with the selected account's bearer token. Read models[].tokenLimits.maxInputTokens, map resolved Kiro model IDs back to plugin aliases, cache results for 5 minutes, and fall back to the existing -1m alias heuristic when discovery is unavailable.
  • src/core/request/request-handler.ts: Refresh the catalog after selecting/refreshing the active account and before preparing the request, so response usage estimation uses the active account's actual limits.
  • src/__tests__/models.test.ts: Verify live API limits and alias mapping.

Verification

  • bun run build passed
  • bun test src/__tests__/models.test.ts passed (1 pass, 0 fail)
  • Live authenticated request verified ListAvailableModels returns tokenLimits.maxInputTokens (including 1M Sonnet 4.6 and account-specific limits for DeepSeek/MiniMax).

Note: bun test still has unrelated pre-existing failures in bearer-retry.test.ts and sdk-client.test.ts; the focused model test and build pass.

leecoder and others added 3 commits August 11, 2026 09:24
- Add MODEL_CONTEXT_WINDOWS map with explicit sizes per model
- getContextWindowSize() checks map first, falls back to isLongContextModel
- Add test coverage for context window size resolution
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
@leecoder leecoder changed the title feat(models): add per-model context window size lookup table feat(models): discover context windows from Kiro catalog Aug 11, 2026
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
lamiskin pushed a commit to lamiskin/opencode-kiro-auth that referenced this pull request Sep 3, 2026
Instead of hardcoding context sizes, query ListAvailableModels API to get
actual tokenLimits.maxInputTokens per model. Cached for 5 minutes per
account/region. Falls back to heuristics when unavailable.

Implementation from tickernelz PR tickernelz#121
RvVeen added a commit to Servoy/opencode-kiro-auth that referenced this pull request Sep 12, 2026
Implements upstream tickernelz#121.

Ask GET /ListAvailableModels?origin=AI_EDITOR for each model's
tokenLimits.maxInputTokens with the active account's token, and prefer
that over the hardcoded table in getModelContextLimit. Never awaited: it
only sharpens a token estimate, so no request waits on it. A failed or
empty catalog leaves the built-in limits untouched.

This matters because what Kiro advertises publicly has not always matched
what it serves — upstream tickernelz#128 measured opus-5 plateauing at ~166K
against a documented 1M — and the limits are account-specific. Asking the
service ends the guessing for whoever is running it.

Held for 30 minutes rather than the 5 the upstream PR used: context
windows change a few times a year, not a few times an hour. The cache is
keyed by region and profile ARN, so a different account refetches at once
instead of inheriting limits that are not its own.

Note the advertised registry limit is still built at plugin init, before
any token exists, so discovery corrects the token estimate immediately
and the advertised value only on a later start.

Also complete the logger mock in every test file that stubs it. bun's
mock.module is global, so a stub missing logApiRequest broke whichever
file happened to run after it.
RvVeen added a commit to Servoy/opencode-kiro-auth that referenced this pull request Sep 12, 2026
…kernelz#121)

Implements upstream tickernelz#121.

Ask GET /ListAvailableModels?origin=AI_EDITOR for each model's
tokenLimits.maxInputTokens with the active account's token, and prefer
that over the hardcoded table in getModelContextLimit. Never awaited: it
only sharpens a token estimate, so no request waits on it. A failed or
empty catalog leaves the built-in limits untouched.

This matters because what Kiro advertises publicly has not always matched
what it serves — upstream tickernelz#128 measured opus-5 plateauing at ~166K
against a documented 1M — and the limits are account-specific. Asking the
service ends the guessing for whoever is running it.

Held for 30 minutes rather than the 5 the upstream PR used: context
windows change a few times a year, not a few times an hour. The cache is
keyed by region and profile ARN, so a different account refetches at once
instead of inheriting limits that are not its own.

Note the advertised registry limit is still built at plugin init, before
any token exists, so discovery corrects the token estimate immediately
and the advertised value only on a later start.

Also complete the logger mock in every test file that stubs it. bun's
mock.module is global, so a stub missing logApiRequest broke whichever
file happened to run after it.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant