Skip to content

Invocation log: count repeated failures on one row, short payloads past an hourly budget - #893

Merged
keysersoft merged 1 commit into
mainfrom
keysersoft/invocation-log-dedupe
Oct 6, 2026
Merged

keysersoft merged 1 commit into
mainfrom
keysersoft/invocation-log-dedupe

Conversation

@keysersoft

Copy link
Copy Markdown
Contributor

Why

tool_invocations on cloud: 859,784 rows, 4.4 GB. One workspace (1885Data) is 802,628 rows and ~1.8 GB of row data. In the last 24 h: 53,894 calls from their backend API key, peak 197/min, and 21,365 identical 429s from their own gateway (over_request_rate_limit). They have been emailed with the numbers and a caching recommendation.

What

  • Repeated failures counted, not stored again. Same connector + tool + status + error text (digits blanked) within INVOCATION_REPEAT_WINDOW_SECONDS (60) → repeat_count += 1 on the row already stored. Counts are buffered in memory and written when the window closes, every 30 s, on overflow (10k keys) and on shutdown. Successes are always stored (activation, KG, usage).
  • Hourly full-payload budget per organisation. Past INVOCATION_FULL_PAYLOADS_PER_HOUR (1000), input/output keep a 512-byte excerpt; status, duration and error stay complete.
  • Numbers stay "calls". getStats, getAnalytics, getBreakdowns and usageByServer sum repeat_count instead of counting rows.
  • Migration 20261006090000_tool_invocation_repeat_count: ADD COLUMN repeat_count INTEGER NOT NULL DEFAULT 1. Constant default = catalog-only change on the cloud's Postgres 17, no rewrite of the 4.4 GB table.

Expected effect on 1885's traffic: ~21k error rows/day become a few hundred (one per tool per minute), and their ~32k success rows/day keep full payloads only for the first 1000 per hour.

Not included: deleting their existing rows (256,768 errors, ~800k rows total). Retention (90 d) and payload trim (14 d) keep running; a one-off cleanup needs a separate OK.

Tests

  • New: repeats counted on the first row, numbers-only differences treated as the same, different error / success stored separately, window expiry carries the count over, payload budget at row 1001.
  • Updated specs for aggregate-based counts. Backend suite: 6692 passed (HANA suite needs hdb, missing locally only). tsc + eslint clean.

…st an hourly budget

One workspace (1885Data) is 802,628 of the 859,784 rows in tool_invocations
and about 1.8 GB of it: their backend calls at up to 197 a minute and their
own gateway answers 40% of those calls with the same 429.

- A failure that repeats within 60 s with the same connector, tool and error
  (numbers blanked) is counted on the row already stored, in a new
  repeat_count column, instead of being stored again. Counts are written when
  the window closes, every 30 s and on shutdown. Successes are always stored.
- Past INVOCATION_FULL_PAYLOADS_PER_HOUR (default 1000) rows per organisation
  per hour, input and output keep a 512-byte excerpt; status, timing and
  error stay complete.
- Stats, analytics, usage breakdowns and server usage sum repeat_count, so
  the numbers shown stay the number of calls.

The migration adds a column with a constant default: catalog-only on the
cloud's Postgres 17, no table rewrite.
@keysersoft
keysersoft merged commit 5d0ac4a into main Oct 6, 2026
14 checks passed
@keysersoft
keysersoft deleted the keysersoft/invocation-log-dedupe branch October 6, 2026 08:16
@github-actions github-actions Bot locked and limited conversation to collaborators Oct 6, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant