Repository navigation
Invocation log: count repeated failures on one row, short payloads past an hourly budget - #893
Merged
Merged
Conversation
…st an hourly budget One workspace (1885Data) is 802,628 of the 859,784 rows in tool_invocations and about 1.8 GB of it: their backend calls at up to 197 a minute and their own gateway answers 40% of those calls with the same 429. - A failure that repeats within 60 s with the same connector, tool and error (numbers blanked) is counted on the row already stored, in a new repeat_count column, instead of being stored again. Counts are written when the window closes, every 30 s and on shutdown. Successes are always stored. - Past INVOCATION_FULL_PAYLOADS_PER_HOUR (default 1000) rows per organisation per hour, input and output keep a 512-byte excerpt; status, timing and error stay complete. - Stats, analytics, usage breakdowns and server usage sum repeat_count, so the numbers shown stay the number of calls. The migration adds a column with a constant default: catalog-only on the cloud's Postgres 17, no table rewrite.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
tool_invocationson cloud: 859,784 rows, 4.4 GB. One workspace (1885Data) is 802,628 rows and ~1.8 GB of row data. In the last 24 h: 53,894 calls from their backend API key, peak 197/min, and 21,365 identical 429s from their own gateway (over_request_rate_limit). They have been emailed with the numbers and a caching recommendation.What
INVOCATION_REPEAT_WINDOW_SECONDS(60) →repeat_count += 1on the row already stored. Counts are buffered in memory and written when the window closes, every 30 s, on overflow (10k keys) and on shutdown. Successes are always stored (activation, KG, usage).INVOCATION_FULL_PAYLOADS_PER_HOUR(1000), input/output keep a 512-byte excerpt; status, duration and error stay complete.getStats,getAnalytics,getBreakdownsandusageByServersumrepeat_countinstead of counting rows.20261006090000_tool_invocation_repeat_count:ADD COLUMN repeat_count INTEGER NOT NULL DEFAULT 1. Constant default = catalog-only change on the cloud's Postgres 17, no rewrite of the 4.4 GB table.Expected effect on 1885's traffic: ~21k error rows/day become a few hundred (one per tool per minute), and their ~32k success rows/day keep full payloads only for the first 1000 per hour.
Not included: deleting their existing rows (256,768 errors, ~800k rows total). Retention (90 d) and payload trim (14 d) keep running; a one-off cleanup needs a separate OK.
Tests
hdb, missing locally only). tsc + eslint clean.