fix(heartbeat): make the company WIP limit a reservation, not a count - #5
Conversation
The company-wide limit was enforced by counting running runs before claimQueuedRun, roughly eleven round trips before the conditional UPDATE that actually flips a run to running. The only mutex in that window, withAgentStartLock, is in-process and keyed by agent id, so two agents of one company take two different locks and both sail through: both read zero running runs, both claim, and the company spends two runs where the operator configured one. Pushing the count into the UPDATE's WHERE clause does not fix it either. Under READ COMMITTED both implicit transactions take their snapshot before the other commits, so both subqueries still see zero. The serialization point has to be a lock, not a query. withCompanyWipSlot takes a per-company pg_advisory_xact_lock, re-counts under it, and only then runs the caller's claim UPDATE inside the same transaction. It lives in the database so it holds across server processes and Postgres drops it automatically when the transaction ends, including when a process dies. Disabled by default (limit <= 0), so upstream behaviour is unchanged when the limit isn't configured. Wrapping the claim in a transaction means the losing agent's run stays queued instead of erroring, so it needs to be redispatched when the slot frees up. startNextQueuedRunsForCompanyPeers sweeps the company's other agents on run completion; without it the freed slot would sit idle until the next resumeQueuedRuns timer tick, turning the over-spend bug into a latency bug instead.
|
Reviewed before merging. The reservation itself is right, and the two things I checked most carefully hold up:
One gap worth recording rather than fixing here. That is not a regression: before this change a run blocked by the company limit stayed |
The completion path was the only place the freed company slot was handed to a waiting peer. Two other paths free the same slot: reapOrphanedRuns, when a run's process died and the run is finalized as failed, and cancelRun, when an operator stops a running run. Both re-dispatched only the agent whose run ended, and the agent waiting on the company slot is by definition a different one, so the slot sat idle until the next resumeQueuedRuns tick. Self-healing on a timer, so not a correctness bug — but it is exactly the latency the peer sweep exists to avoid, and leaving two of the three paths uncovered would have made the sweep look complete when it wasn't.
|
Follow-up commit The peer sweep was wired into one slot-freeing path — the normal completion path at the end of a run execution. Two others free the same company slot and were left uncovered:
In both cases the freed slot sat idle until the next Re-validated after the change: |
The bug
The company-wide WIP limit was enforced by counting
runningruns incountRunningRunsForCompany, called fromstartNextQueuedRunForAgentroughly eleven database round trips before the conditionalUPDATEinclaimQueuedRunthat actually flips a run torunning. The only mutex in that window,withAgentStartLock, is in-process and keyed by agent id — two agents belonging to the same company take two different locks and both sail straight through. Both read zero running runs, both claim, and the company spends two runs where the operator configured a limit of one. It fails on money, not just correctness.Pushing the count into the
UPDATE'sWHEREclause does not fix it either. Under READ COMMITTED, each implicit transaction takes its snapshot before the other commits, so both subqueries still see zero running runs regardless of where the count is evaluated. The serialization point has to be a lock, not a smarter query.The fix
withCompanyWipSlot(new,server/src/services/company-wip-limit.ts) takes a per-companypg_advisory_xact_lock(hashtextextended(key, 0)), re-counts running runs under that lock, and only then runs the caller's claimUPDATEinside the same transaction. It lives in the database, so it holds across server processes, and Postgres releases it automatically when the transaction ends — including when a process dies, which a counter table would not do. Disabled by default (limit <= 0), so upstream behaviour is unchanged when the limit isn't configured.heartbeat.ts'sclaimQueuedRunnow runs its claimUPDATEthroughwithCompanyWipSlot. Since a losing claim now returnsnull(run staysqueued) instead of racing to a duplicaterunningrow, the loser needs to be redispatched once the winner's slot frees up.startNextQueuedRunsForCompanyPeerssweeps the company's other agents with queued runs when a run completes; without it the freed slot would sit idle until the nextresumeQueuedRunstimer tick, turning the over-spend bug into a latency bug instead.Tests
server/src/__tests__/heartbeat-company-wip-limit-race.test.ts(new, 3 tests, embedded Postgres):Reverted to the pre-fix code (
git stashon both changed files, keeping the test file) and re-ran the suite: all 3 failed —withCompanyWipSlot is not a function(function doesn't exist yet)expected false to be true(both agents ended up running)expected null not to be null(loser never got re-queued)Restored the fix and re-ran: all 3 green again. The test is not a tautology — it fails without the fix and passes with it.
Verified
pnpm --filter @paperclipai/server typecheckheartbeat-company-wip-limit-race.test.ts)company-wip-limit.test.ts(existing unit tests)Risks
Low. The limit is opt-in (disabled unless
companyMaxConcurrentRuns> 0), so this only changes behavior for operators who already configured a company-wide cap — and for them it closes an over-spend hole rather than opening one. The advisory-lock transaction body is deliberately three statements (lock, re-count, delegate to caller's UPDATE) with no budget checks or adapter calls inside it, to avoid serializing slow work across the whole company or deadlocking against the longer-livedpaperclip:folders:*locks.