Skip to content

Derive RoadRunner worker defaults from the container CPU limit - #2662

Closed
torvalstrom wants to merge 1 commit into
shlinkio:developfrom
torvalstrom:container-aware-worker-defaults
Closed

torvalstrom wants to merge 1 commit into
shlinkio:developfrom
torvalstrom:container-aware-worker-defaults

Conversation

@torvalstrom

Copy link
Copy Markdown

Fixes the default worker count described in #2661.

RoadRunner resolves num_workers: 0 against the number of CPUs the host
reports, which is unaffected by a CPU limit set on the container. On a 24-core
host with the container limited to 1 CPU that is 24 web + 24 task workers at
~50-60MB each — about 1.2GB before a single request is served, and 48 workers
that can never run concurrently.

This reads the cgroup CPU quota in the entrypoint and derives the defaults from
it. Behaviour is unchanged when there is no CPU limit, and an explicitly set
WEB_WORKER_NUM / TASK_WORKER_NUM still wins.

Web workers get a floor of 2, so a single slow request cannot block the whole
container when the limit is one CPU. Happy to drop that floor if you would
rather keep strict one-per-CPU.

Tested

Both cgroup v1 and v2 paths are read; verified in busybox sh (alpine:3.20) and
dash (debian:12-slim):

container detected CPUs WEB_WORKER_NUM TASK_WORKER_NUM
no limit (none) unset unset
--cpus=0.5 1 2 1
--cpus=1 1 2 1
--cpus=3 3 3 3
--cpus=4 4 4 4
--cpus=1 -e WEB_WORKER_NUM=9 1 9 1

On the instance that prompted #2661, setting these by hand took the same
container from 50 processes / 1262MB to 7 / 183MB with no change in behaviour.

If you would rather solve this with documentation than with a behaviour change,
I am happy to convert this into a docs PR instead.

RoadRunner resolves num_workers: 0 against the number of CPUs the host
reports, which is unaffected by any CPU limit set on the container. On a
24-core host with the container limited to 1 CPU that starts 24 web plus 24
task workers, at roughly 50-60MB each, so the container idles around 1.2GB
before serving a request and 48 of those workers can never run concurrently.

When a CPU limit is present, derive the defaults from it instead. Without a
limit nothing changes, and an explicitly set WEB_WORKER_NUM or TASK_WORKER_NUM
still wins.

Web workers get a floor of 2 so that a single slow request cannot block the
whole container when the limit is one CPU.

Refs shlinkio#2661
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants