What happened
make check — the gate that stands between a change and a release — failed on:
FAILED tests/integration/test_reports.py::test_report_json_flag_emits_clean_json_on_stdout
=========== 1 failed, 1633 passed, 4 skipped, 23 warnings in 40.34s ============
make: *** [Makefile:145: test] Error 1
It then passed on every subsequent attempt. Measured: 1 failure in 8 full-suite
executions (3 × make check, 5 × uv run python -m pytest tests/), all on the same
commit and the same machine.
The test passes reliably in isolation:
uv run python -m pytest "tests/integration/test_reports.py::test_report_json_flag_emits_clean_json_on_stdout" -q
# 1 passed
Why this is worth fixing even though it is rare
A gate that fails roughly one run in eight will eventually fail on a release, and the
honest response to a red gate whose cause is unknown is to stop and investigate — so this
costs a release cycle each time it fires. It also teaches the habit this project has
explicitly tried to avoid elsewhere: re-running a red gate until it goes green.
What is known
The test asserts that aobench report json <dir> --json writes exactly one line to
stdout:
result = runner.invoke(app, ["report", "json", str(run_dir), "--json"])
assert result.exit_code == 0, result.output
# No "Report written:" banner, no human summary lines, no trailing blank line.
assert result.output.count("\n") == 1
data = json.loads(result.output)
assert data["run_id"] == run_id
assert data["task_count"] >= 10
(tests/integration/test_reports.py:110-126)
So any single extra line on stdout fails it. Note what the test does and does not isolate:
it uses tmp_path for the run directory, but CliRunner captures process-wide stdout, and
logging configuration is global to the interpreter.
Which of the four assertions failed was not captured — that is the first thing to find
out, and it splits the problem cleanly:
output.count("\n") == 1 failing means something else wrote to stdout. The leading
hypothesis is global state leaking from an earlier test — a logging handler pointed at
stdout, or a warning emitted once per interpreter under a condition an earlier test sets
up. This would explain passing in isolation and failing in a suite.
task_count >= 10 failing means the helper _run_all_tasks produced fewer scored tasks
than expected, which is a different bug entirely and a more serious one.
Suggested approach
- Reproduce with the assertion visible. Running the suite with
-p no:randomly is not
needed — ordering here is already deterministic (only anyio and cov are
installed), which is itself informative: same order, different outcome means real
state or environment dependence rather than order dependence.
- A cheap way to catch it: run the suite in a loop until it fails, keeping the long
traceback, e.g. `for i in $(seq 20); do uv run python -m pytest tests/ -q --tb=long
run_$i.log 2>&1 || break; done`.
- If it is stdout pollution, the fix is likely to make the test assert on its own
captured output rather than trusting process-wide capture, and to stop whatever
leaks the handler. Please fix the leak, not only the assertion — a test hardened against
pollution still leaves the pollution for the next test.
- Do not mark it
xfail, add a retry plugin, or loosen the count("\n") == 1 assertion.
That assertion is the entire point of the test: --json exists so the output can be
piped into jq, and one stray banner line breaks every downstream consumer.
Acceptance criteria
Effort
Medium, and genuinely uncertain — the hard part is reproduction, not the patch. This is
deliberately not labelled a good first issue: an intermittent failure with no known
trigger is a poor first experience. If you have chased a flaky test before, this is a good
one, and #65 and #70 in this repo are the same family of problem (a failure whose location
tells you nothing about its cause).
What happened
make check— the gate that stands between a change and a release — failed on:It then passed on every subsequent attempt. Measured: 1 failure in 8 full-suite
executions (3 ×
make check, 5 ×uv run python -m pytest tests/), all on the samecommit and the same machine.
The test passes reliably in isolation:
Why this is worth fixing even though it is rare
A gate that fails roughly one run in eight will eventually fail on a release, and the
honest response to a red gate whose cause is unknown is to stop and investigate — so this
costs a release cycle each time it fires. It also teaches the habit this project has
explicitly tried to avoid elsewhere: re-running a red gate until it goes green.
What is known
The test asserts that
aobench report json <dir> --jsonwrites exactly one line tostdout:
(
tests/integration/test_reports.py:110-126)So any single extra line on stdout fails it. Note what the test does and does not isolate:
it uses
tmp_pathfor the run directory, butCliRunnercaptures process-wide stdout, andloggingconfiguration is global to the interpreter.Which of the four assertions failed was not captured — that is the first thing to find
out, and it splits the problem cleanly:
output.count("\n") == 1failing means something else wrote to stdout. The leadinghypothesis is global state leaking from an earlier test — a logging handler pointed at
stdout, or a warning emitted once per interpreter under a condition an earlier test sets
up. This would explain passing in isolation and failing in a suite.
task_count >= 10failing means the helper_run_all_tasksproduced fewer scored tasksthan expected, which is a different bug entirely and a more serious one.
Suggested approach
-p no:randomlyis notneeded — ordering here is already deterministic (only
anyioandcovareinstalled), which is itself informative: same order, different outcome means real
state or environment dependence rather than order dependence.
traceback, e.g. `for i in $(seq 20); do uv run python -m pytest tests/ -q --tb=long
captured output rather than trusting process-wide capture, and to stop whatever
leaks the handler. Please fix the leak, not only the assertion — a test hardened against
pollution still leaves the pollution for the next test.
xfail, add a retry plugin, or loosen thecount("\n") == 1assertion.That assertion is the entire point of the test:
--jsonexists so the output can bepiped into
jq, and one stray banner line breaks every downstream consumer.Acceptance criteria
extra line (or the low count).
the failure no longer reproducible that way. A flake declared fixed without ever
having been reproduced is not fixed.
make checkgreen, anduv run python -m pytest tests/green over at least 20consecutive runs.
xfail, no retry/rerun plugin, andoutput.count("\n") == 1still asserted.Effort
Medium, and genuinely uncertain — the hard part is reproduction, not the patch. This is
deliberately not labelled a good first issue: an intermittent failure with no known
trigger is a poor first experience. If you have chased a flaky test before, this is a good
one, and #65 and #70 in this repo are the same family of problem (a failure whose location
tells you nothing about its cause).