Skip to content

Flaky release gate: test_report_json_flag_emits_clean_json_on_stdout fails ~1 run in 8 #72

Description

@MSKazemi

What happened

make check — the gate that stands between a change and a release — failed on:

FAILED tests/integration/test_reports.py::test_report_json_flag_emits_clean_json_on_stdout
=========== 1 failed, 1633 passed, 4 skipped, 23 warnings in 40.34s ============
make: *** [Makefile:145: test] Error 1

It then passed on every subsequent attempt. Measured: 1 failure in 8 full-suite
executions
(3 × make check, 5 × uv run python -m pytest tests/), all on the same
commit and the same machine.

The test passes reliably in isolation:

uv run python -m pytest "tests/integration/test_reports.py::test_report_json_flag_emits_clean_json_on_stdout" -q
# 1 passed

Why this is worth fixing even though it is rare

A gate that fails roughly one run in eight will eventually fail on a release, and the
honest response to a red gate whose cause is unknown is to stop and investigate — so this
costs a release cycle each time it fires. It also teaches the habit this project has
explicitly tried to avoid elsewhere: re-running a red gate until it goes green.

What is known

The test asserts that aobench report json <dir> --json writes exactly one line to
stdout:

result = runner.invoke(app, ["report", "json", str(run_dir), "--json"])
assert result.exit_code == 0, result.output
# No "Report written:" banner, no human summary lines, no trailing blank line.
assert result.output.count("\n") == 1
data = json.loads(result.output)
assert data["run_id"] == run_id
assert data["task_count"] >= 10

(tests/integration/test_reports.py:110-126)

So any single extra line on stdout fails it. Note what the test does and does not isolate:
it uses tmp_path for the run directory, but CliRunner captures process-wide stdout, and
logging configuration is global to the interpreter.

Which of the four assertions failed was not captured — that is the first thing to find
out, and it splits the problem cleanly:

  • output.count("\n") == 1 failing means something else wrote to stdout. The leading
    hypothesis is global state leaking from an earlier test — a logging handler pointed at
    stdout, or a warning emitted once per interpreter under a condition an earlier test sets
    up. This would explain passing in isolation and failing in a suite.
  • task_count >= 10 failing means the helper _run_all_tasks produced fewer scored tasks
    than expected, which is a different bug entirely and a more serious one.

Suggested approach

  1. Reproduce with the assertion visible. Running the suite with -p no:randomly is not
    needed — ordering here is already deterministic (only anyio and cov are
    installed), which is itself informative: same order, different outcome means real
    state or environment dependence rather than order dependence.
  2. A cheap way to catch it: run the suite in a loop until it fails, keeping the long
    traceback, e.g. `for i in $(seq 20); do uv run python -m pytest tests/ -q --tb=long

    run_$i.log 2>&1 || break; done`.

  3. If it is stdout pollution, the fix is likely to make the test assert on its own
    captured output rather than trusting process-wide capture, and to stop whatever
    leaks the handler. Please fix the leak, not only the assertion — a test hardened against
    pollution still leaves the pollution for the next test.
  4. Do not mark it xfail, add a retry plugin, or loosen the count("\n") == 1 assertion.
    That assertion is the entire point of the test: --json exists so the output can be
    piped into jq, and one stray banner line breaks every downstream consumer.

Acceptance criteria

  • The cause is identified and stated — which assertion failed, and what produced the
    extra line (or the low count).
  • Red-green evidence: a way to make the failure happen on demand, then the fix, then
    the failure no longer reproducible that way. A flake declared fixed without ever
    having been reproduced is not fixed.
  • make check green, and uv run python -m pytest tests/ green over at least 20
    consecutive runs.
  • No xfail, no retry/rerun plugin, and output.count("\n") == 1 still asserted.

Effort

Medium, and genuinely uncertain — the hard part is reproduction, not the patch. This is
deliberately not labelled a good first issue: an intermittent failure with no known
trigger is a poor first experience. If you have chased a flaky test before, this is a good
one, and #65 and #70 in this repo are the same family of problem (a failure whose location
tells you nothing about its cause).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area: infraBuild, CI, packaging, deploymentbugSomething isn't workingeffort: mediumRoughly a dayhelp wantedExtra attention is needed

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions