Skip to content

feat: say why a scenario's processes died - #652

Open
ykhrustalev wants to merge 3 commits into
sourcefrog:mainfrom
ykhrustalev:ykhrustalev/report-process-death
Open

ykhrustalev wants to merge 3 commits into
sourcefrog:mainfrom
ykhrustalev:ykhrustalev/report-process-death

Conversation

@ykhrustalev

Copy link
Copy Markdown

Problem
A mutant caught because the kernel killed its tests looks, in the counts, exactly like one caught by a failing assertion — which is the thing you want to know when diagnosing a run.

Solution

  • Carries the process-group sweep's findings on each phase result.
  • Renders the killing signal (by name) and anything reaped in parentheses on the outcome line, in the scenario log, and as a sweep field in outcomes.json.
  • Classification is unchanged.

Testing

  • Unit tests for the rendering; the sweep integration test now also checks that the outcome line names the strays.

Stacked on #648#650; only the last commit is new here. Split out of #647 as requested.

On a timeout, terminate_child() sent one SIGTERM and then blocked in wait()
with no bound. A cargo process that ignores the signal, or that has been
stopped and so never receives it, hung the whole run at that point.

Wait only for a short grace period, then SIGKILL the process group. This
needs a second signal, so the errno handling moves into a signal_group()
helper rather than being duplicated.
Tests can leave processes running after they exit: a helper a test spawned and
forgot, or a test binary that was never reaped. Until now cargo-mutants only
signalled the child's process group on a timeout, so on a normal exit those
processes survived and kept allocating while later mutants were tested, in a
window where no cargo phase is running at all.

Sweep the group in the process layer, so every phase gets it: once the direct
child exits with any status, SIGTERM the group, wait a bounded grace period,
then SIGKILL whatever is left. What was reaped goes into the scenario log, and
the pids into the debug log, enumerated from /proc on Linux.

Windows has no process groups and the job object equivalent is out of scope,
so the sweep is a no-op there.
A mutant caught because its tests were killed is indistinguishable, in the
counts, from one caught by a failing assertion -- which is exactly what you
want to know when diagnosing a run.

Carry the process group sweep's findings on each phase result, and render them,
with the name of any signal that killed the phase, in parentheses on the
outcome line, in the scenario log, and in outcomes.json.

Classification is deliberately untouched: this only makes the reason visible.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant