Combining Genetic Algorithms and User Interaction Recordings to Enhance Automated App Testing — Replication Package
This private repository is the replication package for the master thesis Combining Genetic Algorithms and User Interaction Recordings to Enhance Automated App Testing. It converts the thesis working archive into a documented, machine-checkable workflow containing curated CSV inputs, validation rules, statistical analyses, figures, tables, and tests.
The manuscript itself is intentionally not distributed in this repository.
| Component | Included evidence |
|---|---|
| Subject applications | 7 open-source Android applications |
| Record–replay observations | 53 scenarios × 10 replays = 530 rows |
| Search experiment observations | 2,240 independently seeded runs |
| Repetitions | 10 random seeds per experimental condition |
| Coverage collected | Activity, line, branch, and method coverage |
| Primary search-analysis outcome | Final branch coverage |
| Statistical methods | Median/IQR, two-sided Mann–Whitney U, Vargha–Delaney A12 |
| Generated outputs | 6 figures, 9 CSV/LaTeX table pairs, validation report, run manifest |
Together, the curated inputs contain 2,770 run-level observations. They contain no personal filesystem paths, workstation nicknames, virtual environments, emulator images, APK binaries, or thesis PDF.
The order and scope below follow the thesis.
| ID | Research question | Primary evidence |
|---|---|---|
| RQ1 | Does replaying a recorded user interaction trace result in a UI state similar to the originally recorded state? | Cosine similarity for 530 replays, summarized first by scenario and then by application |
| RQ2 | How does seeding the GA's population with user traces impact test coverage in MIO and MOSA? | Six seed fractions for MIO and MOSA; final branch coverage used for comparisons and global selection |
| RQ3 | What is the best probability of replaying user traces during mutation given a fixed similarity threshold in MIO? | Six replay probabilities at fixed τ = 0.5 |
| RQ4 | What is the best mutation configuration in MIO in terms of similarity threshold and replay probability? | Stage 1 selects the threshold; Stage 2 selects replay probability at the chosen threshold |
| RQ5 | How does the best seeding-and-mutation configuration compare to the baseline GA in MIO? | Per-application comparison of the tuned configuration and trace-free baseline |
Detailed operational definitions are in docs/methodology.md.
flowchart LR
A["Final experiment exports"] --> B["Curate and give stable names"]
B --> C["Validate schemas and experimental keys"]
C --> D["Summarize all four coverage measures"]
D --> E["Analyze final branch coverage"]
E --> F["Mann–Whitney U and A12"]
F --> G["Figures and CSV/LaTeX tables"]
G --> H["Checksummed run manifest"]
Every tracked result can be regenerated from data/raw/. Generated artifacts are
also versioned so reviewers can inspect the evidence without first configuring a
Python environment.
The 53 scenarios were each replayed 10 times. Cosine similarity is summarized by the scenario median before application-level aggregation, matching the thesis. All seven application medians are 1.0; lower individual replays remain visible in the machine-readable summaries.
Activity, line, branch, and method coverage are retained in descriptive summaries. The inferential comparisons and the global parameter-selection rule use final branch coverage. The selected global seed fractions are 0.8 for MIO and 1.0 for MOSA.
Stage 1 selects cosine threshold τ = 0.2. Stage 2 then compares replay
probabilities while holding that threshold fixed and selects 0.4. Selection uses
the highest mean of the seven per-application median final branch coverages, with
the smaller parameter value resolving an exact tie.
The tuned MIO configuration uses seed fraction 0.8, cosine threshold 0.2, and replay probability 0.4. Its pooled final branch-coverage median is 54.58%, compared with 50.06% for the baseline, and its per-application median is higher in six of the seven applications. Application-level tests and A12 effect sizes prevent the pooled values from hiding heterogeneity.
The pipeline follows the thesis analysis:
- each experimental condition contains 10 independently seeded runs;
- central tendency and spread are reported with median and IQR;
- defined conditions are compared using a two-sided, unpaired Mann–Whitney U test;
- Vargha–Delaney A12 reports effect size and direction, with 0.5 as neutral; and
- statistical significance uses the thesis threshold
p < 0.05.
No Benjamini–Hochberg/FDR adjustment is added because it is not part of the thesis method. P-values should be interpreted together with A12, median differences, and per-application plots rather than as a substitute for practical importance.
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e ".[dev]"
thesis-testing-artifact validate
thesis-testing-artifact reproduce
pytestEquivalent commands are available as make validate, make reproduce, and
make test. Outputs are written to data/processed/, results/figures/,
results/tables/, and results/run_manifest.json.
.
├── data/
│ ├── raw/ # Curated, immutable analysis inputs
│ └── processed/ # Reproducible summaries, selections, and tests
├── docs/
│ ├── data_dictionary.md
│ ├── methodology.md
│ ├── provenance.md
│ └── reproducibility.md
├── results/
│ ├── figures/ # Six reproducible PNG figures
│ ├── tables/ # Reviewer-friendly CSV and LaTeX tables
│ └── run_manifest.json # Versions, row counts, and SHA-256 checksums
├── src/gui_testing_artifact/
│ ├── analysis.py # RQ-specific analysis orchestration
│ ├── plots.py # Figure generation
│ ├── statistics.py # Mann–Whitney U and A12
│ └── validation.py # Dataset contracts and integrity checks
├── tests/ # Unit, integrity, privacy, and artifact checks
└── tools/curate_from_archive.py
The working archive used date-based filenames and local labels. Those names are not reproduced here. Tracked filenames are stable, lowercase, English, and state the scientific role of each dataset. The source archive is never modified. See docs/provenance.md for the curation and exclusion policy.
- The same seven applications appear throughout RQ1–RQ5.
- The repository reproduces the final reported analyses; large execution logs, emulator images, APKs, decompiled code, failed runs, and exploratory copies are excluded.
- All four coverage types are retained, but RQ2–RQ5 hypothesis tests and global parameter selection use final branch coverage as specified in the thesis.
- Mann–Whitney U preserves continuity with the thesis. A paired sensitivity analysis could be added separately if random seeds are treated as strict pairs.
- Findings concern the studied applications and configurations and should not be generalized to all Android applications without further replication.
Tests cover known A12 cases, expected row counts, 10-replay/10-seed condition matrices, experimental-key uniqueness, RQ4's fixed threshold, generated outputs, and removal of personal path leakage. GitHub Actions validates the curated data, runs the tests, and reproduces every result on each push and pull request.
See docs/reproducibility.md for the reviewer checklist.
Citation metadata is provided in CITATION.cff. Add the final institution, award year, and persistent identifier before public release if those details have been assigned.
Code is released under the MIT License. Curated experimental observations are provided for research reproducibility; third-party Android applications remain subject to their original licenses.



