Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions .github/workflows/windows-beta.yml
Original file line number Diff line number Diff line change
Expand Up @@ -50,9 +50,15 @@ jobs:
- name: Build Electron application
run: npm run build

- name: Run Electron overlay smoke tests
run: npx playwright test

- name: Evaluate Milestone 4 retrieval offline
run: npm run eval:m4

- name: Validate Milestone 6 live-campaign budget offline
run: npm run eval:m6:preflight

- name: Create unsigned NSIS installer
run: npx electron-builder --win nsis --publish never
env:
Expand All @@ -61,11 +67,17 @@ jobs:
- name: Verify packaged Electron SQLite FTS5
run: npm run test:packaged-fts

- name: Verify packaged Windows helper health
run: npm run test:packaged-helper

- name: Upload Windows installer and Milestone 4 report
uses: actions/upload-artifact@v4
with:
name: PresenterAI-Windows-beta
path: |
release/PresenterAI-*-setup.exe
artifacts/m4/m4-retrieval-report.json
docs/validation/milestones-0-2.md
docs/validation/milestone-5.md
docs/validation/milestone-6.md
if-no-files-found: error
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,7 @@ dist/
release/
.tsbuild/
coverage/
test-results/
artifacts/
resources/windows-helper/
native/**/bin/
Expand Down
16 changes: 16 additions & 0 deletions docs/capture-compatibility/matrix.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,22 @@

Record Windows build, PresenterAI/Electron version, capture application version, GPU setting, monitor layout, and date for every run. First record the **protection OFF control**, then repeat with protection ON.

Current campaign status: **Untested. No capture path is verified on the current build.**

## Campaign environment

- Campaign ID:
- Date/time and tester:
- Git commit / PresenterAI version:
- Windows edition/build:
- Electron version:
- Chrome / Meet version:
- OBS version and capture backend:
- GPU / driver / graphics preference:
- Monitor count, layout, scaling, connection types:

If any environment value changes, start a new campaign block rather than silently updating completed rows.

| Capture path | OFF control | ON result | Version / setup | Notes |
|---|---|---|---|---|
| Google Meet — entire screen | Untested | Untested | | Primary display-capture test |
Expand Down
28 changes: 28 additions & 0 deletions docs/manual/windows-beta-validation.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,19 @@

Do not infer any result from the protection request or Electron's reported state. Record exact versions and test protection OFF before protection ON.

This is the operator runbook. Record immutable results in:

- `docs/validation/milestones-0-2.md` for early Windows and capture prerequisites;
- `docs/validation/milestone-5.md` for push-to-listen;
- `docs/validation/milestone-6.md` for transcription-to-response;
- `docs/capture-compatibility/matrix.md` for capture-path observations.

Milestones 3 and 4 are already accepted in their respective reports. Do not spend API credits rerunning M3 or repeat the offline M4 corpus as manual evidence.

Milestones 5 and 6 are **not accepted**. The current non-billable gate passes 196 Vitest tests in 28 files, 29/29 .NET tests, 5/5 Playwright Electron tests, an audit with zero high-severity findings, and the 50/50 M4 corpus. Those results do not replace the fullscreen, capture-path, Meet, physical-device, or installed-app rows below.

> **M6 paid campaign safety stop:** the zero-network preflight estimates $0.145567, but its documented worst-case bound is $0.641867. Because that exceeds the immutable $0.15 cap, it reports `strictCampaignFeasible=false`, sets `billableExecutionEnabled=false`, and refuses live execution. M6 has spent $0. Do not supply the API key or run billable cases until the user separately revises the case count or cap.

## Environment

- Date and tester:
Expand Down Expand Up @@ -31,6 +44,21 @@ Run the matrix once for speakers, wired headphones, and each available Bluetooth

The local command `npm run test:helper-smoke` is an automated format and lifecycle check. It is not a substitute for intelligibility testing with a live Meet session.

The latest smoke succeeded on the Realtek default endpoint (12.94 seconds, 414,126 bytes, 16 kHz mono, WAV removed). A first attempt against a stale Bluetooth endpoint was invalidated by Windows; enumeration then exposed Realtek as the default and allowed retry. Record this as a diagnostic only—do not mark device-removal, default-switch, or Meet rows passed from it.

The final campaign is stricter than the compact table above: perform 50 physical shortcut cycles, designate 20 for real Meet intelligibility/transcription review, and designate ten of those for the complete response pipeline. Enter individual case IDs and measurements in the M5/M6 records. Do not store audio, transcripts, prompts, or answers as evidence.

## Required execution order

1. Run the complete non-billable automated gate and packaged helper smoke.
2. Complete multi-monitor, fullscreen, shortcut-conflict, and capture OFF/ON rows for M0–M2.
3. Complete all 50 M5 shortcut cycles and the physical endpoint matrix.
4. Sign off M5 only if its strict thresholds pass.
5. Stop before the paid M6 phase. Obtain a separate user decision that changes the case count or budget cap enough for the strict documented bound; the existing $0.15 authorization cannot start the current 20-case campaign.
6. After a revised plan is explicitly authorized, run only that approved campaign sequentially with zero SDK retries. Stop on the first infrastructure/budget failure and leave later rows `Untested`.
7. Calculate p50/p95 only from operation-scoped production-app renderer acknowledgements for the pre-designated full-pipeline cases. A standalone/imported timing value or internal pipeline `total` is diagnostic and cannot satisfy release-to-visible acceptance.
8. Sign off M6 only if every gate passes within the newly authorized strict budget.

## Installer matrix

Use a non-production test account or VM for upgrade testing.
Expand Down
73 changes: 73 additions & 0 deletions docs/validation/milestone-5.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,73 @@
# Milestone 5 push-to-listen validation record

Status: **NOT ACCEPTED — automated gate passed; M0–M2 and physical-device validation remain pending.**

M5 may be accepted only after M0–M2 are signed off in `milestones-0-2.md`. This record validates bounded system WASAPI loopback and a restricted hold shortcut; it does not validate process-specific capture, microphone fallback, or continuous listening.

## Build and environment

- Campaign ID:
- Git commit / installer SHA-256:
- Date/time and tester:
- Windows edition/build:
- Helper protocol/features:
- Chrome/Meet version:
- Output endpoints and drivers:
- Available endpoint classes: speakers / wired / Bluetooth / USB
- Unavailable endpoint classes and reason:

Latest full regression run: 2026-07-16 on Windows 11 build 26200, Electron 43.1.0. The complete non-billable gate passed 196 Vitest tests in 28 files, 29/29 .NET tests, 5/5 Playwright Electron tests, a high-severity dependency audit with zero findings, and the accepted 50/50 M4 retrieval corpus. The packaged probes and physical helper smoke described below were recorded on 2026-07-14.

The latest helper smoke enumerated the available render endpoints and produced a 12,910 ms, 413,166-byte, 16 kHz mono PCM recording from the current Realtek default endpoint, then removed the WAV. An initial attempt against a previously selected Bluetooth endpoint failed when Windows invalidated that endpoint; subsequent visible enumeration identified the Realtek default and allowed a successful retry. This is useful device-change diagnostic evidence only. It does **not** pass the required in-app endpoint-removal/default-switch recovery row or prove Meet intelligibility.

## Automated evidence

| Case ID | Gate | Result | Evidence / notes |
|---|---|---|---|
| M5-AUTO-01 | Audit, typecheck, Vitest, helper tests, Playwright, M4 regression, production build | Pass | Audit: 0 high-severity findings; Vitest: 28 files / 196 tests; .NET: 29/29; Playwright: 5/5; M4: 50/50; typecheck and production build pass |
| M5-AUTO-02 | Packaged helper v2 handshake and required features | Pass | The final unpacked PresenterAI executable launched and reported protocol v2 plus all 9 required features from its bundled helper |
| M5-AUTO-03 | Rapid release during startup is latched exactly once | Pass | Deterministic controller race tests |
| M5-AUTO-04 | Cancel during startup/finalization reaches one terminal cleanup | Pass | Coordinator/controller cleanup and cancellation-boundary tests |
| M5-AUTO-05 | Stale/duplicate events cannot change a newer operation | Pass | Operation-ID and duplicate-terminal tests |
| M5-AUTO-06 | Idle crash restarts once; active/second crash fails safely | Pass | Helper crash/restart controller tests |
| M5-AUTO-07 | Missing endpoint falls back once to the current default with warning | Pass (automated) | Endpoint-loss and single-fallback tests; the Bluetooth-to-Realtek smoke transition is diagnostic only and does not replace the in-app manual row |
| M5-AUTO-08 | WAV is 16 kHz mono PCM, 250 ms–90 s, with one final file | Pass | Native conversion/limit tests, byte-level validator tests, and real helper smoke |
| M5-AUTO-09 | Cancel, failure, exit, and stale-startup cleanup remove owned files | Pass | Cleanup-path tests; real smoke WAV confirmed absent after terminal cleanup |

## Fifty-cycle shortcut campaign

Record one row per physical cycle. `Start events` and `terminal events` must both equal one. `Indicator ms` is measured from confirmed helper capture start to the renderer's first-frame acknowledgement.

| Trial | Endpoint | Scenario | Start events | Terminal events | Indicator ms | WAV valid | Temp removed | Result / notes |
|---:|---|---|---:|---:|---:|---|---|---|
| 01–50 | | normal / rapid release / autorepeat / Esc / recovery | | | | | | Untested |

Required aggregate:

- 50/50 trials have exactly one start and one terminal event.
- Listening indicator is confirmed within 150 ms in all timed normal-start trials.
- Esc and every failure/cancel path remove temporary audio.
- Listening is OFF after every launch and restart.

## Meet and endpoint matrix

Use 20 designated shortcut trials with real reviewer speech in Google Meet. Test each available endpoint class; unavailable hardware is recorded as unavailable, never passed.

| Endpoint / trial IDs | Trials | Intelligible | Device removal/default switch | Recoverable UI | Result / notes |
|---|---:|---:|---|---|---|
| Speakers | | | | | Untested |
| Wired headphones | | | | | Untested |
| Bluetooth | | | | | Untested |
| USB | | | | | Untested |

Acceptance requires intelligible reviewer speech in at least 19/20 designated Meet trials and a visible recoverable outcome for endpoint removal/default switching.

## Decision

- Aggregate starts/stops:
- Meet intelligibility:
- Maximum confirmed indicator latency:
- Temporary-file failures:
- Blocking case IDs:
- Human sign-off / date:
- Decision: **Untested**
72 changes: 72 additions & 0 deletions docs/validation/milestone-6.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,72 @@
# Milestone 6 transcription-to-response validation record

Status: **NOT ACCEPTED — automated gate passed; live validation is safety-blocked by the strict budget and Milestone 5/manual prerequisites.**

This gate reuses the immutable 20 audio captures designated in the M5 campaign. It validates bounded transcription, M4 retrieval, one M3-compatible Responses request, cancellation, privacy, and latency. It does not evaluate Terra, Realtime transcription, diarization, continuous listening, or process-specific capture.

## Budget and privacy controls

- Model: `gpt-4o-mini-transcribe`; answer mode: Normal/Luna only.
- Lifetime M6 live cap: **$0.15**; SDK retries: **zero**.
- The paid validation corpus uses short reviewer questions and rejects clips over **20 seconds before upload**. That reduces the practical estimate but cannot create a strict provider-token ceiling; the product's separate M5 capture bound remains 90 seconds.
- M3's historical evaluation report recorded **$0.206648**. This M6 campaign has made **zero network requests and spent $0**; PresenterAI does not inspect or claim the provider account's current balance, which may reflect unrelated usage.
- Stop before each request if conservative projected spend would exceed the cap.
- Stop on authentication, quota, rate-limit, timeout, network, or budget failure; do not rerun automatically.
- Persist only case IDs, pass/fail flags, model IDs, timings, token usage, price-version metadata, estimated cost, and failed IDs.
- Never persist credentials, audio, transcripts, prompts, answers, or reasoning content.

Latest offline preflight: 2026-07-16, 20-case corpus / 10 full-pipeline cases, **zero network requests**.

- Practical projected estimate: **$0.145567**.
- Documented worst-case bound: **$0.641867**.
- Immutable live cap: **$0.15**.
- `strictCampaignFeasible=false`.
- `billableExecutionEnabled=false`.

The current [model specification](https://developers.openai.com/api/docs/models/gpt-4o-mini-transcribe) permits up to 2,000 output tokens and lists token pricing; the transcription request does not expose a smaller caller-controlled output-token ceiling. The resulting documented worst case exceeds the cap, so the evaluator refuses billable execution before making a request. A separate user decision must revise the case count or cap; an estimate below $0.15 is not sufficient authority to spend.

Bounded audio is transmitted to the transcription endpoint and selected M4 context is transmitted to Responses for the ten full-pipeline cases. OpenAI's current [data-controls table](https://developers.openai.com/api/docs/guides/your-data#default-usage-policies-by-endpoint) reports no application-state or abuse-monitoring retention for audio transcription; ordinary Responses API abuse-monitoring rules can still apply.

## Automated evidence

| Case ID | Gate | Result | Evidence / notes |
|---|---|---|---|
| M6-AUTO-01 | Typed and audio operations cannot overlap; Busy remains distinct | Pass | Shared coordinator and UI tests |
| M6-AUTO-02 | One operation ID/cancellation signal spans every stage | Pass | Coordinator/controller tests |
| M6-AUTO-03 | Cancellation at each boundary prevents later stages/stale updates | Pass | Deterministic boundary and race tests |
| M6-AUTO-04 | Upload accepts only owned, bounded, valid 16 kHz mono PCM WAVs | Pass | Byte-level ownership/RIFF/format/duration/size tests |
| M6-AUTO-05 | Transcript normalization rejects empty/control-only/>4,000 characters | Pass | Transcription validation tests |
| M6-AUTO-06 | Hint is deduplicated and bounded from approved vocabulary/doc titles | Pass | Vocabulary/settings/transcription tests |
| M6-AUTO-07 | WAV is deleted in transcription `finally`, before retrieval | Pass | Pipeline cleanup-order and failure tests |
| M6-AUTO-08 | Exactly one Responses request performs classification and answering | Pass | Mocked pipeline/request-shape tests; `store:false` preserved |
| M6-AUTO-09 | Only selected M4 chunks are sent; citations remain validated | Pass | Cross-reference accepted `milestone-4.md`; context/citation tests remain green |
| M6-AUTO-10 | Stage timings/usage contain no audio, transcript, prompt, or answer | Pass | Timing/usage/redaction tests and offline report scan |

Aggregate non-billable evidence: 196 Vitest tests in 28 files, 29/29 .NET tests, 5/5 Playwright Electron tests, zero high-severity audit findings, and the accepted 50/50 M4 corpus.

## Renderer-visible latency evidence

Release-to-visible-answer acceptance must use the production app's operation-scoped renderer frame acknowledgement. An imported or manually entered standalone timing value is not bound strongly enough to the same audio operation and therefore cannot close the p50/p95 gate by itself. During the user-assisted campaign, transiently verify that the acknowledgement belongs to the active operation, then persist only its case ID and timing. Internal pipeline `total` timing remains diagnostic only.

## Live twenty-case transcription campaign

Human meaning review is transient. Record only the outcome flag; do not copy transcript or answer text into this file.

| Case ID | M5 trial | Structured | Meaning correct | Continued E2E | Transcription ms | Retrieval ms | Generation ms | Internal total ms | Release→visible ms | Temp gone | Tokens / cost | Result |
|---|---:|---|---|---|---:|---:|---:|---:|---:|---|---|---|
| M6-LIVE-01–20 | | | | | | | | | | | | Untested |

Designate exactly ten cases for the complete transcription-to-answer path before starting the campaign. The other ten stop after transcription and deletion verification.

## Acceptance calculations

- Valid structured transcription results: ___/20; required 20/20.
- Correct reviewer-question meaning: ___/20; required at least 18/20.
- Complete E2E cases: ___/10; required 10/10 terminal outcomes.
- Release-to-visible-answer p50: ___ ms; required ≤5,000 ms.
- Release-to-visible-answer p95: ___ ms; required ≤8,000 ms.
- Temporary WAV absent after all outcomes: ___/20; required 20/20.
- Actual estimated spend: $___; required ≤$0.15.
- Failed/blocking case IDs:
- Human sign-off / date:
- Decision: **Blocked before billable execution; $0 spent**
Loading
Loading