Repository navigation
usertest: voice-moderated sessions with a Claude moderator - #185
Merged
Merged
Conversation
Testers get one invite link to /usertest/ on the test instance. The page asks consent for what is recorded and where it goes, sets the test catalog, opens a fresh drive in a second window, records the shared screen and the microphone, listens with the browser's speech recognition and speaks the moderator's answers with speech synthesis. The moderator (moderator/server.mjs, @anthropic-ai/sdk, claude-opus-5 at effort low, server-side refusal fallback) follows moderator/script.md: short spoken questions, mostly listening ([WAIT]), no help unless the tester is stuck and asks. Every turn also carries the collector's error, warning, feedback and sync lines since the previous one. Per session it keeps the transcript and the recording. The endpoints need the invite code, and sessions and turns are capped, because they spend API money. Checked on the droplet: about 3 s per turn, [WAIT] while the tester thinks aloud, a question instead of a hint when asked which button to press, and prompt-cache reads from the second turn on. Not yet run by a real tester. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…#184) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
michielbdejong
force-pushed
the
claude/usertest-interview
branch
from
September 28, 2026 12:26
22e5700 to
68b967f
Compare
From the first real session (2026-09-28, Michiel): - Dutch was chosen at the start, but the moderator switched language on its own while the voice stayed; and before Chrome's voice list loaded, the first line had a different voice. Now English only (page, recognizer, script) and one voice picked once the list has loaded. - Questions did not hand over: Chrome's recognizer adds no "?", so every question waited the full 9 s. Questions are now recognized by their words (2 s), other speech hands over after 5 s, and an "Ask the moderator" button hands over at once. - The moderator could not see the screen. Each turn now carries a 1280-pixel JPEG of the shared screen (the history keeps only the text), saved per turn as screen-NNN.jpg; the consent text says so. Checked on the droplet: a turn with a JPEG is accepted and answered. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… 'wait' Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds voice-moderated user-testing sessions to the test instance from #182, so testers can do a session on their own laptop without anyone from the team on a call.
What a tester gets
One invite link:
https://plugins.<base-domain>/usertest/?code=…, Chrome or Edge, headphones. The page (usertest/page/):SpeechRecognitionand hands a turn to the moderator after a pause (3 s after a question, 9 s otherwise), after 60 s of silence, or when the collector logs a new error; then speaks the answer withspeechSynthesis. Recognition pauses while the moderator speaks.The moderator
usertest/moderator/server.mjs, with@anthropic-ai/sdk0.128.0:claude-opus-5,output_config.effort: "low"for short pauses, server-side refusal fallback (fallbacks: "default"), top-level prompt caching;usertest/moderator/script.md: spoken one-or-two-sentence turns, mostly listening ([WAIT]), no hints unless the tester is stuck and asks, neutral about the app, four tasks (import calendar events, change one and send it back, explore other apps, wrap-up), then[END];/var/lib/usertest-sessions/<id>/:meta.json,transcript.jsonl(both sides, the log lines each turn saw, token usage) andrecording.webm.The endpoints spend API money and are public, so they need the invite code (
x-usertest-code, from/etc/usertest-moderator.env). There are at most 120 turns per session and 20 sessions per UTC day. The API key stays in/etc/anthropic.env, read only by the moderator container.Checked
On the droplet, over the real API:
[WAIT](nothing spoken);cache_read_input_tokens989 to 1088 from the second turn on;oxlintandoxfmtover the new files are clean.node --checkpasses for both scripts.Not verified yet: a full session with a real tester, screen and microphone permissions in the page, speech recognition quality in Dutch, and Edge.
🤖 Generated with Claude Code