Owner's concept, 2026-07-24. Status: approved direction, staged build; nothing started. This phase crosses the project's founding line — the AI moves from advising the doctor to acting on the patient — so the hard rules below are constitutional, not preferences.
A small Auto toggle beside Start consultation. In auto mode, Consultation AI conducts the history-taking conversation with the (synthetic/actor) patient by voice — listening like a good GP, asking questions derived from its own CDS agenda — while the doctor supervises, intervenes at will, and performs the examination. An expressive robot face gives the system a bedside presence.
- The AI asks questions and acknowledges; it NEVER advises, reassures, interprets, breaks news, or hints at diagnosis to the patient. Any drafted utterance that is not a question or a minimal acknowledgement is suppressed in code, not just prompt.
- Urgency pause protocol (owner-refined 2026-07-24): when the urgency alarm fires, auto mode PAUSES — it stops speaking and asking, but keeps listening and transcribing — and the alert is presented to the supervising doctor. Pause, not termination: alarms can be transient (the script-02 dengue boundary case), and the stateless urgency officer re-evaluates fresh on every update, so clarification can clear a false alarm on its own. On acknowledging the alert the doctor chooses: RESUME AUTO (the question agenda is reprioritised toward the alarm's clarifying questions) or TAKE OVER (drop to standard mode). Ratchet: a re-fire of the same urgent action after a resume pauses again and requires a fresh acknowledgement — no automatic ack-resume loops — and an alarm still unresolved at Stop carries into the review page's acknowledge-gated banner exactly as today.
- The doctor always wins: doctor speech (barge-in) cancels the AI's current and queued utterances instantly. Toggling out of auto mode is one tap and immediate.
- Disclosure: the patient/actor is told they are talking to a machine. The robot face (deliberately non-human) reinforces rather than replaces this.
- Everything is logged: every question asked, its CDS rationale, its timing, and every doctor intervention — CDSSnapshot-style, for research and audit.
- English only — now the project-wide v1 position, not a Phase 7 restriction: Consultation AI is English-only for v1 by owner decision 2026-07-25 (PROJECT_PLAN.md §4, Scope decision). Phase 7 inherits it rather than imposing it.
- Synthetic consultations only, as everywhere in this project.
Mimic real GP craft, not interrogation:
- Golden minutes. Open with a single invitation ("Please, tell me what's brought you in") then stay silent for the first 2–3 minutes while the patient talks freely. During this phase the system produces only minimal encouragers ("mm-hm", "I see", "go on") triggered after natural pauses (~1.5–2 s of silence), and NO questions.
- Open-to-closed cone. When the patient's free narrative dries up (sustained silence or explicit hand-back), begin with open questions ("Can you tell me more about the chest pain?") before narrowing to specific/closed questions (onset, radiation, exacerbating factors...) drawn from the CDS questions_to_ask agenda.
- One question at a time, wait for the answer, let the CDS revise the agenda on each answer (the existing revision rules already remove answered questions).
- Examination handover. The system never pretends to examine. When the agenda is exhausted or the doctor intervenes, it hands over explicitly: "Thank you — Dr X will examine you now."
Each of these is measurable (see eval design) — the policy is written to be scored.
HARD REQUIREMENT (added 2026-07-25), alongside the hard rules above and equally non-negotiable.
System-spoken utterances must be excluded from the patient transcript BY CONSTRUCTION — never by prompt instruction and never by post-hoc filtering. The server knows precisely when it is speaking; that knowledge is the mechanism, and it is the only acceptable one.
Why this is a hard rule rather than an implementation detail: 7a introduces audio output into a system whose entire safety guarantee is that the transcript is faithful to the room. If the system's own voice can enter the transcript, the system can put words in the patient's mouth — and every downstream defence inherits the corruption, because the note grounding gate would faithfully cite a fabricated turn. That is a new fabrication vector of exactly the class the transcript-quality gate was built to close (consultation #70: a note faithful to a transcript that was not faithful to the audio).
Echo handling and ASR gating during playback are part of this requirement, not an optimisation. Muting or AEC-gating the recogniser while the system speaks is how the guarantee is kept in the presence of an open microphone; it is not a quality improvement to be deferred to a later pass.
A prompt asking the model to ignore its own utterances does not satisfy this rule. Neither does stripping matched strings from the transcript afterwards. Both are detection; the requirement is prevention. CDS panel questions become tappable. Doctor taps → local TTS speaks the question to the patient → answer flows through the existing ASR/CDS pipeline. No turn-taking AI, no end-of-turn detection, no barge-in problem — the doctor IS the turn-taker. Builds and battle-tests: TTS integration (Piper or equivalent, fully local), audio output path alongside capture (echo handling: mute/AEC-gate the ASR while the system speaks, and exclude system utterances from the patient transcript by construction — the system knows when it is speaking), spoken-question logging, and patient reaction to a machine voice.
https://github.com/smithandrewjohn/kindalive — MIT (attribution in NOTICE), Python, drives emotion from local Ollama-compatible models, renders a retro LED dot-matrix face via 12 FACS-based muscles from a simulated-neurochemistry state that evolves smoothly.
Build brief:
PHASE_7B_KINDALIVE.md(owner's own conclusions from reviewing upstream, 2026-07-25 — to be followed as written). It covers the upstream assessment, the architecture as understood, and six build decisions: vendor the zero-dependency core at a pinned commit; skip the upstream LLM interpreter and inject impulses deterministically; a damped[clinical]personality preset plus hard caps in our own mapping layer;face3d.jsover the existing WebSocket with no new transport; bedside-device styling; and an expected calibration pass. It also records the CARE-measure study design behind the face-off requirement below.
- Embed the web face renderer in the live page (renderer-agnostic per upstream; a canvas/web component beside or replacing the idle CDS space in auto mode).
- Drive it cheaply: add one optional field to the existing CDS assessment JSON (e.g. patient_affect_hint or situation summary line) and feed that to kindalive's impulse input — zero additional model calls on the 4090. Direct small-model mode via Ollama is the fallback if the hint quality disappoints.
- The face listens in 7a already (reacts while the patient talks, blinks, attends) even though the doctor still taps the questions — presence before autonomy.
- Face is a runtime toggle, independent of auto mode: OFF must remain a first-class state (it is the control arm of the face study, and some demos will want no face).
The system taps its own buttons, governed by the behaviour policy above:
- End-of-turn detection: VAD silence threshold + a lightweight "has the patient finished the thought?" check; err toward waiting (a slow system is polite; an interrupting one is clinically wrong and fails the eval).
- Latency budget: encourager < 1 s; question (MedGemma selection + TTS synthesis) ≤ ~2 s from end-of-turn. Pre-synthesise the top agenda question during the patient's turn to hide latency.
- Barge-in: doctor VAD-detected speech cancels playback immediately (rule 3). Doctor utterances route to the transcript as doctor turns as today.
- Phase state machine in code (invitation → golden-minutes → open → closed → handover), with the LLM choosing content WITHIN the current phase, never the phase itself.
Reuse the existing method: mock scripts already carry expected-clinical-content marking schemes; the patient side can be played by an actor from the script while auto mode conducts the interview.
Primary metrics, all computable from logs:
- Elicitation coverage: fraction of the script's marking-scheme content the auto interview surfaced (the auto-mode analogue of note coverage).
- Talk-time ratio (patient speaking time / total) and time-to-first-question — the golden-minutes compliance pair.
- Open:closed ratio over consultation thirds (should fall over time — the cone).
- Interruption count (system spoke while patient mid-turn) — target ~0.
- Urgency latency in auto mode: script-14-style buried red flags must still fire the alarm and halt questioning.
- Face study (7b): face on/off as randomised arm across matched script runs — effect on actor disclosure length and content. (With real participants one day, this is an ethics-approved study; with actors it is a rehearsal of the method.)
Sinhala voice interaction; AI examination of any kind; AI communication of findings/diagnosis/plan to the patient; unsupervised operation (no doctor present); real patients (as everywhere).
Phase 7 starts only after the current docket obligations are stable (Phase 5 adjudication/decision, Phase 0 recordings, Docker/demo packaging).
Gate update 2026-07-25. The Phase 5 precondition is satisfied — by closure, not by deferral: the step 6 adjudication completed 2026-07-25 and Phase 5 was closed the same day with a negative result, Sinhala being out of scope for v1 rather than postponed (evals/2026-07-12_sinhala_asr_recordings_eval.md § Step 6; decision header in evals/2026-07-17_finetune_plan.md; PROJECT_PLAN.md §4). The recordings precondition means 05_epigastric_pain_en only — 01_chest_pain_si is not being recorded for v1, since it needs a Sinhala-speaking second reader and only feeds an out-of-scope arm.
The remaining gate is three items:
- Record
05_epigastric_pain_en. - Docker Compose packaging plus the two-role demo script.
- The finalisation transcript-quality gate (docket build item — flag the review as unreliable on language mismatch or low average confidence instead of presenting a normal draft). Declaring the project English-only makes this load-bearing: a non-English speaker at the publicly reachable demo would otherwise receive a fluent fabricated note, which is exactly what happened with consultation #70.
Gate amendment 2026-07-25 — PHASE 7 IS OPEN. Owner decision: Phase 7 opens now, with two of the three gate items deliberately carried rather than met. This is a decision to proceed knowingly, not a decision that the items stopped mattering.
| Gate item | State | Why it is carried |
|---|---|---|
1. 05_epigastric_pain_en |
carried | Cannot be recorded for roughly 8 days (reader availability). |
| 2. Docker Compose + demo script | carried | Spec exists but awaits the owner's decisions. |
| 3. Transcript-quality gate | MET, for its refuse path | The refuse tier is built and live: S2 below 0.60 and S4 above 20 s each refuse independently, no note is drafted, status unreliable_transcript. See TRANSCRIPT_QUALITY_GATE_SPEC.md §11. |
Outstanding work carried alongside, recorded so none of it is lost — see HANDOVER's "Phase 7 opened" section for the full list: the gate's flag tier is specified but unbuilt; the S1 multi-window and S3 within-segment redesigns are pending, and must share one measurement function with the calibration harness; the schema-drift recommendation is agreed in principle but not built; Phase 2 real-audio validation still needs 05_epigastric_pain_en; and consultations #78 and #162 await the owner's review.
7a+7b make a strong demo milestone on their own and are the recommended first commitment; 7c is committed separately after 7a/7b review.