diff --git a/README.md b/README.md index ce88cc7..8eb75b3 100644 --- a/README.md +++ b/README.md @@ -76,6 +76,11 @@ py -m darwin_v50.predictive_planning_evaluation py -m darwin_v50.learned_context_evaluation ``` +The isolated conversational development surface requires an explicitly +selected backend and model. It has no automatic fallback and keeps its bounded +transcript in memory only. See +[`docs/v50/CONVERSATIONAL_DEVELOPMENT_GUIDE.md`](docs/v50/CONVERSATIONAL_DEVELOPMENT_GUIDE.md). + Do not rerun a final seed set to tune a failed experiment. Seed contamination is part of the research record. diff --git a/docs/v50/CONVERSATIONAL_DEVELOPMENT_GUIDE.md b/docs/v50/CONVERSATIONAL_DEVELOPMENT_GUIDE.md new file mode 100644 index 0000000..3736feb --- /dev/null +++ b/docs/v50/CONVERSATIONAL_DEVELOPMENT_GUIDE.md @@ -0,0 +1,84 @@ +# Conversational development guide + +This guide starts the isolated E044 terminal session. It does not modify or +replace the E043 desktop runtime. + +## What this surface does + +Each successful turn makes two model calls through `DarwinLanguageGateway`: + +```text +free text -> UNDERSTAND -> unverified observation + -> deterministic conversation policy + -> EXPRESS -> new natural-language reply +``` + +The transcript is kept only in process memory, is bounded to 60 messages, and +is cleared when the command exits. The runtime has no event-store, goal, +executor, consent, capability, RZS, sigma, identity, or autobiographical-memory +handle. + +This is still a development adapter around a language model. The current +conversation policy preserves the authority boundary, but it is not a general +cognitive deliberation mechanism. + +## OpenAI configuration + +Install the project, then set all three values explicitly in the same terminal: + +```powershell +py -m pip install -e . +$env:DARWIN_LLM_BACKEND = "openai" +$env:DARWIN_LLM_MODEL = "" +$env:OPENAI_API_KEY = "" +darwin-conversation-dev +``` + +You can also run: + +```powershell +py -m darwin_v50.conversation.cli +``` + +There is no default model. The command checks the exact configured identifier +through `GET /v1/models/{model}` before admitting the session. A missing key, +missing model, rejected probe, malformed result, or provider error never causes +a switch to a local model. + +Every model request uses the Responses API with `store: false`, strict +Structured Outputs, no `previous_response_id`, and no tools. The transcript is +manually resent as bounded input. `store: false` means the runtime does not rely +on provider-side application state; it does not override the provider's +separate abuse-monitoring retention or account data-control policy. + +Relevant official references: + +- +- +- + +## No backend and local development + +The default is: + +```powershell +$env:DARWIN_LLM_BACKEND = "none" +``` + +In that state the conversational command reports `backend_not_requested` and +exits. It does not generate a canned conversation. + +`DARWIN_LLM_BACKEND=local` is an explicit code-level integration seam. E044 +does not install, discover, start, or assume compatibility with Ollama, LM +Studio, or any other local server. A local adapter must be supplied explicitly, +declare the exact `DARWIN_LLM_MODEL` it serves, implement the frozen gateway +contract, and clear its pending context on close. No local adapter has been +validated by E044 yet. + +## Ending a session + +Use `/exit`, `/quit`, Ctrl+C, or close the terminal. The runtime clears its +temporary transcript in all normal exit paths. It writes no conversation log. + +Do not paste secrets into the conversation. Model inputs are still sent to the +explicitly selected provider, subject to that provider's data controls. diff --git a/docs/v50/EXPERIMENT_044_CONVERSATIONAL_DEVELOPMENT_RUNTIME.md b/docs/v50/EXPERIMENT_044_CONVERSATIONAL_DEVELOPMENT_RUNTIME.md new file mode 100644 index 0000000..1ddcc76 --- /dev/null +++ b/docs/v50/EXPERIMENT_044_CONVERSATIONAL_DEVELOPMENT_RUNTIME.md @@ -0,0 +1,250 @@ +# Experiment 044 - conversational development runtime + +Status: automated development admission passed after pre-registration. The live +provider and 20-30 minute usability probes are unexecuted. No language-quality +or scientific capability claim is registered. + +## Purpose + +Build a deliberately separate development surface for open-ended conversation +without changing the E043 persistent desktop candidate or granting a language +model authority over Darwin's persistent state. + +This is a usability and boundary experiment. It is not a calibration +experiment, a consciousness test, evidence of personhood, or evidence that +Darwin is more than a language model. A fluent result can establish only that a +configured model can operate behind the existing language boundary while the +development runtime preserves its stated restrictions. + +## Frozen starting point + +- Base commit: `2602c57f21dc930b6fbe4426610469b39357119f` +- Development branch: `codex/conversational-darwin-dev` +- E043 runtime SHA-256: + `fd0d8aaf2bd1011addea581eaef172ce940157164dc013681d5326476d49f7e8` +- E043 protocol SHA-256: + `beea70ce6bcfd177a18d36feca016f5f93ed0bed7743fc73c31152cc19eaefad` + +The files covered by the two digests above must remain byte-identical in this +experiment: + +- `src/darwin_v50/desktop_runtime.py` +- `docs/v50/EXPERIMENT_043_PERSISTENT_DESKTOP_RUNTIME.md` + +E043 remains PURE and keeps its own incomplete 14-day admission campaign. +E044 cannot complete, revise, or add evidence to E043. + +## Architecture under test + +```text +explicit backend configuration + | + v +OpenAI backend / explicit local seam / no backend + | + v +DarwinLanguageGateway + UNDERSTAND -> candidate observation + | + v +deterministic conversation policy + candidate remains unverified + no persistent mutation path + | + v +DarwinLanguageGateway + EXPRESS -> natural-language rendering +``` + +The deterministic conversation policy is a narrow core-side policy for this +development surface. It does not claim to be the full Darwin cognitive kernel. +The present v50 kernel has no general conversational deliberation mechanism; +pretending otherwise would overstate the implementation. + +## Pre-registered invariants + +### Backend selection + +1. `DARWIN_LLM_BACKEND` is the only backend selector. +2. The default is `none`. +3. Valid selections are `none`, `openai`, and `local`. +4. `openai` never falls back to `local`, and `local` never falls back to + `openai`. +5. Missing configuration, a failed model probe, a network error, a refusal, or + a malformed response makes the requested backend unavailable for that + operation. It does not select another model or provider. +6. An unavailable runtime does not manufacture a conversational reply through + a hidden fallback. +7. The local seam is usable only when a local backend is supplied explicitly. + E044 will not discover, install, launch, or guess a local server. + +### Model and transport + +1. `DARWIN_LLM_MODEL` is required for every model-backed configuration. +2. No API model identifier is embedded as a default in source code. +3. OpenAI configuration additionally requires `OPENAI_API_KEY`. +4. The OpenAI adapter uses the Responses API. +5. Every Responses request contains `store: false`. +6. The adapter does not send `previous_response_id`, tools, function calls, + web-search configuration, or computer-use configuration. +7. The model identifier is checked through the Models API before an interactive + OpenAI session is admitted. +8. The adapter uses strict Structured Outputs for both operations and rejects + refusals, incomplete output, missing output text, invalid JSON, unknown + fields, and authority fields. +9. API keys must not appear in snapshots, errors, prompts, test fixtures, or + logs. + +`store: false` prevents E044 from relying on provider-side application state. +It is not documented or represented as a guarantee of zero provider retention; +provider abuse-monitoring and account data-control policies remain separate. + +### Conversation flow + +One successful user turn performs exactly this sequence: + +1. create a bounded `UnderstandingRequest` from the current text and temporary + session transcript; +2. invoke the configured model for `UNDERSTAND`; +3. parse the result as a candidate `LanguageObservation`; +4. pass that candidate to the deterministic conversation policy; +5. create an `ExpressionPlan` that preserves the candidate status and the + no-authority boundary; +6. invoke the same explicitly selected model for `EXPRESS`; +7. append the user and Darwin text to the in-memory transcript only after both + operations succeed. + +An `EXPRESS` request without a successful immediately preceding `UNDERSTAND` +request is invalid. A partial or failed turn is not appended to the transcript. + +### Temporary context + +1. Context is held in memory for one process session only. +2. The transcript is bounded by count and per-message length. +3. The runtime manually resends the bounded transcript on each turn and does + not depend on provider-side conversation state. +4. Closing the runtime clears the transcript and any pending backend context. +5. E044 does not write conversation text or extracted candidates to SQLite, + files, autobiographical memory, a vector store, or a provider conversation. +6. No automatic memory-candidate feature is in scope. + +### Authority + +The language model receives no callable path that can: + +- write persistent or autobiographical memory; +- create, start, cancel, or modify a goal; +- change RZS, sigma, motivation, preference, identity, or world-model state; +- dispatch or execute an action; +- issue consent or capability grants; +- write directly to the Darwin event store. + +The model output remains subject to the existing exact response schemas and +forbidden-authority-field check in `DarwinLanguageGateway`. + +## Acceptance checks + +The implementation can be admitted as E044 development infrastructure only if: + +1. the two frozen E043 file digests remain unchanged; +2. automated tests prove all backend-selection branches and the absence of + cross-provider fallback; +3. captured OpenAI requests prove that the configured model is used and that + every request has `store: false`; +4. captured requests contain no tools or provider-side continuation identifier; +5. a model-backed successful turn makes one `UNDERSTAND` call followed by one + `EXPRESS` call; +6. strict response parsing rejects malformed and authority-bearing results; +7. failed and partial turns do not enter session context; +8. closing a session erases its temporary transcript; +9. the full existing automated suite still passes; and +10. no live result is reported unless a real configured provider was actually + called and the raw run metadata was recorded without secrets. + +## Live usability gate + +The intended later live probe is a 20-30 minute conversation in Brazilian +Portuguese that includes unplanned subjects, at least two topic changes, and at +least one return to an earlier topic. No response sentence may come from a +registered response library. + +The probe must separately record: + +- successful and failed turns; +- end-to-end latency per operation; +- whether topic returns were handled correctly; +- human-noted contradictions or fabricated memories; +- model and backend identifiers; +- confirmation that persistent-authority mutation counts remained zero; and +- confirmation that the transcript disappeared when the session closed. + +Automated mocks cannot pass this live usability gate. If no API credential or +explicit local backend is available, the correct result is `UNEXECUTED`, not a +simulated success. + +## Claims explicitly excluded + +E044 cannot establish: + +- calibrated natural-language understanding; +- semantic fidelity of generated replies; +- long-term memory or autobiographical continuity; +- autonomous goal formation; +- general reasoning by the Darwin kernel; +- consciousness, sentience, personhood, AGI, or similarity to Diana beyond a + superficial conversational impression; or +- that Darwin is more than an LLM-centered conversational system. + +The last exclusion matters. Until durable cognition, learning, self-model, +goal, and evidence mechanisms causally shape conversation under independent +tests, a successful E044 result is still best described as an LLM conversation +adapter behind a strict authority boundary. + +## Observed implementation result + +The protocol above was committed as `61d6f97` before any E044 implementation +or test was added. + +The admitted implementation adds: + +- explicit environment parsing with `none` as the default backend; +- an OpenAI Responses adapter built on an injectable, bounded JSON transport; +- mandatory exact-model probing through the Models API; +- strict `UNDERSTAND` and `EXPRESS` JSON schemas; +- a deterministic policy that marks the interpretation as an unverified + candidate and creates no persistent mutation handle; +- a bounded 60-message in-memory session; +- an explicit local-backend seam with model matching and mandatory ephemeral + context cleanup; and +- a terminal command that starts only when invoked. + +Automated observations on the local Windows machine: + +- 24 of 24 focused E044 tests passed; +- the full suite discovered 472 tests; +- 471 tests executed and passed; +- one pre-existing workspace-executor test was skipped because Windows symlink + creation was unavailable; +- zero tests failed; +- both frozen E043 SHA-256 digests remained exact; +- captured OpenAI request fixtures used the configured model, `store: false`, + strict Structured Outputs, and no tools or `previous_response_id`; and +- the no-backend terminal check reported `backend_not_requested` and made no + conversational reply. + +These are hermetic implementation checks. The provider responses were test +fixtures, not OpenAI responses. No `OPENAI_API_KEY` was configured on the test +machine, no explicit local backend was installed or validated, and no billable +API request was made. Therefore: + +```text +live OpenAI model probe UNEXECUTED +live UNDERSTAND call UNEXECUTED +live EXPRESS call UNEXECUTED +20-30 minute usability probe UNEXECUTED +language quality UNKNOWN +topic-return quality UNKNOWN +``` + +Automated development admission does not pass the live usability gate and does +not alter the independent E041/E042 human-annotation block. diff --git a/docs/v50/README.md b/docs/v50/README.md index 5b53d88..d7e7bc9 100644 --- a/docs/v50/README.md +++ b/docs/v50/README.md @@ -267,6 +267,19 @@ real-machine campaign has not started. No cognitive-continuity claim is registered. The language-calibration line remains independently blocked on human annotation. +## Conversational development + +[Experiment 044](EXPERIMENT_044_CONVERSATIONAL_DEVELOPMENT_RUNTIME.md) +pre-registers a separate, non-persistent conversation surface. A backend and +model must be selected explicitly; OpenAI and local modes never replace one +another silently. Successful turns require model-backed `UNDERSTAND` and +`EXPRESS` operations, while the deterministic policy keeps interpretations +unverified and exposes no persistent-memory, goal, RZS, sigma, identity, or +action mutation path. This is development infrastructure, not language +calibration or evidence that Darwin is more than an LLM-centered system. Setup +is documented in the +[conversational development guide](CONVERSATIONAL_DEVELOPMENT_GUIDE.md). + Local passes are E1 evidence produced by this repository's own evaluator. They are useful engineering results, but they are not independent replication. diff --git a/docs/v50/results/E044_CI_PORTABILITY_INCIDENT_01.md b/docs/v50/results/E044_CI_PORTABILITY_INCIDENT_01.md new file mode 100644 index 0000000..c6c1849 --- /dev/null +++ b/docs/v50/results/E044_CI_PORTABILITY_INCIDENT_01.md @@ -0,0 +1,81 @@ +# E044 CI portability incident 01 + +Status: first failing run recorded. The failure is classified as a portability +defect in the freeze measurement, not a mutation of E043. This document records +the first run only; it does not claim that a later corrective run passed. + +## Run + +- Pull request: +- Workflow run: +- Head commit: `ca04a2c23615e3713700a56269126ab264cacab2` +- Event: `pull_request` +- Result: failed +- Total tests: 472 +- Passed tests: 470 +- Failed tests: 2 +- Failed step: `Run test suite` + +The only failures were: + +- `test_e043_runtime_remains_byte_identical` +- `test_e043_protocol_remains_byte_identical` + +## Observed hashes + +The first test version hashed raw working-tree bytes. The local checkout used +LF, while the Windows Actions checkout materialized the same lines as CRLF. + +| Frozen file | Registered LF SHA-256 | CI raw CRLF SHA-256 | +| --- | --- | --- | +| `src/darwin_v50/desktop_runtime.py` | `fd0d8aaf2bd1011addea581eaef172ce940157164dc013681d5326476d49f7e8` | `b3cfca4be41a59a6a80fe0ea6c16855b6c73846abf3f9081bf88470a5f22bc7f` | +| `docs/v50/EXPERIMENT_043_PERSISTENT_DESKTOP_RUNTIME.md` | `beea70ce6bcfd177a18d36feca016f5f93ed0bed7743fc73c31152cc19eaefad` | `86cf2b12a05ead4c124bd9b48df50765915067135a77ec78fd476c1be0489148` | + +Converting the frozen LF content to CRLF reproduces both CI hashes exactly. +Normalizing only CRLF back to LF reproduces the originally registered hashes. +No expected SHA-256 value was recalibrated from the E044 head. + +## Independent Git identity check + +The Git blob identifiers are identical at the frozen base and the failing head: + +| Frozen file | Base blob | Head blob | +| --- | --- | --- | +| `src/darwin_v50/desktop_runtime.py` | `01687c8e57aa1a867b18a66df9442b8745c08566` | `01687c8e57aa1a867b18a66df9442b8745c08566` | +| `docs/v50/EXPERIMENT_043_PERSISTENT_DESKTOP_RUNTIME.md` | `5becbc0177c7fc1fc1cfe0e2903e7a8323a7bd7d` | `5becbc0177c7fc1fc1cfe0e2903e7a8323a7bd7d` | + +The frozen base is +`2602c57f21dc930b6fbe4426610469b39357119f`. A direct Git diff from that base +to the failing head is empty for both files. + +## Classification + +The first test conflated two properties: + +1. semantic and Git-object identity of the frozen files; and +2. the checkout's platform-specific line-ending materialization. + +E043 satisfied the first property. The test failed on the second. The failure +therefore does not show a change to the E043 runtime or protocol. + +## Authorized correction boundary + +The corrective test may: + +- replace only `CRLF` byte pairs with `LF` before SHA-256 calculation; +- retain the original registered SHA-256 values; +- reject any remaining lone carriage return; +- require the frozen-base blob identifier to equal its registered value; and +- require the current `HEAD` blob identifier to equal the frozen-base blob. + +It may not: + +- modify either frozen E043 file; +- derive an expected digest or blob from the E044 head; +- normalize any content other than CRLF/LF line endings; +- remove the digest check; +- remove the Git blob identity check; or +- weaken another E044 admission criterion. + +The first red run remains part of the record even if the bounded correction +later passes. diff --git a/pyproject.toml b/pyproject.toml index 8282b90..16f64f2 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -29,6 +29,7 @@ darwin-information-directed-audit = "darwin_v50.information_directed_diagnostics darwin-transfer-benchmark = "darwin_v50.cross_world_transfer_evaluation:main" darwin-language-conformance = "darwin_v50.language_evaluation:main" darwin-language-annotation = "darwin_v50.language_annotation_evaluation:main" +darwin-conversation-dev = "darwin_v50.conversation.cli:main" [tool.setuptools] package-dir = {"" = "src"} diff --git a/src/darwin_v50/conversation/__init__.py b/src/darwin_v50/conversation/__init__.py new file mode 100644 index 0000000..115d4d1 --- /dev/null +++ b/src/darwin_v50/conversation/__init__.py @@ -0,0 +1,49 @@ +"""Darwin conversational development surface.""" + +from .config import ( + DEFAULT_LOCALE, + DEFAULT_OPENAI_API_BASE, + ConversationBackendKind, + ConversationSettings, +) +from .openai_responses import ( + EXPRESSION_SCHEMA, + UNDERSTANDING_SCHEMA, + JSONTransport, + OpenAIResponsesBackend, + OpenAITransportError, + UrllibJSONTransport, +) +from .runtime import ( + AuthorityMutationCounts, + ConversationAvailability, + ConversationPolicy, + ConversationRuntime, + ConversationRuntimeError, + ConversationSnapshot, + ConversationTurnResult, + ConversationUnavailableError, + ExplicitLocalBackend, +) + +__all__ = [ + "AuthorityMutationCounts", + "ConversationAvailability", + "ConversationBackendKind", + "ConversationPolicy", + "ConversationRuntime", + "ConversationRuntimeError", + "ConversationSettings", + "ConversationSnapshot", + "ConversationTurnResult", + "ConversationUnavailableError", + "DEFAULT_LOCALE", + "DEFAULT_OPENAI_API_BASE", + "EXPRESSION_SCHEMA", + "ExplicitLocalBackend", + "JSONTransport", + "OpenAIResponsesBackend", + "OpenAITransportError", + "UNDERSTANDING_SCHEMA", + "UrllibJSONTransport", +] diff --git a/src/darwin_v50/conversation/cli.py b/src/darwin_v50/conversation/cli.py new file mode 100644 index 0000000..aebaa60 --- /dev/null +++ b/src/darwin_v50/conversation/cli.py @@ -0,0 +1,74 @@ +"""Explicitly launched terminal surface for conversational development.""" + +from __future__ import annotations + +import sys +from typing import TextIO + +from ..language import LanguageBoundaryError +from ..models import ValidationError +from .config import ConversationSettings +from .runtime import ConversationAvailability, ConversationRuntime + + +def _write(stream: TextIO, message: str) -> None: + stream.write(message + "\n") + stream.flush() + + +def main() -> int: + try: + settings = ConversationSettings.from_environment() + except ValidationError as exc: + _write(sys.stderr, f"Darwin configuration error: {exc}") + return 2 + + runtime = ConversationRuntime.create(settings) + snapshot = runtime.snapshot() + if snapshot.availability is not ConversationAvailability.AVAILABLE: + _write( + sys.stderr, + "Darwin conversational backend is unavailable: " + f"{snapshot.unavailable_reason}.", + ) + _write( + sys.stderr, + "Select a backend explicitly. For OpenAI, set " + "DARWIN_LLM_BACKEND=openai, DARWIN_LLM_MODEL, and OPENAI_API_KEY.", + ) + runtime.close() + return 2 + + _write( + sys.stdout, + f"Darwin conversational dev is active via {snapshot.language_source}.", + ) + _write( + sys.stdout, + "This session is temporary. Type /exit to erase it and leave.", + ) + try: + while True: + try: + text = input("You: ").strip() + except EOFError: + break + if text.lower() in {"/exit", "/quit"}: + break + if not text: + continue + try: + result = runtime.turn(text) + except LanguageBoundaryError as exc: + _write(sys.stderr, f"Darwin turn failed closed: {exc}") + continue + _write(sys.stdout, f"Darwin: {result.expression.text}") + except KeyboardInterrupt: + _write(sys.stdout, "") + finally: + runtime.close() + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/darwin_v50/conversation/config.py b/src/darwin_v50/conversation/config.py new file mode 100644 index 0000000..1be6d0b --- /dev/null +++ b/src/darwin_v50/conversation/config.py @@ -0,0 +1,97 @@ +"""Explicit configuration for the conversational development surface.""" + +from __future__ import annotations + +from dataclasses import dataclass, field +from enum import StrEnum +import os +from typing import Mapping + +from ..models import ValidationError, require_text + + +DEFAULT_OPENAI_API_BASE = "https://api.openai.com/v1" +DEFAULT_LOCALE = "pt-BR" + + +class ConversationBackendKind(StrEnum): + NONE = "none" + OPENAI = "openai" + LOCAL = "local" + + +def _optional_environment_text( + environment: Mapping[str, str], + name: str, +) -> str | None: + value = environment.get(name) + if value is None: + return None + normalized = value.strip() + return normalized or None + + +@dataclass(frozen=True, slots=True) +class ConversationSettings: + """Configuration values without any implicit provider or model choice.""" + + backend: ConversationBackendKind + model: str | None + api_key: str | None = field(default=None, repr=False) + locale: str = DEFAULT_LOCALE + request_timeout_seconds: float = 45.0 + + def __post_init__(self) -> None: + if not isinstance(self.backend, ConversationBackendKind): + raise ValidationError("conversation backend kind is invalid") + if self.model is not None: + require_text(self.model, "conversation model") + if len(self.model) > 200: + raise ValidationError("conversation model exceeds 200 characters") + if self.api_key is not None: + require_text(self.api_key, "OpenAI API key") + require_text(self.locale, "conversation locale") + if len(self.locale) > 32: + raise ValidationError("conversation locale exceeds 32 characters") + if isinstance(self.request_timeout_seconds, bool) or not isinstance( + self.request_timeout_seconds, + (int, float), + ): + raise ValidationError("request timeout must be a number") + timeout = float(self.request_timeout_seconds) + if not 1.0 <= timeout <= 120.0: + raise ValidationError("request timeout must be from 1 to 120 seconds") + object.__setattr__(self, "request_timeout_seconds", timeout) + + @classmethod + def from_environment( + cls, + environment: Mapping[str, str] | None = None, + ) -> "ConversationSettings": + source = os.environ if environment is None else environment + raw_backend = source.get("DARWIN_LLM_BACKEND", "none").strip().lower() + try: + backend = ConversationBackendKind(raw_backend) + except ValueError as exc: + raise ValidationError( + "DARWIN_LLM_BACKEND must be none, openai, or local" + ) from exc + + raw_timeout = _optional_environment_text( + source, + "DARWIN_LLM_TIMEOUT_SECONDS", + ) + try: + timeout = 45.0 if raw_timeout is None else float(raw_timeout) + except ValueError as exc: + raise ValidationError( + "DARWIN_LLM_TIMEOUT_SECONDS must be numeric" + ) from exc + + return cls( + backend=backend, + model=_optional_environment_text(source, "DARWIN_LLM_MODEL"), + api_key=_optional_environment_text(source, "OPENAI_API_KEY"), + locale=source.get("DARWIN_CONVERSATION_LOCALE", DEFAULT_LOCALE).strip(), + request_timeout_seconds=timeout, + ) diff --git a/src/darwin_v50/conversation/openai_responses.py b/src/darwin_v50/conversation/openai_responses.py new file mode 100644 index 0000000..511888e --- /dev/null +++ b/src/darwin_v50/conversation/openai_responses.py @@ -0,0 +1,363 @@ +"""OpenAI Responses adapter for Darwin's provider-neutral language gateway.""" + +from __future__ import annotations + +from copy import deepcopy +import json +from types import MappingProxyType +from typing import Any, Mapping, Protocol +from urllib.error import HTTPError, URLError +from urllib.parse import quote +from urllib.request import Request, urlopen + +from ..language import ( + LanguageBackendError, + LanguageModelRequest, + LanguageOperation, +) +from ..models import ValidationError, require_text +from .config import DEFAULT_OPENAI_API_BASE + + +MAX_HTTP_RESPONSE_BYTES = 2_000_000 + + +class JSONTransport(Protocol): + """Injectable JSON transport; tests never need a network connection.""" + + def request_json( + self, + *, + method: str, + url: str, + headers: Mapping[str, str], + body: Mapping[str, object] | None, + timeout_seconds: float, + ) -> Mapping[str, Any]: + """Perform one JSON request and return a JSON object.""" + + +class OpenAITransportError(LanguageBackendError): + """A sanitized provider transport error that never includes credentials.""" + + +class UrllibJSONTransport: + """Small standard-library HTTPS transport with bounded response reads.""" + + def request_json( + self, + *, + method: str, + url: str, + headers: Mapping[str, str], + body: Mapping[str, object] | None, + timeout_seconds: float, + ) -> Mapping[str, Any]: + data = None + if body is not None: + data = json.dumps( + body, + ensure_ascii=False, + separators=(",", ":"), + ).encode("utf-8") + request = Request( + url=url, + data=data, + headers=dict(headers), + method=method, + ) + try: + with urlopen(request, timeout=timeout_seconds) as response: + raw = response.read(MAX_HTTP_RESPONSE_BYTES + 1) + except HTTPError as exc: + raise OpenAITransportError(f"openai_http_status_{exc.code}") from exc + except (URLError, TimeoutError, OSError) as exc: + raise OpenAITransportError("openai_transport_unavailable") from exc + if len(raw) > MAX_HTTP_RESPONSE_BYTES: + raise OpenAITransportError("openai_response_too_large") + try: + parsed = json.loads(raw.decode("utf-8")) + except (UnicodeDecodeError, json.JSONDecodeError) as exc: + raise OpenAITransportError("openai_response_not_json") from exc + if not isinstance(parsed, dict): + raise OpenAITransportError("openai_response_not_object") + return parsed + + +UNDERSTANDING_SCHEMA: Mapping[str, object] = MappingProxyType( + { + "type": "object", + "properties": { + "intent": {"type": "string"}, + "entities": { + "type": "array", + "items": { + "type": "object", + "properties": { + "kind": {"type": "string"}, + "value": {"type": "string"}, + }, + "required": ["kind", "value"], + "additionalProperties": False, + }, + }, + "reported_signals": { + "type": "array", + "items": { + "type": "object", + "properties": { + "name": {"type": "string"}, + "value": { + "type": "number", + "minimum": 0, + "maximum": 1, + }, + }, + "required": ["name", "value"], + "additionalProperties": False, + }, + }, + "temporal_reference": { + "anyOf": [{"type": "string"}, {"type": "null"}] + }, + "explicit_preference": { + "anyOf": [{"type": "string"}, {"type": "null"}] + }, + "confidence": {"type": "number", "minimum": 0, "maximum": 1}, + }, + "required": [ + "intent", + "entities", + "reported_signals", + "temporal_reference", + "explicit_preference", + "confidence", + ], + "additionalProperties": False, + } +) + + +EXPRESSION_SCHEMA: Mapping[str, object] = MappingProxyType( + { + "type": "object", + "properties": { + "text": {"type": "string"}, + "acknowledged_fact_ids": { + "type": "array", + "items": {"type": "string"}, + }, + }, + "required": ["text", "acknowledged_fact_ids"], + "additionalProperties": False, + } +) + + +_UNDERSTAND_INSTRUCTIONS = """You are a replaceable language parser for Darwin. +Return only the requested structured result. Interpret the current user text in +light of the bounded session transcript. Every result is an unverified language +candidate. Never issue commands, claim that state changed, or request authority +over memory, goals, identity, motivation, RZS, sigma, actions, or the world +model. Signal values are coarse normalized linguistic indicators, not measured +probabilities.""" + + +_EXPRESS_INSTRUCTIONS = """You are Darwin's replaceable language renderer, not +its cognitive authority. Reply naturally in the requested locale to the current +user message, using only the bounded session transcript and the expression plan. +Do not claim persistent memory, a changed goal, an executed action, or an +internal state change. Do not pretend that an unverified language candidate is +a durable fact. Return every required fact id in acknowledged_fact_ids. The text +should be a direct conversational reply, not a description of this protocol.""" + + +def _json_object(value: object, field: str) -> Mapping[str, Any]: + if not isinstance(value, Mapping): + raise LanguageBackendError(f"{field}_not_object") + return value + + +def _plain_json(value: object) -> object: + """Detach gateway mapping proxies and tuples into plain JSON containers.""" + + if isinstance(value, Mapping): + return {str(key): _plain_json(child) for key, child in value.items()} + if isinstance(value, (list, tuple)): + return [_plain_json(child) for child in value] + if value is None or isinstance(value, (bool, int, float, str)): + return value + raise LanguageBackendError("language_request_contains_non_json_value") + + +def _extract_structured_output(response: Mapping[str, Any]) -> Mapping[str, Any]: + if response.get("status") != "completed": + raise LanguageBackendError("openai_response_not_completed") + output = response.get("output") + if not isinstance(output, list): + raise LanguageBackendError("openai_output_not_list") + texts: list[str] = [] + for item in output: + if not isinstance(item, Mapping) or item.get("type") != "message": + continue + content = item.get("content") + if not isinstance(content, list): + raise LanguageBackendError("openai_message_content_not_list") + for part in content: + if not isinstance(part, Mapping): + raise LanguageBackendError("openai_content_part_not_object") + if part.get("type") == "refusal": + raise LanguageBackendError("openai_response_refused") + if part.get("type") == "output_text": + text = part.get("text") + if not isinstance(text, str) or not text.strip(): + raise LanguageBackendError("openai_output_text_invalid") + texts.append(text) + if len(texts) != 1: + raise LanguageBackendError("openai_output_text_count_invalid") + try: + parsed = json.loads(texts[0]) + except json.JSONDecodeError as exc: + raise LanguageBackendError("openai_output_text_not_json") from exc + return _json_object(parsed, "openai_structured_output") + + +class OpenAIResponsesBackend: + """Two-call UNDERSTAND/EXPRESS backend with ephemeral turn context.""" + + def __init__( + self, + *, + model: str, + api_key: str, + request_timeout_seconds: float = 45.0, + transport: JSONTransport | None = None, + api_base: str = DEFAULT_OPENAI_API_BASE, + ) -> None: + self.model = require_text(model, "OpenAI model") + self._api_key = require_text(api_key, "OpenAI API key") + self._api_base = require_text(api_base, "OpenAI API base").rstrip("/") + if self._api_base != DEFAULT_OPENAI_API_BASE: + raise ValidationError("OpenAI backend requires the official API base") + if isinstance(request_timeout_seconds, bool) or not isinstance( + request_timeout_seconds, + (int, float), + ): + raise ValidationError("request timeout must be numeric") + self._request_timeout_seconds = float(request_timeout_seconds) + if not 1.0 <= self._request_timeout_seconds <= 120.0: + raise ValidationError("request timeout must be from 1 to 120 seconds") + self._transport = transport or UrllibJSONTransport() + self._pending_understanding: dict[str, Any] | None = None + self.name = f"openai-responses:{self.model}" + + def _headers(self, *, json_body: bool) -> Mapping[str, str]: + headers = { + "Accept": "application/json", + "Authorization": f"Bearer {self._api_key}", + } + if json_body: + headers["Content-Type"] = "application/json" + return headers + + def probe_model(self) -> None: + response = self._transport.request_json( + method="GET", + url=f"{self._api_base}/models/{quote(self.model, safe='')}", + headers=self._headers(json_body=False), + body=None, + timeout_seconds=self._request_timeout_seconds, + ) + if response.get("object") != "model" or response.get("id") != self.model: + raise LanguageBackendError("openai_model_probe_mismatch") + + def _structured_response( + self, + *, + operation: LanguageOperation, + instructions: str, + payload: Mapping[str, Any], + schema_name: str, + schema: Mapping[str, object], + ) -> Mapping[str, Any]: + plain_payload = _plain_json(payload) + if not isinstance(plain_payload, dict): + raise LanguageBackendError("language_request_payload_not_object") + body: Mapping[str, object] = { + "model": self.model, + "store": False, + "instructions": instructions, + "input": [ + { + "role": "user", + "content": [ + { + "type": "input_text", + "text": json.dumps( + { + "contract_version": "darwin-language-v1", + "operation": operation.value, + "payload": plain_payload, + }, + ensure_ascii=False, + separators=(",", ":"), + ), + } + ], + } + ], + "text": { + "format": { + "type": "json_schema", + "name": schema_name, + "strict": True, + "schema": deepcopy(dict(schema)), + } + }, + "max_output_tokens": 2_000, + } + response = self._transport.request_json( + method="POST", + url=f"{self._api_base}/responses", + headers=self._headers(json_body=True), + body=body, + timeout_seconds=self._request_timeout_seconds, + ) + return _extract_structured_output(response) + + def invoke(self, request: LanguageModelRequest) -> Mapping[str, Any]: + if not isinstance(request, LanguageModelRequest): + raise ValidationError("OpenAI backend requires LanguageModelRequest") + if request.operation is LanguageOperation.UNDERSTAND: + self._pending_understanding = None + result = self._structured_response( + operation=request.operation, + instructions=_UNDERSTAND_INSTRUCTIONS, + payload=request.payload, + schema_name="darwin_understanding_v1", + schema=UNDERSTANDING_SCHEMA, + ) + pending = _plain_json(request.payload) + if not isinstance(pending, dict): + raise LanguageBackendError("language_request_payload_not_object") + self._pending_understanding = pending + return result + if request.operation is LanguageOperation.EXPRESS: + pending = self._pending_understanding + self._pending_understanding = None + if pending is None: + raise LanguageBackendError("express_requires_prior_understand") + return self._structured_response( + operation=request.operation, + instructions=_EXPRESS_INSTRUCTIONS, + payload={ + "conversation_request": pending, + "expression_plan": request.payload, + }, + schema_name="darwin_expression_v1", + schema=EXPRESSION_SCHEMA, + ) + raise LanguageBackendError("openai_consult_not_enabled") + + def clear_ephemeral_context(self) -> None: + self._pending_understanding = None diff --git a/src/darwin_v50/conversation/runtime.py b/src/darwin_v50/conversation/runtime.py new file mode 100644 index 0000000..5139f98 --- /dev/null +++ b/src/darwin_v50/conversation/runtime.py @@ -0,0 +1,337 @@ +"""Session-only conversational orchestration outside the frozen E043 runtime.""" + +from __future__ import annotations + +from dataclasses import dataclass +from enum import StrEnum +from threading import RLock +from typing import Protocol + +from ..language import ( + DarwinLanguageGateway, + ExpressionPlan, + GroundedFact, + LanguageExpression, + LanguageMode, + LanguageModelBackend, + LanguageObservation, + UnderstandingRequest, +) +from ..models import DarwinV50Error, ValidationError, require_text +from .config import ConversationBackendKind, ConversationSettings +from .openai_responses import JSONTransport, OpenAIResponsesBackend + + +MAX_SESSION_MESSAGES = 60 +MAX_CONTEXT_MESSAGE_LENGTH = 2_000 + + +class ConversationRuntimeError(DarwinV50Error): + """Base error for the isolated conversational development runtime.""" + + +class ConversationUnavailableError(ConversationRuntimeError): + """Raised when no explicitly selected backend can answer a turn.""" + + +class ConversationAvailability(StrEnum): + PURE = "pure" + AVAILABLE = "available" + UNAVAILABLE = "unavailable" + CLOSED = "closed" + + +class ExplicitLocalBackend(LanguageModelBackend, Protocol): + """A local backend must declare the exact configured model it serves.""" + + model: str + + def clear_ephemeral_context(self) -> None: + """Erase any pending turn data.""" + + +@dataclass(frozen=True, slots=True) +class AuthorityMutationCounts: + memory_writes: int = 0 + goal_changes: int = 0 + rzs_changes: int = 0 + sigma_changes: int = 0 + identity_changes: int = 0 + world_model_changes: int = 0 + actions_dispatched: int = 0 + actions_executed: int = 0 + + +@dataclass(frozen=True, slots=True) +class ConversationSnapshot: + requested_backend: ConversationBackendKind + availability: ConversationAvailability + language_mode: LanguageMode + language_source: str + configured_model: str | None + unavailable_reason: str | None + completed_turns: int + temporary_messages: int + persistent_history_enabled: bool + automatic_memory_enabled: bool + authority_mutations: AuthorityMutationCounts + + +@dataclass(frozen=True, slots=True) +class ConversationTurnResult: + observation: LanguageObservation + plan: ExpressionPlan + expression: LanguageExpression + authority_mutations: AuthorityMutationCounts + + +class ConversationPolicy: + """Pure policy that keeps model interpretation explicitly provisional.""" + + def plan(self, observation: LanguageObservation, *, locale: str) -> ExpressionPlan: + if not isinstance(observation, LanguageObservation): + raise ValidationError("conversation policy requires LanguageObservation") + require_text(locale, "conversation locale") + return ExpressionPlan( + speech_act="conversation_reply", + facts=( + GroundedFact( + fact_id="candidate-status", + statement=( + "The current language interpretation is an unverified " + f"candidate with proposed intent {observation.intent!r}." + ), + ), + GroundedFact( + fact_id="authority-status", + statement=( + "This turn has not changed persistent memory, goals, " + "identity, motivation, RZS, sigma, world-model state, " + "or executed an action." + ), + ), + ), + fallback_text="The conversational backend is unavailable.", + style_hints=( + f"reply naturally in {locale}", + "do not narrate the protocol unless the user asks", + "do not claim persistent memory or completed actions", + ), + ) + + +def _bounded_context_message(role: str, text: str) -> str: + prefix = f"{role}:\n" + available = MAX_CONTEXT_MESSAGE_LENGTH - len(prefix) + if len(text) <= available: + return prefix + text + marker = "\n[truncated from temporary context]" + return prefix + text[: available - len(marker)] + marker + + +class ConversationRuntime: + """Open-ended, non-persistent conversation behind the language gateway.""" + + def __init__( + self, + *, + settings: ConversationSettings, + gateway: DarwinLanguageGateway, + availability: ConversationAvailability, + unavailable_reason: str | None, + backend_controller: object | None, + policy: ConversationPolicy | None = None, + ) -> None: + if not isinstance(settings, ConversationSettings): + raise ValidationError("conversation settings are invalid") + if availability is ConversationAvailability.AVAILABLE: + if gateway.mode is not LanguageMode.MODEL or unavailable_reason is not None: + raise ValidationError("available conversation runtime is inconsistent") + elif availability in { + ConversationAvailability.PURE, + ConversationAvailability.UNAVAILABLE, + }: + if gateway.mode is not LanguageMode.PURE or unavailable_reason is None: + raise ValidationError("inactive conversation runtime is inconsistent") + else: + raise ValidationError("new conversation runtime cannot start closed") + self._settings = settings + self._gateway = gateway + self._availability = availability + self._unavailable_reason = unavailable_reason + self._backend_controller = backend_controller + self._policy = policy or ConversationPolicy() + self._messages: list[str] = [] + self._completed_turns = 0 + self._lock = RLock() + + @classmethod + def create( + cls, + settings: ConversationSettings, + *, + openai_transport: JSONTransport | None = None, + local_backend: ExplicitLocalBackend | None = None, + ) -> "ConversationRuntime": + if not isinstance(settings, ConversationSettings): + raise ValidationError("conversation settings are invalid") + + if settings.backend is ConversationBackendKind.NONE: + return cls( + settings=settings, + gateway=DarwinLanguageGateway(), + availability=ConversationAvailability.PURE, + unavailable_reason="backend_not_requested", + backend_controller=None, + ) + + if settings.model is None: + return cls( + settings=settings, + gateway=DarwinLanguageGateway(), + availability=ConversationAvailability.UNAVAILABLE, + unavailable_reason="model_not_configured", + backend_controller=None, + ) + + if settings.backend is ConversationBackendKind.OPENAI: + if settings.api_key is None: + return cls( + settings=settings, + gateway=DarwinLanguageGateway(), + availability=ConversationAvailability.UNAVAILABLE, + unavailable_reason="openai_api_key_not_configured", + backend_controller=None, + ) + backend = OpenAIResponsesBackend( + model=settings.model, + api_key=settings.api_key, + request_timeout_seconds=settings.request_timeout_seconds, + transport=openai_transport, + ) + try: + backend.probe_model() + except Exception: + backend.clear_ephemeral_context() + return cls( + settings=settings, + gateway=DarwinLanguageGateway(), + availability=ConversationAvailability.UNAVAILABLE, + unavailable_reason="openai_model_probe_failed", + backend_controller=None, + ) + return cls( + settings=settings, + gateway=DarwinLanguageGateway(backend), + availability=ConversationAvailability.AVAILABLE, + unavailable_reason=None, + backend_controller=backend, + ) + + if local_backend is None: + return cls( + settings=settings, + gateway=DarwinLanguageGateway(), + availability=ConversationAvailability.UNAVAILABLE, + unavailable_reason="explicit_local_backend_not_supplied", + backend_controller=None, + ) + if getattr(local_backend, "model", None) != settings.model: + return cls( + settings=settings, + gateway=DarwinLanguageGateway(), + availability=ConversationAvailability.UNAVAILABLE, + unavailable_reason="local_model_mismatch", + backend_controller=None, + ) + if not callable(getattr(local_backend, "clear_ephemeral_context", None)): + return cls( + settings=settings, + gateway=DarwinLanguageGateway(), + availability=ConversationAvailability.UNAVAILABLE, + unavailable_reason="local_backend_missing_ephemeral_clear", + backend_controller=None, + ) + return cls( + settings=settings, + gateway=DarwinLanguageGateway(local_backend), + availability=ConversationAvailability.AVAILABLE, + unavailable_reason=None, + backend_controller=local_backend, + ) + + def snapshot(self) -> ConversationSnapshot: + with self._lock: + return ConversationSnapshot( + requested_backend=self._settings.backend, + availability=self._availability, + language_mode=self._gateway.mode, + language_source=self._gateway.source_name, + configured_model=self._settings.model, + unavailable_reason=self._unavailable_reason, + completed_turns=self._completed_turns, + temporary_messages=len(self._messages), + persistent_history_enabled=False, + automatic_memory_enabled=False, + authority_mutations=AuthorityMutationCounts(), + ) + + def temporary_context(self) -> tuple[str, ...]: + with self._lock: + return tuple(self._messages) + + def turn(self, text: str) -> ConversationTurnResult: + require_text(text, "conversation text") + with self._lock: + if self._availability is ConversationAvailability.CLOSED: + raise ConversationRuntimeError("conversation_runtime_closed") + if self._availability is not ConversationAvailability.AVAILABLE: + raise ConversationUnavailableError( + self._unavailable_reason or "conversation_backend_unavailable" + ) + request = UnderstandingRequest( + text=text, + locale=self._settings.locale, + recent_turns=tuple(self._messages), + ) + try: + observation = self._gateway.understand(request) + plan = self._policy.plan(observation, locale=self._settings.locale) + expression = self._gateway.express(plan) + except BaseException: + self._clear_pending_backend_context() + raise + + new_messages = [ + _bounded_context_message("user", text), + _bounded_context_message("darwin", expression.text), + ] + self._messages.extend(new_messages) + if len(self._messages) > MAX_SESSION_MESSAGES: + del self._messages[: len(self._messages) - MAX_SESSION_MESSAGES] + self._completed_turns += 1 + return ConversationTurnResult( + observation=observation, + plan=plan, + expression=expression, + authority_mutations=AuthorityMutationCounts(), + ) + + def _clear_pending_backend_context(self) -> None: + clear = getattr(self._backend_controller, "clear_ephemeral_context", None) + if callable(clear): + clear() + + def close(self) -> None: + with self._lock: + if self._availability is ConversationAvailability.CLOSED: + return + self._messages.clear() + self._clear_pending_backend_context() + self._availability = ConversationAvailability.CLOSED + + def __enter__(self) -> "ConversationRuntime": + return self + + def __exit__(self, *_: object) -> None: + self.close() diff --git a/tests/test_v50_conversation_freeze.py b/tests/test_v50_conversation_freeze.py new file mode 100644 index 0000000..f4dee21 --- /dev/null +++ b/tests/test_v50_conversation_freeze.py @@ -0,0 +1,97 @@ +from __future__ import annotations + +import hashlib +from pathlib import Path +import subprocess +from tempfile import TemporaryDirectory +import unittest + + +REPOSITORY_ROOT = Path(__file__).resolve().parents[1] +FROZEN_BASE_COMMIT = "2602c57f21dc930b6fbe4426610469b39357119f" + + +def _git_blob(revision: str, repository_path: str) -> str: + result = subprocess.run( + ["git", "rev-parse", f"{revision}:{repository_path}"], + cwd=REPOSITORY_ROOT, + check=True, + capture_output=True, + text=True, + encoding="utf-8", + ) + return result.stdout.strip() + + +def _normalized_line_ending_digest(path: Path) -> str: + normalized = path.read_bytes().replace(b"\r\n", b"\n") + if b"\r" in normalized: + raise AssertionError("frozen file contains a non-CRLF carriage return") + return hashlib.sha256(normalized).hexdigest() + + +class ConversationDevelopmentFreezeTests(unittest.TestCase): + def assert_frozen_file( + self, + *, + repository_path: str, + frozen_base_blob: str, + expected_normalized_digest: str, + ) -> None: + head_blob = _git_blob("HEAD", repository_path) + + self.assertEqual( + head_blob, + frozen_base_blob, + msg=( + f"{repository_path} at HEAD is not identical to its blob at " + f"frozen base {FROZEN_BASE_COMMIT}" + ), + ) + self.assertEqual( + _normalized_line_ending_digest(REPOSITORY_ROOT / repository_path), + expected_normalized_digest, + ) + + def test_e043_runtime_remains_identical_to_frozen_base(self) -> None: + self.assert_frozen_file( + repository_path="src/darwin_v50/desktop_runtime.py", + frozen_base_blob="01687c8e57aa1a867b18a66df9442b8745c08566", + expected_normalized_digest=( + "fd0d8aaf2bd1011addea581eaef172ce" + "940157164dc013681d5326476d49f7e8" + ), + ) + + def test_e043_protocol_remains_identical_to_frozen_base(self) -> None: + self.assert_frozen_file( + repository_path=( + "docs/v50/EXPERIMENT_043_PERSISTENT_DESKTOP_RUNTIME.md" + ), + frozen_base_blob="5becbc0177c7fc1fc1cfe0e2903e7a8323a7bd7d", + expected_normalized_digest=( + "beea70ce6bcfd177a18d36feca016f5f9" + "3ed0bed7743fc73c31152cc19eaefad" + ), + ) + + def test_freeze_digest_normalizes_only_crlf_and_lf(self) -> None: + with TemporaryDirectory() as temporary: + root = Path(temporary) + lf_path = root / "lf.txt" + crlf_path = root / "crlf.txt" + lone_cr_path = root / "lone-cr.txt" + lf_path.write_bytes(b"alpha\nbeta\n") + crlf_path.write_bytes(b"alpha\r\nbeta\r\n") + lone_cr_path.write_bytes(b"alpha\rbeta\r") + + self.assertEqual( + _normalized_line_ending_digest(lf_path), + _normalized_line_ending_digest(crlf_path), + ) + with self.assertRaisesRegex(AssertionError, "non-CRLF"): + _normalized_line_ending_digest(lone_cr_path) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_v50_conversation_runtime.py b/tests/test_v50_conversation_runtime.py new file mode 100644 index 0000000..8d5c044 --- /dev/null +++ b/tests/test_v50_conversation_runtime.py @@ -0,0 +1,306 @@ +from __future__ import annotations + +from dataclasses import asdict +from pathlib import Path +from tempfile import TemporaryDirectory +import unittest +from typing import Any, Mapping + +from darwin_v50.conversation import ( + ConversationAvailability, + ConversationBackendKind, + ConversationRuntime, + ConversationRuntimeError, + ConversationSettings, + ConversationUnavailableError, +) +from darwin_v50.language import ( + LanguageAuthorityError, + LanguageBackendError, + LanguageMode, + LanguageModelRequest, + LanguageOperation, +) +from darwin_v50.models import ValidationError + + +def understanding_response() -> dict[str, Any]: + return { + "intent": "open_conversation", + "entities": [], + "reported_signals": [], + "temporal_reference": None, + "explicit_preference": None, + "confidence": 0.6, + } + + +class ScriptedLocalBackend: + name = "explicit-local-test-backend" + + def __init__( + self, + *, + model: str = "local-test-model", + fail_expression: bool = False, + authority_violation: bool = False, + ) -> None: + self.model = model + self.fail_expression = fail_expression + self.authority_violation = authority_violation + self.requests: list[LanguageModelRequest] = [] + self.clear_calls = 0 + + def invoke(self, request: LanguageModelRequest) -> Mapping[str, Any]: + self.requests.append(request) + if request.operation is LanguageOperation.UNDERSTAND: + result = understanding_response() + if self.authority_violation: + result["memory"] = {"write": "forbidden"} + return result + if request.operation is LanguageOperation.EXPRESS: + if self.fail_expression: + raise LanguageBackendError("scripted_expression_failure") + facts = request.payload["facts"] + return { + "text": "Uma resposta nova, produzida para este turno.", + "acknowledged_fact_ids": [fact["fact_id"] for fact in facts], + } + raise AssertionError("consult must not be called") + + def clear_ephemeral_context(self) -> None: + self.clear_calls += 1 + + +def local_settings(**overrides: object) -> ConversationSettings: + values: dict[str, object] = { + "backend": ConversationBackendKind.LOCAL, + "model": "local-test-model", + "locale": "pt-BR", + } + values.update(overrides) + return ConversationSettings(**values) # type: ignore[arg-type] + + +class ConversationConfigurationTests(unittest.TestCase): + def test_environment_defaults_to_no_backend_and_no_model(self) -> None: + settings = ConversationSettings.from_environment({}) + + self.assertEqual(settings.backend, ConversationBackendKind.NONE) + self.assertIsNone(settings.model) + self.assertIsNone(settings.api_key) + + def test_environment_never_infers_backend_from_available_api_key(self) -> None: + settings = ConversationSettings.from_environment( + { + "OPENAI_API_KEY": "secret-key", + "DARWIN_LLM_MODEL": "account-model", + } + ) + + self.assertEqual(settings.backend, ConversationBackendKind.NONE) + + def test_invalid_or_blank_backend_is_rejected(self) -> None: + for backend in ("ollama", "auto", ""): + with self.subTest(backend=backend): + with self.assertRaises(ValidationError): + ConversationSettings.from_environment( + {"DARWIN_LLM_BACKEND": backend} + ) + + def test_api_key_is_excluded_from_settings_repr(self) -> None: + settings = ConversationSettings( + backend=ConversationBackendKind.OPENAI, + model="account-model", + api_key="never-print-this-secret", + ) + + self.assertNotIn("never-print-this-secret", repr(settings)) + + def test_missing_openai_configuration_is_unavailable_not_local(self) -> None: + runtime = ConversationRuntime.create( + ConversationSettings( + backend=ConversationBackendKind.OPENAI, + model="account-model", + ), + local_backend=ScriptedLocalBackend(model="account-model"), + ) + + snapshot = runtime.snapshot() + self.assertEqual(snapshot.availability, ConversationAvailability.UNAVAILABLE) + self.assertEqual(snapshot.language_mode, LanguageMode.PURE) + self.assertEqual(snapshot.unavailable_reason, "openai_api_key_not_configured") + with self.assertRaises(ConversationUnavailableError): + runtime.turn("Olá") + + def test_local_backend_requires_explicit_instance_and_matching_model(self) -> None: + absent = ConversationRuntime.create(local_settings()) + mismatch = ConversationRuntime.create( + local_settings(), + local_backend=ScriptedLocalBackend(model="different-model"), + ) + + self.assertEqual(absent.snapshot().availability, ConversationAvailability.UNAVAILABLE) + self.assertEqual( + absent.snapshot().unavailable_reason, + "explicit_local_backend_not_supplied", + ) + self.assertEqual(mismatch.snapshot().availability, ConversationAvailability.UNAVAILABLE) + self.assertEqual(mismatch.snapshot().unavailable_reason, "local_model_mismatch") + + def test_local_backend_must_expose_ephemeral_clear(self) -> None: + class MissingClearBackend: + name = "missing-clear" + model = "local-test-model" + + def invoke(self, request: LanguageModelRequest) -> Mapping[str, Any]: + raise AssertionError("unavailable backend must not be invoked") + + runtime = ConversationRuntime.create( + local_settings(), + local_backend=MissingClearBackend(), # type: ignore[arg-type] + ) + + self.assertEqual(runtime.snapshot().availability, ConversationAvailability.UNAVAILABLE) + self.assertEqual( + runtime.snapshot().unavailable_reason, + "local_backend_missing_ephemeral_clear", + ) + + def test_snapshot_never_contains_api_key(self) -> None: + class ProbeTransport: + def request_json(self, **_: object) -> Mapping[str, Any]: + return {"object": "model", "id": "account-model"} + + runtime = ConversationRuntime.create( + ConversationSettings( + backend=ConversationBackendKind.OPENAI, + model="account-model", + api_key="never-print-this-secret", + ), + openai_transport=ProbeTransport(), # type: ignore[arg-type] + ) + + self.assertNotIn("never-print-this-secret", repr(runtime.snapshot())) + self.assertNotIn("never-print-this-secret", repr(asdict(runtime.snapshot()))) + + +class ConversationRuntimeTests(unittest.TestCase): + def test_successful_turn_is_understand_then_express(self) -> None: + backend = ScriptedLocalBackend() + runtime = ConversationRuntime.create( + local_settings(), + local_backend=backend, + ) + + result = runtime.turn("Vamos falar sobre astronomia?") + + self.assertEqual( + [request.operation for request in backend.requests], + [LanguageOperation.UNDERSTAND, LanguageOperation.EXPRESS], + ) + self.assertEqual(result.observation.intent, "open_conversation") + self.assertEqual( + result.expression.text, + "Uma resposta nova, produzida para este turno.", + ) + self.assertEqual( + result.expression.acknowledged_fact_ids, + ("candidate-status", "authority-status"), + ) + self.assertEqual(result.authority_mutations.memory_writes, 0) + self.assertEqual(result.authority_mutations.goal_changes, 0) + self.assertEqual(result.authority_mutations.rzs_changes, 0) + self.assertEqual(result.authority_mutations.sigma_changes, 0) + self.assertEqual(result.authority_mutations.actions_executed, 0) + + def test_later_turn_receives_bounded_temporary_transcript(self) -> None: + backend = ScriptedLocalBackend() + runtime = ConversationRuntime.create(local_settings(), local_backend=backend) + runtime.turn("Primeiro assunto") + runtime.turn("Agora outro assunto") + + second_understand = backend.requests[2] + self.assertEqual( + tuple(second_understand.payload["recent_turns"]), + ( + "user:\nPrimeiro assunto", + "darwin:\nUma resposta nova, produzida para este turno.", + ), + ) + + def test_failed_expression_does_not_commit_partial_turn(self) -> None: + backend = ScriptedLocalBackend(fail_expression=True) + runtime = ConversationRuntime.create(local_settings(), local_backend=backend) + + with self.assertRaisesRegex( + LanguageBackendError, + "scripted_expression_failure", + ): + runtime.turn("Este turno deve falhar") + + self.assertEqual(runtime.temporary_context(), ()) + self.assertEqual(runtime.snapshot().completed_turns, 0) + self.assertEqual(backend.clear_calls, 1) + + def test_authority_bearing_understanding_is_rejected_without_context_write(self) -> None: + backend = ScriptedLocalBackend(authority_violation=True) + runtime = ConversationRuntime.create(local_settings(), local_backend=backend) + + with self.assertRaises(LanguageAuthorityError): + runtime.turn("Grave isso diretamente na memória") + + self.assertEqual(runtime.temporary_context(), ()) + self.assertEqual(runtime.snapshot().authority_mutations.memory_writes, 0) + + def test_context_is_bounded_to_sixty_messages(self) -> None: + backend = ScriptedLocalBackend() + runtime = ConversationRuntime.create(local_settings(), local_backend=backend) + + for index in range(31): + runtime.turn(f"turno {index}") + + context = runtime.temporary_context() + self.assertEqual(len(context), 60) + self.assertNotIn("user:\nturno 0", context) + self.assertIn("user:\nturno 30", context) + + def test_long_messages_are_marked_when_clipped_from_future_context(self) -> None: + backend = ScriptedLocalBackend() + runtime = ConversationRuntime.create(local_settings(), local_backend=backend) + runtime.turn("x" * 3_000) + + first_message = runtime.temporary_context()[0] + self.assertEqual(len(first_message), 2_000) + self.assertTrue(first_message.endswith("[truncated from temporary context]")) + + def test_close_erases_transcript_and_pending_state(self) -> None: + backend = ScriptedLocalBackend() + runtime = ConversationRuntime.create(local_settings(), local_backend=backend) + runtime.turn("Conversa temporária") + + runtime.close() + + snapshot = runtime.snapshot() + self.assertEqual(snapshot.availability, ConversationAvailability.CLOSED) + self.assertEqual(snapshot.temporary_messages, 0) + self.assertEqual(runtime.temporary_context(), ()) + self.assertEqual(backend.clear_calls, 1) + with self.assertRaisesRegex(ConversationRuntimeError, "closed"): + runtime.turn("não deve funcionar") + + def test_conversation_runtime_creates_no_files(self) -> None: + backend = ScriptedLocalBackend() + with TemporaryDirectory() as temporary: + root = Path(temporary) + before = tuple(root.rglob("*")) + runtime = ConversationRuntime.create(local_settings(), local_backend=backend) + runtime.turn("Nada deve ser persistido") + runtime.close() + after = tuple(root.rglob("*")) + + self.assertEqual(before, after) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_v50_openai_responses_backend.py b/tests/test_v50_openai_responses_backend.py new file mode 100644 index 0000000..d8e2b01 --- /dev/null +++ b/tests/test_v50_openai_responses_backend.py @@ -0,0 +1,276 @@ +from __future__ import annotations + +from copy import deepcopy +import json +import unittest +from typing import Any, Mapping + +from darwin_v50.conversation import ( + ConversationAvailability, + ConversationBackendKind, + ConversationRuntime, + ConversationSettings, + OpenAIResponsesBackend, +) +from darwin_v50.language import ( + DarwinLanguageGateway, + ExpressionPlan, + GroundedFact, + LanguageBackendError, + LanguageModelRequest, + LanguageOperation, + UnderstandingRequest, +) + + +def completed_output(payload: Mapping[str, Any]) -> dict[str, Any]: + return { + "status": "completed", + "output": [ + { + "type": "message", + "role": "assistant", + "content": [ + { + "type": "output_text", + "text": json.dumps(payload, ensure_ascii=False), + } + ], + } + ], + } + + +def understanding_payload() -> dict[str, Any]: + return { + "intent": "ask_question", + "entities": [{"kind": "topic", "value": "estrelas"}], + "reported_signals": [], + "temporal_reference": None, + "explicit_preference": None, + "confidence": 0.75, + } + + +class CapturingTransport: + def __init__(self, *responses: object) -> None: + self.responses = list(responses) + self.calls: list[dict[str, Any]] = [] + + def request_json( + self, + *, + method: str, + url: str, + headers: Mapping[str, str], + body: Mapping[str, object] | None, + timeout_seconds: float, + ) -> Mapping[str, Any]: + self.calls.append( + { + "method": method, + "url": url, + "headers": dict(headers), + "body": deepcopy(body), + "timeout_seconds": timeout_seconds, + } + ) + if not self.responses: + raise AssertionError("unexpected transport call") + response = self.responses.pop(0) + if isinstance(response, BaseException): + raise response + if not isinstance(response, Mapping): + raise AssertionError("scripted response must be a mapping") + return response + + +def expression_plan() -> ExpressionPlan: + return ExpressionPlan( + speech_act="answer", + facts=(GroundedFact("fact-1", "No persistent state changed."),), + fallback_text="Backend unavailable.", + ) + + +class OpenAIResponsesBackendTests(unittest.TestCase): + def test_probe_uses_exact_configured_model(self) -> None: + transport = CapturingTransport( + {"object": "model", "id": "account/model:revision"} + ) + backend = OpenAIResponsesBackend( + model="account/model:revision", + api_key="test-secret", + transport=transport, + ) + + backend.probe_model() + + call = transport.calls[0] + self.assertEqual(call["method"], "GET") + self.assertTrue(call["url"].endswith("/models/account%2Fmodel%3Arevision")) + self.assertIsNone(call["body"]) + + def test_understand_and_express_use_responses_store_false(self) -> None: + transport = CapturingTransport( + completed_output(understanding_payload()), + completed_output( + { + "text": "As estrelas nascem em nuvens moleculares.", + "acknowledged_fact_ids": ["fact-1"], + } + ), + ) + backend = OpenAIResponsesBackend( + model="configured-model", + api_key="test-secret", + transport=transport, + ) + gateway = DarwinLanguageGateway(backend) + + gateway.understand( + UnderstandingRequest( + "Como nascem as estrelas?", + locale="pt-BR", + recent_turns=("user:\nFalávamos sobre o céu.",), + ) + ) + expression = gateway.express(expression_plan()) + + self.assertEqual(expression.text, "As estrelas nascem em nuvens moleculares.") + self.assertEqual(len(transport.calls), 2) + for call in transport.calls: + body = call["body"] + self.assertEqual(call["method"], "POST") + self.assertTrue(call["url"].endswith("/responses")) + self.assertEqual(body["model"], "configured-model") + self.assertIs(body["store"], False) + self.assertNotIn("previous_response_id", body) + self.assertNotIn("tools", body) + self.assertNotIn("tool_choice", body) + self.assertEqual(body["text"]["format"]["type"], "json_schema") + self.assertIs(body["text"]["format"]["strict"], True) + + understand_input = json.loads( + transport.calls[0]["body"]["input"][0]["content"][0]["text"] + ) + express_input = json.loads( + transport.calls[1]["body"]["input"][0]["content"][0]["text"] + ) + self.assertEqual(understand_input["operation"], "understand") + self.assertEqual(express_input["operation"], "express") + self.assertEqual( + express_input["payload"]["conversation_request"]["text"], + "Como nascem as estrelas?", + ) + + def test_express_requires_immediately_prior_understanding(self) -> None: + backend = OpenAIResponsesBackend( + model="configured-model", + api_key="test-secret", + transport=CapturingTransport(), + ) + gateway = DarwinLanguageGateway(backend) + + with self.assertRaisesRegex( + LanguageBackendError, + "express_requires_prior_understand", + ): + gateway.express(expression_plan()) + + def test_pending_context_is_one_use_only(self) -> None: + transport = CapturingTransport( + completed_output(understanding_payload()), + completed_output( + {"text": "Resposta.", "acknowledged_fact_ids": ["fact-1"]} + ), + ) + gateway = DarwinLanguageGateway( + OpenAIResponsesBackend( + model="configured-model", + api_key="test-secret", + transport=transport, + ) + ) + gateway.understand(UnderstandingRequest("Pergunta")) + gateway.express(expression_plan()) + + with self.assertRaisesRegex( + LanguageBackendError, + "express_requires_prior_understand", + ): + gateway.express(expression_plan()) + + def test_refusal_incomplete_and_malformed_outputs_fail_closed(self) -> None: + cases = ( + { + "status": "completed", + "output": [ + { + "type": "message", + "content": [{"type": "refusal", "refusal": "no"}], + } + ], + }, + {"status": "incomplete", "output": []}, + { + "status": "completed", + "output": [ + { + "type": "message", + "content": [{"type": "output_text", "text": "not json"}], + } + ], + }, + ) + for response in cases: + with self.subTest(response=response): + backend = OpenAIResponsesBackend( + model="configured-model", + api_key="test-secret", + transport=CapturingTransport(response), + ) + with self.assertRaises(LanguageBackendError): + backend.invoke( + LanguageModelRequest( + contract_version="darwin-language-v1", + operation=LanguageOperation.UNDERSTAND, + payload={"text": "Olá", "locale": "pt-BR", "recent_turns": []}, + ) + ) + + def test_openai_probe_failure_does_not_select_supplied_local_backend(self) -> None: + class LocalTrap: + name = "must-not-be-used" + model = "configured-model" + + def __init__(self) -> None: + self.calls = 0 + + def invoke(self, request: LanguageModelRequest) -> Mapping[str, Any]: + self.calls += 1 + raise AssertionError("silent local fallback occurred") + + def clear_ephemeral_context(self) -> None: + pass + + local = LocalTrap() + runtime = ConversationRuntime.create( + ConversationSettings( + backend=ConversationBackendKind.OPENAI, + model="configured-model", + api_key="test-secret", + ), + openai_transport=CapturingTransport( + LanguageBackendError("model_unavailable") + ), + local_backend=local, + ) + + self.assertEqual(runtime.snapshot().availability, ConversationAvailability.UNAVAILABLE) + self.assertEqual(runtime.snapshot().unavailable_reason, "openai_model_probe_failed") + self.assertEqual(local.calls, 0) + + +if __name__ == "__main__": + unittest.main()