Repository navigation
Conversation
|
Self-review note: #16 now proves a real single-turn OMP RPC execution path, but the current Live Eval runner still creates a fresh OMP RPC process for each Therefore this PR must not yet claim same-session multi-turn continuity. Multi-turn Career state can persist through the isolated DB, but model conversational/session state is restarted per turn. Current honest scope:
I recommend treating persistent RPC-session reuse as the next acceptance slice before marking the multi-turn gate PASS. |
|
PR #17 introduces Tool Surface V2 (263 raw Registry vs 112 Skill-addressable tools before cleanup). Once #16's real OMP RPC harness is locally runnable, use it for before/after Agent-native evaluation before reducing large Skill allowlists further. Do not accept Skill-level consolidation solely from lower tool counts. |
…ripts/live_eval/agent_executor.py)
…ripts/live_eval/runner.py)
…/E2E-EVAL-REPORT.md)
…/REAL-AGENT-E2E.md)
7b404f9 to
7a7f2af
Compare
|
CI note after rebuilding the branch: GitHub reports the PR as mergeable/clean and 0 commits behind main. The current build.yml workflow run still terminates as failure before any job is exposed; this PR does not modify .github/workflows/build.yml, and the same workflow-level failure was already present on the preceding main baseline. Treat that as a separate CI-baseline issue, not evidence that the seven Eval files failed their focused tests. The PR remains Draft until focused tests + one authenticated real OMP trace are actually run. |
|
Superseded by the newer implementation now landed directly on main (starting with a4e8f8d and follow-up work through current main). Main contains a materially more complete OMP RPC driver: protocol-v2 chunk reassembly, CLI capability probing, WorkBuddy environment isolation, secret/error redaction, stronger identity verification, expanded runner/HITL handling, and much broader focused tests. The current main docs still correctly keep AGENT_NATIVE_E2E = NOT_RUN pending a real authenticated isolated run. Keeping this PR open would now create a stale parallel Eval implementation, so I am closing it. Next work should validate current main, not merge this branch. |
Purpose
This PR is now rebased/rebuilt directly on the current
mainproduct and security baseline.It implements the real external-Agent validation path for OfferU:
It is no longer stacked on #14 and does not depend on the old branch chain.
Current scope
Exactly seven files differ from current
main:backend/scripts/live_eval/agent_executor.pybackend/scripts/live_eval/runner.pybackend/tests/evals/test_omp_rpc_agent_executor.pydocs/evals/E2E-EVAL-REPORT.mddocs/evals/REAL-AGENT-E2E.mddocs/evals/LIVE_EVAL.mddocs/evals/GRADING.mdReal OMP RPC process
agent_executor.py:omp --mode rpc --no-session --model <model> --thinking <level>;tool_execution_start/endevents;app.cli confirm;The runner supplies environment, isolation and evidence capture. It does not choose the OfferU business Operation sequence.
Trusted execution boundary
The useful part of the old trusted-execution work (#10) is already present on current
main:Trace.operations_usedis trajectory only;OperationAuditLog/ persisted outcome;INVALID / grader_bug.This PR documents that boundary explicitly in
GRADING.md.A command-shaped string, Agent final answer, or
echo "python -m app.cli run ..."is not execution evidence.Agent-native acceptance
The useful product rules from #15 are distilled into the current Eval authority instead of reviving a second 792-line acceptance spec.
LIVE_EVAL.mdnow distinguishes:FRONTEND_PLAYWRIGHT_FLOW;DETERMINISTIC_PIPELINE_SMOKE;AGENT_NATIVE_E2E;AGENT_NATIVE_E2E = PASSrequires real runtime identity, model-issued tool events, trusted OfferU execution evidence, HITL boundaries, isolated data, context isolation and fresh-state reliability (pass^3).Superseded orphan commit
The post-merge extra commit on
feat/action-connector-registryis intentionally not carried here. It prescribed a tool sequence, hard-coded a local Python path, and exposedapp.cli confirmto the Agent; those behaviors conflict with the current Agent-native/HITL contract.Status / non-claim
This PR implements the harness and acceptance boundary.
It does not claim a live SWE-2/OMP pass yet:
until an authenticated live run produces a verifiable model trace plus trusted OfferU outcome.
Suggested focused validation
Then run one authenticated isolated OMP case and inspect identity, model-issued tool events,
audit.json, Proposal/HITL state and final DB outcome before any PASS claim.