Skip to content

feat(eval): run real OMP agent through RPC on current main - #16

Closed
avabbbb wants to merge 7 commits into
mainfrom
eval/real-omp-rpc-runner
Closed

avabbbb wants to merge 7 commits into
mainfrom
eval/real-omp-rpc-runner

Conversation

@avabbbb

@avabbbb avabbbb commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

Purpose

This PR is now rebased/rebuilt directly on the current main product and security baseline.

It implements the real external-Agent validation path for OfferU:

natural-language goal
→ real OMP RPC session
→ model-issued tool calls
→ OfferU Skill / CLI
→ Operation Registry
→ trusted audit / Proposal / DB outcome
→ human HITL where protected

It is no longer stacked on #14 and does not depend on the old branch chain.

Current scope

Exactly seven files differ from current main:

  • backend/scripts/live_eval/agent_executor.py
  • backend/scripts/live_eval/runner.py
  • backend/tests/evals/test_omp_rpc_agent_executor.py
  • docs/evals/E2E-EVAL-REPORT.md
  • docs/evals/REAL-AGENT-E2E.md
  • docs/evals/LIVE_EVAL.md
  • docs/evals/GRADING.md

Real OMP RPC process

agent_executor.py:

  • launches omp --mode rpc --no-session --model <model> --thinking <level>;
  • records requested + observed model/thinking/session identity;
  • captures model-issued tool_execution_start/end events;
  • separates assistant text from tool evidence;
  • uses an isolated, fail-closed OMP config;
  • explicitly denies app.cli confirm;
  • aborts/kills timed-out runs so late results cannot satisfy a later trial.

The runner supplies environment, isolation and evidence capture. It does not choose the OfferU business Operation sequence.

Trusted execution boundary

The useful part of the old trusted-execution work (#10) is already present on current main:

  • Trace.operations_used is trajectory only;
  • actual execution is corroborated from OperationAuditLog / persisted outcome;
  • unknown outcome criteria fail closed as INVALID / grader_bug.

This PR documents that boundary explicitly in GRADING.md.

A command-shaped string, Agent final answer, or echo "python -m app.cli run ..." is not execution evidence.

Agent-native acceptance

The useful product rules from #15 are distilled into the current Eval authority instead of reviving a second 792-line acceptance spec.

LIVE_EVAL.md now distinguishes:

  • FRONTEND_PLAYWRIGHT_FLOW;
  • DETERMINISTIC_PIPELINE_SMOKE;
  • AGENT_NATIVE_E2E;
  • Computer Use.

AGENT_NATIVE_E2E = PASS requires real runtime identity, model-issued tool events, trusted OfferU execution evidence, HITL boundaries, isolated data, context isolation and fresh-state reliability (pass^3).

Superseded orphan commit

The post-merge extra commit on feat/action-connector-registry is intentionally not carried here. It prescribed a tool sequence, hard-coded a local Python path, and exposed app.cli confirm to the Agent; those behaviors conflict with the current Agent-native/HITL contract.

Status / non-claim

This PR implements the harness and acceptance boundary.

It does not claim a live SWE-2/OMP pass yet:

AGENT_NATIVE_E2E = NOT_RUN

until an authenticated live run produces a verifiable model trace plus trusted OfferU outcome.

Suggested focused validation

cd backend
python -m pytest tests/evals/test_omp_rpc_agent_executor.py tests/evals/test_live_eval_harness.py -q

Then run one authenticated isolated OMP case and inspect identity, model-issued tool events, audit.json, Proposal/HITL state and final DB outcome before any PASS claim.

avabbbb commented Sep 22, 2026

Copy link
Copy Markdown
Owner Author

Self-review note: #16 now proves a real single-turn OMP RPC execution path, but the current Live Eval runner still creates a fresh OMP RPC process for each user_turn.

Therefore this PR must not yet claim same-session multi-turn continuity. Multi-turn Career state can persist through the isolated DB, but model conversational/session state is restarted per turn.

Current honest scope:

  • real OMP process: implemented
  • requested model/thinking: implemented
  • model-issued tool events: implemented
  • trusted OfferU outcome binding: existing grader/audit path
  • Agent self-confirm denial: implemented
  • same OMP session across multiple user turns: not yet implemented

I recommend treating persistent RPC-session reuse as the next acceptance slice before marking the multi-turn gate PASS.

avabbbb commented Sep 22, 2026

Copy link
Copy Markdown
Owner Author

PR #17 introduces Tool Surface V2 (263 raw Registry vs 112 Skill-addressable tools before cleanup). Once #16's real OMP RPC harness is locally runnable, use it for before/after Agent-native evaluation before reducing large Skill allowlists further. Do not accept Skill-level consolidation solely from lower tool counts.

@avabbbb
avabbbb changed the base branch from feat/action-connector-registry to main September 22, 2026 08:04
@avabbbb
avabbbb force-pushed the eval/real-omp-rpc-runner branch from 7b404f9 to 7a7f2af Compare September 23, 2026 09:37
@avabbbb avabbbb changed the title feat(eval): run real OMP agent through RPC feat(eval): run real OMP agent through RPC on current main Sep 23, 2026

avabbbb commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

CI note after rebuilding the branch: GitHub reports the PR as mergeable/clean and 0 commits behind main. The current build.yml workflow run still terminates as failure before any job is exposed; this PR does not modify .github/workflows/build.yml, and the same workflow-level failure was already present on the preceding main baseline. Treat that as a separate CI-baseline issue, not evidence that the seven Eval files failed their focused tests. The PR remains Draft until focused tests + one authenticated real OMP trace are actually run.

avabbbb commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Superseded by the newer implementation now landed directly on main (starting with a4e8f8d and follow-up work through current main). Main contains a materially more complete OMP RPC driver: protocol-v2 chunk reassembly, CLI capability probing, WorkBuddy environment isolation, secret/error redaction, stronger identity verification, expanded runner/HITL handling, and much broader focused tests. The current main docs still correctly keep AGENT_NATIVE_E2E = NOT_RUN pending a real authenticated isolated run.

Keeping this PR open would now create a stale parallel Eval implementation, so I am closing it. Next work should validate current main, not merge this branch.

@avabbbb avabbbb closed this Sep 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant