Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
47 changes: 47 additions & 0 deletions loopx/capabilities/reward_memory/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -144,6 +144,53 @@ verified. Issue Fix continues normally when the experiment is disabled,
unavailable, rejected by guards, or fails exact readback. Invalid or non-v1
configuration resolves unavailable with both automatic flags false.

### Inbox feedback review (explicit ingestion)

When a registry-routed `loopx lark-inbox drain --goal-id <goal> --agent-id
<agent>` returns messages, it also emits `reward_memory_feedback_review` if
Reward Memory is enabled for that agent and has an active, writable
`scoped_feedback` route with an enabled standing policy and exact
`peer_ref=agent:<agent>`. The hint lists configured destinations and a preview
command bound to the same registry, Goal and Agent. It is advisory: there is no
provider call, automatic candidate creation, new permission, or ACK gate.
Both JSON and Markdown drains expose it. Empty/disabled inboxes, disabled or
invalid memory configurations, unconfigured agents, incompatible routes, and
explicit `--config`/`--project` overrides retain their previous output.

The agent reviews the conversation before choosing what, if anything, to learn:

1. Verify the source actor and existing authority, current evidence, conflicts
and applicability. A policy's allowed roles are not proof of a sender's role.
A disagreement is not a universal ban, and one-off task state belongs in
Todo/vision rather than durable preference memory.
2. Distill only confirmed reusable feedback into an applicable configured
`soft_preference`, `procedural_experience`, or independently authorized
`hard_policy` route. Never widen the route or enable automation to make an
event pass. Keep raw chat and credentials out of the event.
3. Prepare `{adapter, event, observed_at}` using the
[scoped feedback fixture](../../../examples/fixtures/reward-memory-scoped-feedback-ingest.public.json)
for field shape only. The event uses
`schema_version=scoped_feedback_reward_memory_event_v0`, a stable
`feedback_ref`, actual `source`, `reasoning`, `guard_context`, compact
`content_summary`, `target_class`, and exact identity/surface/revision/action
scope. Advisory classes require empty `requested_action_scopes`; allowed
policy scopes do not grant advisory memory action authority. Do not copy the
fixture's actor or verified-guard assertions.
4. Replace the hint's input placeholder and preview `ingest-event` without
`--execute`. Inspect its guards. Only then execute within the standing
policy and verify `exact_readback_verified` and
`memory_available_for_recall`; a preview is not a learned memory.
5. Finish normal reply/material-review/ACK. No reusable feedback or unavailable
memory is an honest no-memory outcome, not a reason to block the inbox or
repeatedly ask the user for permission. Actual write failures remain visible
in the existing ingestion receipt.

`automatic_ingest=false` does **not** prohibit this explicit workflow. Enabling
it also does **not** wire raw inbox messages into ingestion. Issue Fix's compact
feedback adapter and recall hooks share the same core but are not a general
inbox-to-memory feedback loop. Disable the hint by disabling Reward Memory for
the agent (or its applicable route); no extra store, queue or scheduler exists.

## Five first-class classes

| Class | Source and scope | Authority and use | Lifecycle |
Expand Down
107 changes: 107 additions & 0 deletions loopx/capabilities/reward_memory/feedback_hint.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,107 @@
"""Effect-free guidance for agents reviewing an inbox's captured feedback."""

from __future__ import annotations

import shlex
from pathlib import Path
from typing import Any

from .experiment import (
resolve_reward_memory_experiment,
resolve_reward_memory_surface_config,
)
from .scoped_feedback import SCOPED_FEEDBACK_ADAPTER


def build_feedback_review_hint(
*, registry_path: Path, goal_id: str | None, agent_id: str | None
) -> dict[str, Any] | None:
"""Describe explicit ingestion, never inspect messages or call a provider.

The caller must first establish an enabled, registry-routed inbox with
returned items. Automation flags govern hooks, not this explicit path.
"""
if not goal_id or not agent_id:
return None
try:
status, config = resolve_reward_memory_experiment(
registry_path=registry_path, goal_id=goal_id, agent_id=agent_id
)
except (OSError, ValueError):
return None
if config is None:
return None
routes = []
for surface_id, surface in sorted(config["surfaces"].items()):
if surface["adapter"] != SCOPED_FEEDBACK_ADAPTER:
continue
route = resolve_reward_memory_surface_config(config, surface_id)
corpus, policy = route["corpus"], route["standing_policy"]
scope = corpus["scope"]
if (
not policy["enabled"]
or scope.get("peer_ref") != f"agent:{agent_id}"
or corpus["lifecycle"]["state"] != "active"
or corpus["write_authority"] in {"read_only", "ephemeral_runtime"}
):
continue
routes.append(
{
"surface_id": surface_id,
"corpus_id": corpus["corpus_id"],
"target_class": corpus["class_id"],
"scope": scope | {"surface_ids": [surface_id]},
"allowed_source_kinds": policy["allowed_source_kinds"],
"allowed_actor_roles": policy["allowed_actor_roles"],
"allowed_action_scopes": policy["allowed_action_scopes"],
}
)
if not routes:
return None
preview = [
"loopx",
"--format",
"json",
"--registry",
str(registry_path.expanduser()),
"reward-memory",
"ingest-event",
"--goal-id",
goal_id,
"--agent-id",
agent_id,
"--input",
"<compact-event.json>",
]
return {
"schema_version": "reward_memory_feedback_review_hint_v0",
"advisory_only": True,
"automatic_ingest": status["automatic_ingest"],
"automatic_ingest_required": False,
"grants_new_action_authority": False,
"blocks_inbox_settlement": False,
"provider_calls_performed": False,
"routes": routes,
"preview_command": shlex.join(preview),
"instruction": (
"While triaging these messages, consider confirmed, reusable feedback for "
"Reward Memory. Distill a compact scoped lesson, not raw chat; verify the "
"source actor, authority, freshness, current artifact and conflicts. "
"Choose only an applicable configured route below; allowed actor roles "
"are constraints, not proof of the sender's authority. Do not turn "
"disagreement or a one-off opinion into a universal prohibition. "
"Use soft preferences or procedural experience where applicable; hard "
"policy still requires independently verified existing authority. "
"For advisory classes, requested_action_scopes must be empty. "
"Prepare {adapter: scoped_feedback, event: scoped_feedback_reward_memory_event_v0, "
"observed_at} using the documented event fields and a stable feedback_ref; "
"replace the input placeholder and preview before adding --execute. "
"Do not copy fixture actor/guard assertions or expand scope to pass guards. "
"Automatic ingest being off does not disable explicit ingest-event. "
"After an authorized write, inspect the ingest receipt and exact readback; "
"do not claim learning from a preview or failed write. If no reusable "
"lesson exists, evidence conflicts, or memory is unavailable, continue "
"normal reply/material-review/ACK with an honest rationale; no new user gate."
),
"event_reference": "loopx/capabilities/reward_memory/README.md",
}
19 changes: 19 additions & 0 deletions loopx/cli_commands/lark_inbox.py
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@
from pathlib import Path

from ..capabilities.issue_fix.provider_hooks import IssueFixReviewerProviderHooks
from ..capabilities.reward_memory.feedback_hint import build_feedback_review_hint
from ..capabilities.reward_memory.outbound import outbound_guidance_hook
from ..control_plane.capability_hooks import (
TURN_START_HOOK_RESULT_SCHEMA_VERSION,
Expand Down Expand Up @@ -555,6 +556,11 @@ def _render(payload: dict[str, object]) -> str:
+ str(guidance.get("review_digest"))
+ ". This is not a request for user approval."
)
feedback = payload.get("reward_memory_feedback_review")
if isinstance(feedback, dict):
lines.append(f"- Reward Memory (advisory): {feedback['instruction']}")
lines.append(f"- preview: {feedback['preview_command']}")
lines.append("- configured routes: " + json.dumps(feedback["routes"]))
return "\n".join(lines).rstrip() + "\n"


Expand Down Expand Up @@ -602,6 +608,19 @@ def handle_lark_inbox_command(
config_path=config_path,
limit=args.limit,
)
if (
payload.get("enabled") is True
and payload.get("items")
and not getattr(args, "config", None)
and not getattr(args, "project", None)
):
hint = build_feedback_review_hint(
registry_path=registry_path,
goal_id=getattr(args, "goal_id", None),
agent_id=getattr(args, "agent_id", None),
)
if hint is not None:
payload["reward_memory_feedback_review"] = hint
elif args.lark_inbox_command == "ack":
payload = acknowledge_routed_lark_event_inbox(
project=project,
Expand Down
9 changes: 9 additions & 0 deletions loopx/extensions/lark/docs/lark-event-inbox.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,15 @@ parent sender is the app id of the configured profile. A reply to a person,
another app, or an unverifiable parent remains captured but does not wake the
agent. The agent does not need to keep a websocket open.

When Lark inbox and the same registered agent's Reward Memory are both enabled,
a non-empty registry-routed drain can also return an advisory
`reward_memory_feedback_review` hint. It asks the agent to review reusable
feedback and preview the existing scoped `reward-memory ingest-event` command;
it does not ingest chat, grant authority, or change settlement/ACK requirements.
This explicit path does not require `automatic_ingest=true`. See the
[Reward Memory inbox workflow](../../../capabilities/reward_memory/README.md#inbox-feedback-review-explicit-ingestion)
for eligibility, source verification, write/readback and default-off behavior.

### Optional turn-start Agent reading hook

Realtime collection is the preferred ingress, but a long-running Agent may also
Expand Down
Loading