From c529bb94f13ecfc1af88ec9685fce88723f2a402 Mon Sep 17 00:00:00 2001 From: Snseam <46218267+Snseam@users.noreply.github.com> Date: Mon, 20 Jul 2026 09:20:50 +0800 Subject: [PATCH] Refresh memory radar for current source evidence Adds primary-source July 20 memory radar entries for prospective memory, memory operations, persistent-memory safety, adaptive memory control, graph correction memory, and git-bound coding-agent memory. Updates managed-memory product notes only where official sources support product behavior, with discovery-only sources kept out of performance evidence. Constraint: Weekly runbook requires accepted items to update notes, indexes, landscape, ledger, discovery log, and README counts together. Rejected: Promoting GitHub/list/catalog signals as evidence | Discovery-only sources cannot support quality or performance conclusions. Confidence: high Scope-risk: moderate Directive: Keep author/vendor-reported numbers labeled until full protocol or independent reproduction is normalized. Tested: ruby scripts/verify_memory_refresh.rb; git diff --check; conflict-marker scan; targeted canonical URL reachability. Not-tested: Full PDF/code/data/license normalization for new seed notes; OpenAI Help Center returns 403 to curl but browser fetch succeeded. --- README.md | 10 +-- README_cn.md | 10 +-- benchmarks/bad-memory-prompt-injection.md | 78 ++++++++++++++++++++++ benchmarks/claims/claims.yaml | 80 ++++++++++++++++++++++ benchmarks/index.md | 4 ++ benchmarks/memops.md | 81 +++++++++++++++++++++++ benchmarks/mempoison.md | 78 ++++++++++++++++++++++ benchmarks/pm-bench.md | 78 ++++++++++++++++++++++ docs/benchmarks-landscape.md | 14 +++- docs/memory-radar-2026-07.md | 44 +++++++++--- docs/product-discovery-log.md | 10 +++ docs/signals.md | 2 + papers/experience-memory-graph.md | 63 ++++++++++++++++++ papers/git-bound-agent-memory.md | 68 +++++++++++++++++++ papers/index.md | 9 ++- papers/memory-as-controlled-process.md | 66 ++++++++++++++++++ products/aws-agentcore-memory.md | 7 +- products/google-memory-bank.md | 11 ++- products/openai-memory.md | 17 +++-- products/oracle-ai-agent-memory.md | 19 +++++- 20 files changed, 718 insertions(+), 31 deletions(-) create mode 100644 benchmarks/bad-memory-prompt-injection.md create mode 100644 benchmarks/memops.md create mode 100644 benchmarks/mempoison.md create mode 100644 benchmarks/pm-bench.md create mode 100644 papers/experience-memory-graph.md create mode 100644 papers/git-bound-agent-memory.md create mode 100644 papers/memory-as-controlled-process.md diff --git a/README.md b/README.md index 7b89679..83f05b9 100644 --- a/README.md +++ b/README.md @@ -10,7 +10,7 @@ [![Papers](https://img.shields.io/badge/papers-989-brightgreen.svg)](papers/index.md) [![PDFs](https://img.shields.io/badge/local_PDFs-534-orange.svg)](papers/pdfs/) [![Memory products](https://img.shields.io/badge/memory%20products-38-purple.svg)](products/) -[![Benchmarks](https://img.shields.io/badge/benchmarks-21-blueviolet.svg)](benchmarks/) +[![Benchmarks](https://img.shields.io/badge/benchmarks-25-blueviolet.svg)](benchmarks/) [![Surveys](https://img.shields.io/badge/meta_surveys-6-yellow.svg)](docs/meta-surveys.md) [![Updated](https://img.shields.io/badge/updated-2026--07-lightgrey.svg)](docs/signals.md) @@ -58,10 +58,10 @@ separate buckets. | Paper index | 989 scraped papers + manual radar additions through 2026-07 | [`papers/index.md`](papers/index.md) | Searchable entry point for agent-memory papers; the 989 count is the 2026-05 scrape baseline. | | Paper stubs | 988 stubs | [`papers/stubs/`](papers/stubs/) | Track papers that are covered but not yet fully read. | | Local PDFs | 534 files | [`papers/pdfs/`](papers/pdfs/) | Re-read sources and audit paper notes. | -| Full / seed paper notes | 7 full + 15 seed | [`papers/`](papers/) | Use human-read notes for architectural decisions. | +| Full / seed paper notes | 7 full + 18 seed | [`papers/`](papers/) | Use human-read notes for architectural decisions. | | Memory product notes | 38 notes | [`products/`](products/) | Compare memory layers, memory SDKs, managed memory, and memory-enabled agents. | | Product page archives | 37 snapshots | [`products/archives/`](products/archives/) | Audit product claims after source pages change. | -| Benchmark catalog | 21 catalog rows | [`benchmarks/index.md`](benchmarks/index.md) | First-class benchmark records plus stub-backed candidate rows and a usage-claim ledger. | +| Benchmark catalog | 25 catalog rows | [`benchmarks/index.md`](benchmarks/index.md) | First-class benchmark records plus stub-backed candidate rows and a usage-claim ledger. | | Claims ledger | Structured YAML ledger | [`benchmarks/claims/claims.yaml`](benchmarks/claims/claims.yaml) | Separate vendor claims, paper evaluations, critiques, and reproductions. | | Cost-savings lane | Seed landscape | [`docs/cost-savings-landscape.md`](docs/cost-savings-landscape.md) | Find agent-memory papers, methods, code paths, and products that reduce token, latency, or runtime cost. | | Survey and taxonomy | 1 living survey + 6 meta-survey records | [`docs/agent-memory-survey.md`](docs/agent-memory-survey.md) · [`docs/meta-surveys.md`](docs/meta-surveys.md) | Build a field-level view before choosing an implementation. | @@ -180,11 +180,11 @@ flowchart LR | Path | Purpose | |---|---| -| [`papers/`](papers/) | 7 full paper notes, 15 seed notes, and the master [`index.md`](papers/index.md). | +| [`papers/`](papers/) | 7 full paper notes, 18 seed notes, and the master [`index.md`](papers/index.md). | | [`papers/stubs/`](papers/stubs/) | 988 generated stubs for papers not yet fully read. | | [`papers/pdfs/`](papers/pdfs/) | 534 archived PDFs, about 1.8 GB. See the archival policy below. | | [`papers/_scrape/`](papers/_scrape/) | Reproducibility artifacts: scrape script and dedup JSON. | -| [`benchmarks/`](benchmarks/) | 21 benchmark catalog rows, including protocol notes, stub-backed candidate rows, and the note template. | +| [`benchmarks/`](benchmarks/) | 25 benchmark catalog rows, including protocol notes, stub-backed candidate rows, and the note template. | | [`benchmarks/claims/`](benchmarks/claims/) | Usage-event ledger for benchmark mentions, vendor claims, critiques, and reproductions. | | [`benchmarks/archives/`](benchmarks/archives/) | Optional source-page snapshots for benchmark pages, repositories, or dataset cards. | | [`docs/benchmarks-landscape.md`](docs/benchmarks-landscape.md) | Benchmark landscape by capability, usage type, and evidence independence. | diff --git a/README_cn.md b/README_cn.md index 0766987..1a78def 100644 --- a/README_cn.md +++ b/README_cn.md @@ -10,7 +10,7 @@ [![Papers](https://img.shields.io/badge/papers-989-brightgreen.svg)](papers/index.md) [![PDFs](https://img.shields.io/badge/local_PDFs-534-orange.svg)](papers/pdfs/) [![记忆产品](https://img.shields.io/badge/memory%20products-38-purple.svg)](products/) -[![Benchmarks](https://img.shields.io/badge/benchmarks-21-blueviolet.svg)](benchmarks/) +[![Benchmarks](https://img.shields.io/badge/benchmarks-25-blueviolet.svg)](benchmarks/) [![Surveys](https://img.shields.io/badge/meta_surveys-6-yellow.svg)](docs/meta-surveys.md) [![Updated](https://img.shields.io/badge/updated-2026--07-lightgrey.svg)](docs/signals.md) @@ -56,10 +56,10 @@ | 论文索引 | 989 篇抓取论文 + 截至 2026-07 的手工 radar 新增 | [`papers/index.md`](papers/index.md) | 搜索 agent-memory 论文和发现线索;989 是 2026-05 抓取基线。 | | 论文 stub | 988 个 stub | [`papers/stubs/`](papers/stubs/) | 跟踪已覆盖但尚未 full 阅读的论文。 | | 本地 PDF | 534 个文件 | [`papers/pdfs/`](papers/pdfs/) | 复读来源和审计论文笔记。 | -| full / seed 论文笔记 | 7 个 full + 15 个 seed | [`papers/`](papers/) | 为架构决策引用人工阅读笔记。 | +| full / seed 论文笔记 | 7 个 full + 18 个 seed | [`papers/`](papers/) | 为架构决策引用人工阅读笔记。 | | 记忆产品笔记 | 38 个笔记 | [`products/`](products/) | 对比 memory layer、memory SDK、managed memory 和带记忆的 agent 产品。 | | 产品页面快照 | 37 个快照 | [`products/archives/`](products/archives/) | 在源页面变化后审计产品 claims。 | -| Benchmark 目录 | 21 个 catalog 行 | [`benchmarks/index.md`](benchmarks/index.md) | 理解 memory benchmark、stub-backed 候选行及其 claims 来源。 | +| Benchmark 目录 | 25 个 catalog 行 | [`benchmarks/index.md`](benchmarks/index.md) | 理解 memory benchmark、stub-backed 候选行及其 claims 来源。 | | Claims ledger | 结构化 YAML ledger | [`benchmarks/claims/claims.yaml`](benchmarks/claims/claims.yaml) | 区分厂商自报、论文评测、方法批评和独立复现。 | | 成本节省专题 | seed landscape | [`docs/cost-savings-landscape.md`](docs/cost-savings-landscape.md) | 查找能减少 token、延迟或运行成本的 agent-memory 论文、方法、代码和产品。 | | 综述与 taxonomy | 1 份活综述 + 6 条 meta-survey 记录 | [`docs/agent-memory-survey.md`](docs/agent-memory-survey.md) · [`docs/meta-surveys.md`](docs/meta-surveys.md) | 在选型或设计前建立领域视角。 | @@ -175,11 +175,11 @@ flowchart LR | 路径 | 用途 | |---|---| -| [`papers/`](papers/) | 7 个 full 论文笔记、15 个 seed 笔记和主索引 [`index.md`](papers/index.md)。 | +| [`papers/`](papers/) | 7 个 full 论文笔记、18 个 seed 笔记和主索引 [`index.md`](papers/index.md)。 | | [`papers/stubs/`](papers/stubs/) | 988 个尚未 full 阅读论文的生成 stub。 | | [`papers/pdfs/`](papers/pdfs/) | 534 个本地 PDF,约 1.8 GB。详见下方存档策略。 | | [`papers/_scrape/`](papers/_scrape/) | 可复现产物:抓取脚本和 dedup JSON。 | -| [`benchmarks/`](benchmarks/) | 21 个 benchmark catalog 行,包含协议笔记、stub-backed 候选行和笔记模板。 | +| [`benchmarks/`](benchmarks/) | 25 个 benchmark catalog 行,包含协议笔记、stub-backed 候选行和笔记模板。 | | [`benchmarks/claims/`](benchmarks/claims/) | benchmark 提及、厂商自报、方法批评和复现的 usage-event ledger。 | | [`benchmarks/archives/`](benchmarks/archives/) | benchmark 页面、repo 或 dataset card 的可选审计快照。 | | [`docs/benchmarks-landscape.md`](docs/benchmarks-landscape.md) | 按能力、使用方式和证据独立性整理的 benchmark 全景。 | diff --git a/benchmarks/bad-memory-prompt-injection.md b/benchmarks/bad-memory-prompt-injection.md new file mode 100644 index 0000000..5d0a789 --- /dev/null +++ b/benchmarks/bad-memory-prompt-injection.md @@ -0,0 +1,78 @@ +--- +title: Bad Memory — prompt-injection risk from persistent agent memory +benchmark_id: bad-memory-prompt-injection +name: Bad Memory +aliases: + - Bad Memory +status: seed +origin_type: paper_origin +origin_source: https://arxiv.org/abs/2607.14611 +first_public_date: 2026-07 +domain: memory_security +modality: text +task_grain: persistent_memory_prompt_injection +capability_axes: + - memory_file_integrity + - persistent_prompt_injection + - multi_session_attack + - memory_update_defense +data_nature: sandboxed_synthetic_workspace +metrics: + - attack_success_rate + - payload_persistence +judge_type: check +code_available: check +data_available: check +license: check +known_limitations: + - seed note; sandbox setup, systems, model versions, and released package need full read +canonical_sources: + - https://arxiv.org/abs/2607.14611 +confidence: medium +memory_modules: + - evaluator-benchmark + - policy-privacy + - ingest-adapter +last_revised: 2026-07-20 +--- + +# Bad Memory + +## What It Measures + +Bad Memory evaluates prompt-injection risks introduced by persistent memory +files, behavioral preferences, and knowledge bases in agentic systems. It is +directly relevant to coding-agent memory because the paper studies planted +payloads that can influence future sessions through durable memory state. + +## Protocol + +The arXiv abstract describes a sandboxed synthetic workspace and evaluations +over Claude Code and OpenAI Codex-style systems. It distinguishes difficulty in +getting an agent to overwrite its own memory from the risk of payloads already +present in persistent memory files. Full workspace construction, adversarial +goals, judge prompts, released package, and model-version details need a deeper +read. + +## Baselines and Reported Results + +No normalized scores are logged here. Reported attack success and payload +persistence remain paper-origin claims until the setup is mapped and compared. + +## Comparability Notes + +Compare with MemPoison and memory-poisoning trajectory-forensics work on threat +model, not raw attack success. Bad Memory is especially useful for file-backed +agent memory and coding-agent memory systems. + +## Related Papers + +- Paper:https://arxiv.org/abs/2607.14611 + +## Impact Use + +- `policy-privacy`:candidate protocol for protecting persistent memory updates + and file-backed memory state. +- `ingest-adapter`:tests whether untrusted content can cross into durable + memory. +- Ready for ImpactReport:no, upgrade after full protocol read. diff --git a/benchmarks/claims/claims.yaml b/benchmarks/claims/claims.yaml index d08a015..2ec12a2 100644 --- a/benchmarks/claims/claims.yaml +++ b/benchmarks/claims/claims.yaml @@ -155,6 +155,26 @@ events: - benchmarks/memsyco-bench.md - https://arxiv.org/abs/2607.01071 extract_confidence: medium + - event_id: pm-bench-origin-2026 + benchmark_id: pm-bench + source_kind: paper + source_id: benchmarks/pm-bench.md + source_date: 2026-07 + actor: PM-Bench authors + actor_type: paper_authors + usage_type: originates_benchmark + claim_direction: methodological_only + compared_systems: [] + reported_metrics: [] + experimental_setup: "Prospective-memory benchmark over a simulated seven-day week; local note is seed quality." + reproduction_status: not_reproduced + evidence_level: primary_pdf + independence_class: survey_only + comparability_notes: "Origin protocol only; author-reported F1 scores, agent configurations, released data, and code need full normalization before use." + evidence_refs: + - benchmarks/pm-bench.md + - https://arxiv.org/abs/2607.12385 + extract_confidence: medium - event_id: memleak-origin-2026 benchmark_id: memleak source_kind: paper @@ -195,6 +215,66 @@ events: - benchmarks/memdelta.md - https://arxiv.org/abs/2606.29914 extract_confidence: medium + - event_id: mempoison-origin-2026 + benchmark_id: mempoison + source_kind: paper + source_id: benchmarks/mempoison.md + source_date: 2026-07 + actor: MemPoison authors + actor_type: paper_authors + usage_type: originates_benchmark + claim_direction: methodological_only + compared_systems: [] + reported_metrics: [] + experimental_setup: "Persistent memory-poisoning benchmark and analysis over hand-validated cases; local note is seed quality." + reproduction_status: not_reproduced + evidence_level: primary_pdf + independence_class: survey_only + comparability_notes: "Origin protocol only; reported attack success and defense rates need full setup normalization before use." + evidence_refs: + - benchmarks/mempoison.md + - https://arxiv.org/abs/2607.14651 + extract_confidence: medium + - event_id: memops-origin-2026 + benchmark_id: memops + source_kind: paper + source_id: benchmarks/memops.md + source_date: 2026-07 + actor: MemOps authors + actor_type: paper_authors + usage_type: originates_benchmark + claim_direction: methodological_only + compared_systems: [] + reported_metrics: [] + experimental_setup: "Memory-operation benchmark over remember, forget, update, reflect, and no-op choices; local note is seed quality." + reproduction_status: not_reproduced + evidence_level: primary_pdf + independence_class: survey_only + comparability_notes: "Origin protocol only; operation traces, scoring rules, and reported model behavior need full normalization before use." + evidence_refs: + - benchmarks/memops.md + - https://arxiv.org/abs/2607.12893 + extract_confidence: medium + - event_id: bad-memory-origin-2026 + benchmark_id: bad-memory + source_kind: paper + source_id: benchmarks/bad-memory-prompt-injection.md + source_date: 2026-07 + actor: Bad Memory authors + actor_type: paper_authors + usage_type: originates_benchmark + claim_direction: methodological_only + compared_systems: [] + reported_metrics: [] + experimental_setup: "Persistent prompt-injection attacks against memory files, preferences, knowledge bases, and defenses; local note is seed quality." + reproduction_status: not_reproduced + evidence_level: primary_pdf + independence_class: survey_only + comparability_notes: "Origin protocol only; attack success rates, defense setups, and agent configurations need full normalization before use." + evidence_refs: + - benchmarks/bad-memory-prompt-injection.md + - https://arxiv.org/abs/2607.14611 + extract_confidence: medium - event_id: fact-memory-vs-long-context-baseline-2026 benchmark_id: fact-memory-vs-long-context source_kind: paper diff --git a/benchmarks/index.md b/benchmarks/index.md index f55ed6f..e76f804 100644 --- a/benchmarks/index.md +++ b/benchmarks/index.md @@ -32,6 +32,8 @@ official repository, dataset card, or independent reproduction. | LoCoMo | seed | paper-origin | very long-term conversational memory | retrieval / temporal / causal | [`locomo.md`](locomo.md) | yes | | ConvoMem | full | paper-origin | conversational memory scaling | multi-evidence / preference / abstention | [`convomem.md`](convomem.md) | yes | | MemoryAgentBench | seed | paper-origin | incremental multi-turn agent memory | retrieval / learning / long-range / forgetting | [`memoryagentbench.md`](memoryagentbench.md) | yes | +| PM-Bench | seed | paper-origin | prospective memory | delayed intention / cue monitoring / ongoing activity | [`pm-bench.md`](pm-bench.md) | yes | +| MemOps | seed | paper-origin | memory operation selection | remember / forget / update / reflect / no-op | [`memops.md`](memops.md) | yes | | GateMem | seed | paper-origin | multi-principal shared memory governance | utility / access control / active forgetting | [`gatemem.md`](gatemem.md) | backlog | | StructMemEval | seed | paper-origin | structured memory organization | structure selection / state tracking / task-specific organization | [`structmemeval.md`](structmemeval.md) | yes | | BEAM | candidate | paper-origin | million-token memory scale | long-scale recall / temporal degradation | [`beam.md`](beam.md) | yes | @@ -49,6 +51,8 @@ official repository, dataset card, or independent reproduction. | MemSyco-Bench | seed | paper-origin | memory-induced sycophancy | scope / conflict resolution / update / valid personalization | [`memsyco-bench.md`](memsyco-bench.md) | yes | | MemLeak | seed | paper-origin | multimodal deletion leakage | deletion compliance / provenance / residual image leakage | [`memleak.md`](memleak.md) | yes | | MemDelta | seed | paper-origin | memory-evaluation baseline control | component delta / model-family sensitivity / write-path cost | [`memdelta.md`](memdelta.md) | yes | +| Bad Memory | seed | paper-origin | persistent memory prompt injection | memory files / preference stores / KB poisoning / mitigation | [`bad-memory-prompt-injection.md`](bad-memory-prompt-injection.md) | yes | +| MemPoison | seed | paper-origin | persistent memory poisoning | compositional corruption / dormant trigger / defense blind spots | [`mempoison.md`](mempoison.md) | yes | ## Evidence Ledgers diff --git a/benchmarks/memops.md b/benchmarks/memops.md new file mode 100644 index 0000000..b2ccf51 --- /dev/null +++ b/benchmarks/memops.md @@ -0,0 +1,81 @@ +--- +title: MemOps — lifecycle memory operations benchmark +benchmark_id: memops +name: MemOps +aliases: + - MemOps +status: seed +origin_type: paper_origin +origin_source: https://arxiv.org/abs/2607.12893 +first_public_date: 2026-07 +domain: memory_lifecycle_operations +modality: text +task_grain: operation_level_conversational_memory +capability_axes: + - remembering + - forgetting + - updating + - reflecting + - evidence_binding + - state_transition +data_nature: generated_long_horizon_conversations +metrics: + - operation_trace_accuracy + - probe_accuracy +judge_type: check +code_available: check +data_available: check +license: check +known_limitations: + - seed note; protocol, generation pipeline, resources, and reported results need full read +canonical_sources: + - https://arxiv.org/abs/2607.12893 +confidence: medium +memory_modules: + - evaluator-benchmark + - memorydiff-generator + - retriever-reranker +last_revised: 2026-07-20 +--- + +# MemOps + +## What It Measures + +MemOps evaluates long-term conversational memory as a lifecycle of explicit +operations rather than final-answer recall alone. It focuses on remembering, +forgetting, updating, reflecting, operation composition, and whether each memory +event is tied to the right trigger, target, scope, state transition, and +supporting evidence. + +## Protocol + +The arXiv abstract describes a controllable generation pipeline that embeds +memory operations into long task-oriented conversations and produces gold +operation traces plus six categories of operation-level probes. Full probe +categories, dataset construction, released resources, and scoring details need a +deeper read. + +## Baselines and Reported Results + +No normalized scores are logged here. The abstract reports comparisons across +long-context, retrieval-based, parametric, and managed-memory systems, but those +remain author-reported results until the setup is normalized. + +## Comparability Notes + +Compare MemOps with A-TMA and MemDelta on failure decomposition, not leaderboard +rank. It is most useful when the design question is whether a memory system can +explain the operation that created, updated, forgot, or reflected on a memory. + +## Related Papers + +- Paper:https://arxiv.org/abs/2607.12893 + +## Impact Use + +- `evaluator-benchmark`:candidate protocol for operation-level memory + diagnosis. +- `memorydiff-generator`:maps naturally to explicit memory event traces and + state transitions. +- Ready for ImpactReport:no, upgrade after full protocol read. diff --git a/benchmarks/mempoison.md b/benchmarks/mempoison.md new file mode 100644 index 0000000..896dec8 --- /dev/null +++ b/benchmarks/mempoison.md @@ -0,0 +1,78 @@ +--- +title: MemPoison — persistent memory poisoning benchmark and analysis +benchmark_id: mempoison +name: MemPoison +aliases: + - MemPoison +status: seed +origin_type: paper_origin +origin_source: https://arxiv.org/abs/2607.14651 +first_public_date: 2026-07 +domain: memory_security +modality: text +task_grain: persistent_memory_poisoning +capability_axes: + - memory_poisoning + - compositional_corruption + - dormant_trigger + - write_time_defense + - retrieval_composition +data_nature: hand_validated_cases +metrics: + - attack_success_rate + - defense_effectiveness +judge_type: check +code_available: check +data_available: check +license: check +known_limitations: + - seed note; protocol, system setup, resources, and reported results need full read +canonical_sources: + - https://arxiv.org/abs/2607.14651 +confidence: medium +memory_modules: + - evaluator-benchmark + - policy-privacy + - memorydiff-generator +last_revised: 2026-07-20 +--- + +# MemPoison + +## What It Measures + +MemPoison studies persistent poisoning in external agent memory. It focuses on +adversarial records that enter through normal interaction channels, persist +across turns, and later distort behavior through retrieval composition or +trigger-conditioned activation. + +## Protocol + +The arXiv abstract describes 1,227 hand-validated cases across attack types, +injection channels, and memory substrates. It introduces L1 direct single-record +corruption, L2 compositional multi-record corruption, and L3 context-triggered +dormant corruption. Full case construction, memory substrates, defenses, and +released resources need a deeper read. + +## Baselines and Reported Results + +No normalized scores are logged here. Reported attack and defense rates remain +paper-origin claims until the setup is mapped into the claims ledger with +comparable model and memory-substrate details. + +## Comparability Notes + +Compare with memory-poisoning trajectory-forensics work only on threat model and +defense surface. MemPoison is about structural blind spots in persistent memory +defense, not a general prompt-injection leaderboard. + +## Related Papers + +- Paper:https://arxiv.org/abs/2607.14651 + +## Impact Use + +- `evaluator-benchmark`:candidate protocol for compositional and dormant + memory poisoning failures. +- `policy-privacy`:useful for write-time versus retrieval-time defense design. +- Ready for ImpactReport:no, upgrade after full protocol read. diff --git a/benchmarks/pm-bench.md b/benchmarks/pm-bench.md new file mode 100644 index 0000000..b377d68 --- /dev/null +++ b/benchmarks/pm-bench.md @@ -0,0 +1,78 @@ +--- +title: PM-Bench — prospective memory benchmark for LLM agents +benchmark_id: pm-bench +name: PM-Bench +aliases: + - PM-Bench + - Prospective Memory Bench +status: seed +origin_type: paper_origin +origin_source: https://arxiv.org/abs/2607.12385 +first_public_date: 2026-07 +domain: prospective_memory +modality: text +task_grain: delayed_intention_execution +capability_axes: + - delayed_intention + - cue_monitoring + - ongoing_activity + - state_change_detection +data_nature: simulated_week +metrics: + - f1 + - intention_execution_accuracy +judge_type: check +code_available: check +data_available: check +license: check +known_limitations: + - seed note; protocol, task counts, resources, and reported results need full read +canonical_sources: + - https://arxiv.org/abs/2607.12385 +confidence: medium +memory_modules: + - evaluator-benchmark + - retriever-reranker + - policy-privacy +last_revised: 2026-07-20 +--- + +# PM-Bench + +## What It Measures + +PM-Bench evaluates prospective memory: whether an LLM agent can remember an +intention, keep doing an ongoing activity, monitor future cues or state changes, +and execute the deferred task when it becomes due. + +## Protocol + +The arXiv abstract describes a text-based simulated seven-day week inspired by +the cognitive-science Virtual Week paradigm. Agents continue an ongoing +activity while deciding whether any delayed intention should fire. Full task +templates, released resources, and configuration details need a deeper read. + +## Baselines and Reported Results + +No normalized scores are logged here. The abstract reports a best F1 score under +one agent configuration, but that remains an author-reported result until the +setup, model versions, and evaluation scripts are checked. + +## Comparability Notes + +Compare PM-Bench with LoCoMo / LongMemEval only on long-horizon memory behavior, +not on factual recall. PM-Bench is most useful when the design question is +whether a memory layer can preserve pending commitments and state-triggered +intentions over time. + +## Related Papers + +- Paper:https://arxiv.org/abs/2607.12385 + +## Impact Use + +- `evaluator-benchmark`:candidate benchmark for pending-intention recall and + cue monitoring. +- `policy-privacy`:helps separate legitimate reminders from stale or + unauthorized remembered intentions. +- Ready for ImpactReport:no, upgrade after full protocol read. diff --git a/docs/benchmarks-landscape.md b/docs/benchmarks-landscape.md index 46d7676..9b68fd1 100644 --- a/docs/benchmarks-landscape.md +++ b/docs/benchmarks-landscape.md @@ -23,6 +23,8 @@ language: zh-CN | 长程对话记忆 | [`LoCoMo`](../benchmarks/locomo.md) | factual recall / temporal / causal / multi-session | seed,产品横评最常见 | | 对话记忆规模曲线 | [`ConvoMem`](../benchmarks/convomem.md) | long-context vs block extraction vs RAG crossover | full,同时批评 LongMemEval/LoCoMo | | agent 记忆能力维度 | [`MemoryAgentBench`](../benchmarks/memoryagentbench.md) | retrieval / test-time learning / long-range / forgetting | seed,适合定义能力轴 | +| prospective memory | [`PM-Bench`](../benchmarks/pm-bench.md) | delayed intention execution / cue monitoring / ongoing activity | seed,origin paper logged; protocol/results not normalized | +| memory-operation routing | [`MemOps`](../benchmarks/memops.md) | remember / forget / update / reflect / no-op operation selection | seed,origin paper logged; operation labels/results not normalized | | shared-memory governance | [`GateMem`](../benchmarks/gatemem.md) | utility / access control / active forgetting | seed,6 月新增治理 benchmark | | 结构化记忆组织 | [`StructMemEval`](../benchmarks/structmemeval.md) | structure selection / state tracking / task-specific organization | seed,FeishuLuo survey 查漏后新增;primary arXiv + working-paper repo | | 百万 token 规模 | [`BEAM`](../benchmarks/beam.md) | 1M/10M 长尺度记忆退化 | candidate,目前主要来自 Mem0 自报 | @@ -35,6 +37,8 @@ language: zh-CN | memory-induced sycophancy | [`MemSyco-Bench`](../benchmarks/memsyco-bench.md) | whether retrieved memory should influence factual reasoning, conflicts, updates, and personalization | seed,origin paper logged; results/resources not normalized | | multimodal deletion leakage | [`MemLeak`](../benchmarks/memleak.md) | residual recovery after deletion via correlated text and retained images | seed,origin paper logged; image/data/setup not normalized | | baseline-control methodology | [`MemDelta`](../benchmarks/memdelta.md) | component-controlled memory-vs-RAG/full-context evaluation and write-path cost discipline | seed,methodology event logged; not an end-agent leaderboard | +| persistent prompt injection | [`Bad Memory`](../benchmarks/bad-memory-prompt-injection.md) | memory-file / preference-store / KB poisoning and mitigation | seed,origin paper logged; setups/results not normalized | +| persistent memory poisoning | [`MemPoison`](../benchmarks/mempoison.md) | direct / compositional / dormant corruption and defense blind spots | seed,origin paper logged; cases/resources not normalized | ## B. 初始交叉统计 @@ -49,6 +53,8 @@ language: zh-CN | ConvoMem | 2 | origin + baseline comparison | | BEAM | 2 | Mem0 BEAM 1M / 10M vendor claims | | MemoryAgentBench | 2 | origin + survey mention | +| PM-Bench | 1 | origin paper logged; prospective-memory protocol not yet normalized | +| MemOps | 1 | origin paper logged; operation-selection protocol not yet normalized | | GateMem | 1 | origin paper logged; metrics/results not yet normalized | | StructMemEval | 1 | origin paper logged; protocol/results not yet normalized | | MemoryRewardBench | 1 | origin paper logged; evaluator benchmark only | @@ -61,6 +67,8 @@ language: zh-CN | MemSyco-Bench | 1 | origin paper logged; memory-induced sycophancy protocol not yet normalized | | MemLeak | 1 | origin paper logged; multimodal deletion-leakage protocol not yet normalized | | MemDelta | 1 | methodology event logged; use for claims discipline, not direct benchmark ranking | +| Bad Memory | 1 | origin paper logged; persistent prompt-injection setups not yet normalized | +| MemPoison | 1 | origin paper logged; persistent memory-poisoning protocol not yet normalized | ### B2. Evaluation uses / baseline comparisons @@ -160,7 +168,11 @@ enters the catalog. becomes a kernel evaluation priority. 9. Upgrade GateMem if shared-memory governance or enterprise scoped recall becomes a kernel priority. -10. Add independent reproduction rows only when the source gives enough setup +10. Upgrade PM-Bench if pending intentions, reminders, or future cue monitoring + become a kernel evaluation priority. +11. Upgrade MemPoison if persistent poisoning and compositional memory defense + become a security evaluation priority. +12. Add independent reproduction rows only when the source gives enough setup detail to distinguish reruns from marketing summaries. ## F. Maintenance Contract diff --git a/docs/memory-radar-2026-07.md b/docs/memory-radar-2026-07.md index 23bba88..6babb8c 100644 --- a/docs/memory-radar-2026-07.md +++ b/docs/memory-radar-2026-07.md @@ -1,15 +1,16 @@ --- title: 2026-07 Memory Radar refresh -date: 2026-07-06 +date: 2026-07-20 status: current-source-refresh language: zh-CN --- # 2026-07 Memory Radar refresh -本页记录 2026-07-06 的 weekly radar refresh。主 agent 从最新 `origin/main` -创建 `codex/weekly-memory-radar-2026-07-06`,并用 paper/product/GitHub discovery -子 agent 做候选检索。所有 GitHub/list/catalog 信号只作为 discovery;最终收录只依赖 +本页记录 2026-07 weekly radar refresh。主 agent 从最新 `origin/main` 创建 +`codex/weekly-memory-radar-2026-07-06` 与 +`codex/weekly-memory-radar-2026-07-20`,并用 paper/product/GitHub discovery 子 +agent 做候选检索。所有 GitHub/list/catalog 信号只作为 discovery;最终收录只依赖 primary paper/product sources。 ## 执行模型 @@ -28,6 +29,13 @@ primary paper/product sources。 | Action | Item | Why it matters | Local anchor | |---|---|---|---| +| must-add | PM-Bench | 把 prospective memory / pending intention 从普通 recall 中拆出来,测试 agent 是否能在 ongoing activity 中监控 future cue 并执行 delayed task | [`../benchmarks/pm-bench.md`](../benchmarks/pm-bench.md) | +| must-add | Memory as a Controlled Process | 将 retrieve / plan reuse / consolidate / forget 统一成 online memory-control policy,避免固定 heuristic 掩盖 memory failure | [`../papers/memory-as-controlled-process.md`](../papers/memory-as-controlled-process.md) | +| must-add | Experience Memory Graph | 用 failed vs successful trajectory graph edit path 作为可检索的纠错记忆,补充 workflow-level behavioral memory diff | [`../papers/experience-memory-graph.md`](../papers/experience-memory-graph.md) | +| must-add | MemOps | 把 remember / forget / update / reflect / no-op 抽成显式 operation selection benchmark,可用于评估 memory manager 是否知道何时写、删、改或不动 | [`../benchmarks/memops.md`](../benchmarks/memops.md) | +| must-add | Bad Memory | 把 persistent prompt injection 放到 memory files、preferences、knowledge bases 和 defense 设置中,补 coding-agent / tool-agent memory safety 维度 | [`../benchmarks/bad-memory-prompt-injection.md`](../benchmarks/bad-memory-prompt-injection.md) | +| must-add | Why Git | 将 git 作为 agentic development lifecycle 的 persistent memory substrate,补 code/workflow memory 的版本化、branch、diff 视角 | [`../papers/git-bound-agent-memory.md`](../papers/git-bound-agent-memory.md) | +| must-add | MemPoison | 记录 persistent memory poisoning 的 L1/L2/L3 分层,补 compositional multi-record corruption 与 dormant trigger 防御盲点 | [`../benchmarks/mempoison.md`](../benchmarks/mempoison.md) | | must-add | A-TMA | 把 ghost memory 拆成 bank / retrieval / answer-time state-resolution failure,强调 current/historical/transition state roles | [`../papers/atma-state-aware-memory-failures.md`](../papers/atma-state-aware-memory-failures.md) | | must-add | MemSyco-Bench | 评估 retrieved memory 什么时候不应影响事实判断,补 memory-induced sycophancy / authority 维度 | [`../benchmarks/memsyco-bench.md`](../benchmarks/memsyco-bench.md) | | must-add | MemLeak | 多模态 memory 删除后仍可从保留图片/相关文本恢复事实,补 deletion / provenance / residual leakage 维度 | [`../benchmarks/memleak.md`](../benchmarks/memleak.md) | @@ -40,6 +48,10 @@ primary paper/product sources。 | Action | Product | Why it matters | Local anchor | |---|---|---|---| +| update-existing | Google Agent Platform Memory Bank | 2026-07-15 release notes 将 memory profiles 与 IngestEvents 标为 GA,并加入 Gemini Embedding 2 similarity-search 配置 | [`../products/google-memory-bank.md`](../products/google-memory-bank.md) | +| update-existing | Oracle AI Agent Memory | 26.6 docs / product page 明确 hybrid vector+keyword search、CRUD/cascade delete、retention policy、context cards、async APIs 和 custom extraction | [`../products/oracle-ai-agent-memory.md`](../products/oracle-ai-agent-memory.md) | +| update-existing | OpenAI ChatGPT Memory | 2026-07-20 复核当前 release notes / Help Center 仍暴露 memory summary 删除、关闭、文本框编辑和高亮纠错控制 | [`../products/openai-memory.md`](../products/openai-memory.md) | +| update-existing | AWS Bedrock AgentCore Memory | 2026-07-20 复核当前 release notes 仍把 Harness built-in/BYO memory 列为 GA 能力面;未发现 post-window memory delta | [`../products/aws-agentcore-memory.md`](../products/aws-agentcore-memory.md) | | update-existing | AWS Bedrock AgentCore Memory | release notes / docs 继续确认 memory record create/update/delete streaming,支持 evented lifecycle 判断 | [`../products/aws-agentcore-memory.md`](../products/aws-agentcore-memory.md) | | update-existing | Google Agent Platform Memory Bank | 2026-06-29 release notes 将 Memory Bank generation 默认模型切到 Gemini 3.5 Flash | [`../products/google-memory-bank.md`](../products/google-memory-bank.md) | | update-existing | OpenAI ChatGPT Memory | 2026-06-25 Business / Enterprise / Edu release notes 扩展 memory summary/source/correction/delete controls | [`../products/openai-memory.md`](../products/openai-memory.md) | @@ -50,6 +62,9 @@ primary paper/product sources。 | Decision | Item | Reason | |---|---|---| +| watchlist | Speculate with Memory | memory-augmented speculator 与 latency/cost 相关,但更像 agent execution acceleration;先放 cost-savings/adjacent 队列,不升级核心 note | +| watchlist | Your Agent's Memories Are Not Its Own / FARMA | 记忆安全方向相关,但与 7 月已入库 trajectory forensics 和 MemPoison 有重叠;先等 full read 决定是否单独 seed | +| watchlist | AgenticSTS | bounded-memory testbed 和可复现实验方法有价值,但场景绑定 Slay the Spire 2;先等 code/data review 后再决定是否进 benchmark catalog | | watchlist | AutoMem | memory as cognitive skill 方向相关,但本轮未完成 project/code/data verification | | watchlist | Governed Shared Memory / Forget to Improve | post-window awesome-list commits surfaced pre-window papers;先等 full source read 后再决定是否升级 | | watchlist | Cloudflare Think harness | Cloudflare agent harness 暴露 persistent memory/context patterns,但与 Cloudflare Agent Memory 产品边界重叠 | @@ -65,6 +80,8 @@ primary paper/product sources。 | Item | Gap | Why not blocking | |---|---|---| +| PM-Bench / MemOps / Bad Memory / MemPoison | 尚未 full read PDF、code/data/license、exact task counts、released resources 和 reported results | 本轮只登记 benchmark-origin event,不登记 score claim | +| Memory as a Controlled Process / Experience Memory Graph / Why Git | 尚未 full read benchmarks、agent frameworks、code availability 与实验设置 | 本轮只登记 seed note,不登记性能结论 | | A-TMA / MemSyco-Bench / MemLeak / MemDelta / Mandol / trajectory signatures | 尚未 full read PDF、code/data/license、leaderboard 和 exact setup | 本轮只登记 seed note / benchmark-origin event,不登记 score claim | | Product updates | release notes and docs support product behavior only | 未写任何 independent benchmark 或 superiority claim | | GitHub discovery | repo activity, releases, catalog hits remain noisy | 保留 watchlist,不影响 README counts 或 core product list | @@ -73,13 +90,22 @@ primary paper/product sources。 - **State validity is overtaking static recall**:A-TMA、MemSyco-Bench、MemDelta 都要求区分 remembered fact 的 authority、scope、currentness 与 evaluation component。 +- **Prospective memory is now a benchmark axis**:PM-Bench 把 delayed intention 与 + future cue monitoring 单独抽成 agent-memory 能力,补足 reminder / commitment 场景。 +- **Memory policy is moving from heuristic to control**:Memory as a Controlled + Process 把 retrieval、plan reuse、consolidation、forgetting 视为可学习 action,提示 + kernel 需要暴露 memory-policy 实验面。 - **Memory privacy now includes multimodal residue**:MemLeak 显示删除 text memory 之后, correlated images / text 仍可能泄漏被删除事实。 +- **Persistent poisoning is becoming compositional and workspace-native**:MemPoison、 + Bad Memory 与 trajectory forensic 方向共同说明 write-time single-record filtering 不足以 + 覆盖多记录组合、workspace memory files 和 dormant trigger。 - **Runtime evidence matters for security**:trajectory-signature work把 memory poisoning 从内容检测扩展到 tool-call forensics。 -- **Managed memory products are tuning lifecycle controls**:AWS streaming、Google model - default、OpenAI org memory controls、Mem0 expiration、Redis two-tier guide 都说明 - vendor products 正在把 retention / lifecycle / state scope 做成公开产品面。 +- **Managed memory products are tuning lifecycle controls**:AWS streaming、Google + profiles/IngestEvents GA、OpenAI memory summary controls、Oracle retention/search + workflows、Mem0 expiration、Redis two-tier guide 都说明 vendor products 正在把 + retention / lifecycle / state scope 做成公开产品面。 ## Review notes @@ -87,5 +113,5 @@ primary paper/product sources。 notes / changelog / blog。 - GitHub stars、README 性能数字、MCP catalog 和 awesome-list placement 只作 discovery。 - Author-reported benchmark results are paper-origin claims;vendor numbers are vendor claims。 -- 本轮 benchmark catalog 从 18 行增至 21 行;paper seed notes 从 12 增至 15;product - note count 不变。 +- 2026-07-20 本轮 benchmark catalog 从 21 行增至 25 行;paper seed notes 从 15 增至 + 18;product note count 不变。 diff --git a/docs/product-discovery-log.md b/docs/product-discovery-log.md index 6ce0064..59eabcb 100644 --- a/docs/product-discovery-log.md +++ b/docs/product-discovery-log.md @@ -156,3 +156,13 @@ language: zh-CN | update-existing | Mem0 | changelog highlights 2026-06-27 | 更新 `products/mem0.md`;expiration controls 作为 lifecycle signal | | update-existing | Redis Agent Memory Server | Redis blog 2026-07-01 | 更新 `products/redis-agent-memory-server.md`;作为产品定位/实现建议,非 benchmark | | watchlist | Cloudflare Think harness | Cloudflare docs | 暂不拆产品;与 Cloudflare Agent Memory 有重叠,先观察是否形成独立 memory product | + +## 11. 2026-07-20 weekly refresh delta + +| Decision | Item | Source | Action | +|---|---|---|---| +| update-existing | Google Agent Platform Memory Bank | Gemini Enterprise Agent Platform release notes 2026-07-15 | 更新 `products/google-memory-bank.md`;memory profiles 与 IngestEvents GA 是托管产品行为证据 | +| update-existing | Oracle AI Agent Memory | Oracle 26.6 docs / product page | 更新 `products/oracle-ai-agent-memory.md`;hybrid search、CRUD/cascade delete、retention policy、context cards、async APIs 和 custom extraction 是产品行为证据 | +| update-existing | OpenAI ChatGPT Memory | current ChatGPT release notes / Help Center recheck | 更新 `products/openai-memory.md`;summary 删除、关闭、文本框编辑和高亮纠错属于 UX/governance 信号 | +| update-existing | AWS Bedrock AgentCore Memory | current AgentCore release notes recheck | 更新 `products/aws-agentcore-memory.md`;Harness built-in/BYO memory surface 为再确认,未发现 post-window memory delta | +| watchlist | Speculate with Memory | arXiv 2607.12236 | memory-augmented speculation 先保留为 cost/latency adjacent,不作为核心 memory 产品或 benchmark | diff --git a/docs/signals.md b/docs/signals.md index e044e5b..e9aaf3a 100644 --- a/docs/signals.md +++ b/docs/signals.md @@ -20,6 +20,8 @@ Radar 动作 enum:`stub` `seed-note` `deep-note` `impact-report` `archive-only` | 日期 | 来源 | 类型 | 一句话 | Radar 动作 | |---|---|---|---|---| +| 2026-07-20 | [PM-Bench](https://arxiv.org/abs/2607.12385) / [MemOps](https://arxiv.org/abs/2607.12893) / [Memory as a Controlled Process](https://arxiv.org/abs/2607.13591) / [Experience Memory Graph](https://arxiv.org/abs/2607.13884) / [Why Git](https://arxiv.org/abs/2607.14390) / [Bad Memory](https://arxiv.org/abs/2607.14611) / [MemPoison](https://arxiv.org/abs/2607.14651) | paper | 周更 radar 追加 prospective memory、memory-operation routing、adaptive control policy、trajectory correction、git-bound memory 和 persistent-memory safety seed | `seed-note` ✅(见 `papers/` / `benchmarks/`) | +| 2026-07-20 | [Google release notes](https://docs.cloud.google.com/gemini-enterprise-agent-platform/release-notes) / [Oracle 26.6 docs](https://docs.oracle.com/en/database/oracle/agent-memory/26.6/guide/whats-new.html) / [OpenAI release notes](https://help.openai.com/en/articles/6825453-chatgpt-release-notes) / [AWS AgentCore release notes](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/release-notes.html) | product | 官方产品源补充 Memory Bank profiles/IngestEvents GA、Oracle hybrid search/retention/workflows、ChatGPT memory summary controls,并复核 AWS Harness memory surface | `deep-note` ✅(更新 `products/`) | | 2026-07-06 | [A-TMA](https://arxiv.org/abs/2607.01935) / [Mandol](https://arxiv.org/abs/2606.29778) / [Forensic Trajectory Signatures](https://arxiv.org/abs/2606.30566) | paper | 周更 radar 追加 state-aware ghost memory、agglomerative memory-native storage、memory-poisoning trajectory forensics 三条 seed note | `seed-note` ✅(见 `papers/`) | | 2026-07-06 | [MemSyco-Bench](https://arxiv.org/abs/2607.01071) / [MemLeak](https://arxiv.org/abs/2606.29788) / [MemDelta](https://arxiv.org/abs/2606.29914) | paper | benchmark catalog 新增 memory-induced sycophancy、多模态删除泄漏、controlled baseline methodology | `seed-note` ✅(见 `benchmarks/`) | | 2026-07-06 | [AWS AgentCore release notes](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/release-notes.html) / [Google release notes](https://docs.cloud.google.com/gemini-enterprise-agent-platform/release-notes) / [OpenAI Business release notes](https://help.openai.com/en/articles/11391654-chatgpt-business-release-notes) / [Mem0 changelog](https://docs.mem0.ai/changelog/highlights) / [Redis guide](https://redis.io/blog/build-smarter-ai-agents-manage-short-term-and-long-term-memory-with-redis/) | product | 官方产品源补充 AgentCore streaming、Memory Bank model default、ChatGPT org memory controls、Mem0 expiration 和 Redis two-tier guidance | `deep-note` ✅(更新 `products/`) | diff --git a/papers/experience-memory-graph.md b/papers/experience-memory-graph.md new file mode 100644 index 0000000..18de494 --- /dev/null +++ b/papers/experience-memory-graph.md @@ -0,0 +1,63 @@ +--- +title: "Experience Memory Graph: One-Shot Error Correction for Agents" +arxiv_id: 2607.13884 +source: arXiv:2607.13884 +date: 2026-07 +domain: memory +core_claim: | + Failed and successful agent trajectories can be converted into an experience + memory graph that retrieves graph-edit corrections for one-shot recovery. +evidence_level: medium +code_available: check +license: check +memory_modules: + - ingest-adapter + - retriever-reranker + - memorydiff-generator + - evaluator-benchmark +status: seed +last_revised: 2026-07-20 +urls: + - https://arxiv.org/abs/2607.13884 +--- + +# Experience Memory Graph(arXiv 2607.13884) + +## Problem statement + +Reflection-style self-correction often requires repeated test-time trials and +stores task-specific lessons that may not transfer. Long-horizon agents need a +memory format that can encode how failed trajectories differ from successful +ones and retrieve a corrective action pattern without a new reflection loop. + +## Core claim + +Experience Memory Graph (EMG) converts failed exploration trajectories and +successful expert trajectories into directed action-decision graphs. It stores +common successful subgraphs and graph edit paths as memory, then retrieves the +relevant correction at test time. The seed note treats EMG as trajectory-memory +evidence, not as an independent benchmark claim. + +## Decision relevance + +- `memorydiff-generator`:graph edit paths are a concrete form of behavioral + memory diff between failed and successful trajectories. +- `retriever-reranker`:retrieval over action-decision subgraphs is a stronger + analogy for agent workflows than pure text similarity. +- `evaluator-benchmark`:ALFWorld / ScienceWorld results are useful only after + the exact setup and baselines are checked. + +## Caveats + +本地笔记是 seed 质量。需要 full read 后确认 graph construction, training-time expert +trajectory assumptions, transfer scope, code/data availability, and whether the +method is memory-native or mainly a reflection replacement. + +## Sources + +- arXiv:https://arxiv.org/abs/2607.13884 + +--- + +> *Ymem 项目对本笔记决策相关性的具体绑定见 +> [`../docs/ymem-binding/relevance-index.md`](../docs/ymem-binding/relevance-index.md)。* diff --git a/papers/git-bound-agent-memory.md b/papers/git-bound-agent-memory.md new file mode 100644 index 0000000..35b7ee6 --- /dev/null +++ b/papers/git-bound-agent-memory.md @@ -0,0 +1,68 @@ +--- +title: "Why Git Is the Memory Solution for the Agentic Development Lifecycle" +arxiv_id: 2607.14390 +source: arXiv:2607.14390 +date: 2026-07 +domain: coding-agent-memory +core_claim: | + Coding-agent memory should be bound to repository lifecycle artifacts such as + commits, branches, review, and merge rather than treated as a separate + transcript-retrieval store. +evidence_level: medium +code_available: yes +license: check +memory_modules: + - ingest-adapter + - retriever-reranker + - memorydiff-generator + - evaluator-benchmark +status: seed +last_revised: 2026-07-20 +urls: + - https://arxiv.org/abs/2607.14390 + - https://github.com/rekal-dev/rekal-cli +--- + +# Why Git Is the Memory Solution for the Agentic Development Lifecycle(arXiv 2607.14390) + +## Problem statement + +Coding-agent decisions often live in assistant transcripts that disappear from +the repository lifecycle. The paper argues that memory for the agentic +development lifecycle should inherit git's ground truth, freshness, review, and +merge boundaries instead of relying only on a separate memory graph or transcript +retrieval layer. + +## Core claim + +The paper proposes git-bound, routed memory: structural map lookups for broad +questions, confidence-gated episodes for pointed questions, and decision +synthesis for rationale. It reports retrieval and answer-sufficiency results on +developer histories, but this seed note records the architecture and evaluation +pressure only. The paper is product-adjacent to Rekal, so all numbers remain +author-reported evidence until independently reproduced. + +## Decision relevance + +- `ingest-adapter`:commit-session links and repository artifacts can be more + reliable memory boundaries than raw transcript ingestion. +- `memorydiff-generator`:git diffs and reviews provide a native provenance + model for coding-agent memory. +- `retriever-reranker`:routing between structure, episodes, and rationale helps + avoid over-injecting transcript memories. + +## Caveats + +本地笔记是 seed 质量。需要 full read 后确认 corpus construction, pre-registration +details, code license, Rekal product boundary, and whether the evaluation can be +replicated outside the author's systems. + +## Sources + +- arXiv:https://arxiv.org/abs/2607.14390 +- Code:https://github.com/rekal-dev/rekal-cli + +--- + +> *Ymem 项目对本笔记决策相关性的具体绑定见 +> [`../docs/ymem-binding/relevance-index.md`](../docs/ymem-binding/relevance-index.md)。* diff --git a/papers/index.md b/papers/index.md index 3a0a8c8..d4ff6e7 100644 --- a/papers/index.md +++ b/papers/index.md @@ -21,9 +21,16 @@ removed. ## 2026-07 manual radar additions -These entries were added by the 2026-07-06 weekly radar refresh. They are not +These entries were added by the 2026-07 weekly radar refreshes. They are not part of the 2026-05-19 nine-list scrape statistics above. +- [PM-Bench: Evaluating Prospective Memory in LLM Agents](../benchmarks/pm-bench.md) — 2026-07 — benchmark seed — [arxiv](https://arxiv.org/abs/2607.12385) +- [Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents](memory-as-controlled-process.md) — 2026-07 — seed — [arxiv](https://arxiv.org/abs/2607.13591) +- [Experience Memory Graph: One-Shot Error Correction for Agents](experience-memory-graph.md) — 2026-07 — seed — [arxiv](https://arxiv.org/abs/2607.13884) +- [MemOps: Memory Operations Benchmark for Memory-Augmented Agents](../benchmarks/memops.md) — 2026-07 — benchmark seed — [arxiv](https://arxiv.org/abs/2607.12893) +- [Bad Memory: Benchmarking and Mitigating Persistent Memory Prompt Injection Attacks](../benchmarks/bad-memory-prompt-injection.md) — 2026-07 — benchmark seed — [arxiv](https://arxiv.org/abs/2607.14611) +- [Why Git: Git as Agentic Memory Layer](git-bound-agent-memory.md) — 2026-07 — seed — [arxiv](https://arxiv.org/abs/2607.14390) +- [MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents](../benchmarks/mempoison.md) — 2026-07 — benchmark seed — [arxiv](https://arxiv.org/abs/2607.14651) - [A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory](atma-state-aware-memory-failures.md) — 2026-07 — seed — [arxiv](https://arxiv.org/abs/2607.01935) - [MemSyco-Bench: Benchmarking Sycophancy in Agent Memory](../benchmarks/memsyco-bench.md) — 2026-07 — benchmark seed — [arxiv](https://arxiv.org/abs/2607.01071) - [MemLeak: Diagnosing Information Leaks in Multimodal Agent Memory](../benchmarks/memleak.md) — 2026-06 — benchmark seed — [arxiv](https://arxiv.org/abs/2606.29788) diff --git a/papers/memory-as-controlled-process.md b/papers/memory-as-controlled-process.md new file mode 100644 index 0000000..5f2bad3 --- /dev/null +++ b/papers/memory-as-controlled-process.md @@ -0,0 +1,66 @@ +--- +title: "Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents" +arxiv_id: 2607.13591 +source: arXiv:2607.13591 +date: 2026-07 +domain: memory +core_claim: | + Agent memory operations should be selected by an online control policy rather + than fixed retrieval, consolidation, and forgetting heuristics. +evidence_level: medium +code_available: check +license: check +memory_modules: + - ingest-adapter + - retriever-reranker + - dream-consolidator + - evaluator-benchmark +status: seed +last_revised: 2026-07-20 +urls: + - https://arxiv.org/abs/2607.13591 +--- + +# Memory as a Controlled Process(arXiv 2607.13591) + +## Problem statement + +Many agent-memory systems expose a fixed memory policy: retrieve by a static +heuristic, inject a fixed number of memories, consolidate on a fixed cadence, or +forget with hand-written rules. The paper argues that this is brittle because +the right memory operation depends on task stage, goal recurrence, stuck-state +signals, and long-run store quality. + +## Core claim + +MemCon models memory management as an online control problem. It wraps existing +memory backends and learns when to retrieve, what to retrieve, how much to +inject, when to reuse distilled plans, and when to consolidate or forget. The +arXiv abstract reports cross-benchmark gains and token reductions, but this +seed note records the control-policy framing only; reported scores remain +author claims until the setup is normalized. + +## Decision relevance + +- `retriever-reranker`:retrieval should be policy-selected from state, not just + nearest-neighbor lookup. +- `dream-consolidator`:consolidation and pruning can be treated as controllable + actions with feedback, not offline cleanup only. +- `evaluator-benchmark`:memory experiments should expose the policy that + decides when to read/write/forget, because fixed policies may hide failure + modes. + +## Caveats + +本地笔记是 seed 质量。需要 full read 后确认 benchmarks, agent frameworks, feedback +signal, code availability, and whether the reported token savings are comparable +to existing cost-savings notes. + +## Sources + +- arXiv:https://arxiv.org/abs/2607.13591 + +--- + +> *Ymem 项目对本笔记决策相关性的具体绑定见 +> [`../docs/ymem-binding/relevance-index.md`](../docs/ymem-binding/relevance-index.md)。* diff --git a/products/aws-agentcore-memory.md b/products/aws-agentcore-memory.md index e615fd9..6c85463 100644 --- a/products/aws-agentcore-memory.md +++ b/products/aws-agentcore-memory.md @@ -12,7 +12,7 @@ memory_modules: - retriever-reranker - policy-privacy status: seed -last_revised: 2026-07-06 +last_revised: 2026-07-20 archive: archives/aws-agentcore-memory-overview.md --- @@ -58,6 +58,11 @@ record streaming 作为当前能力面:memory record create / update / delete Kinesis,用于下游审计、同步或增量处理。该能力是产品行为证据,可支持"AgentCore Memory 暴露 evented memory lifecycle"这一判断,但不支持任何独立性能结论。 +2026-07-20 复核时,当前 AgentCore release notes 仍把 Harness GA 列在 2026-06 +section,并写明 GA harness 支持 built-in memory by default 或 bring-your-own +memory。本轮没有发现新的 post-2026-07-13 memory capability delta;这里只把它作为 +当前官方 release surface 的再确认,不登记为独立性能或质量证据。 + ## 4. 决策相关性 / Decision relevance - **对照点**:AWS 把 memory 当成 agent runtime 的可配置云资源,说明 hyperscaler diff --git a/products/google-memory-bank.md b/products/google-memory-bank.md index a8eb519..21c0c53 100644 --- a/products/google-memory-bank.md +++ b/products/google-memory-bank.md @@ -12,7 +12,7 @@ memory_modules: - retriever-reranker - policy-privacy status: seed -last_revised: 2026-07-06 +last_revised: 2026-07-20 archive: archives/google-memory-bank-overview.md --- @@ -55,6 +55,15 @@ generation 的默认模型从 Gemini 2.5 Flash 改为 Gemini 3.5 Flash。该更 memory extraction / generation 仍会随平台模型配置变化;它是产品行为证据,不代表 memory ranking 机制或质量有独立复现。 +2026-07-15 release notes 又把 Memory Bank memory profiles 标为 GA。Memory +profiles 用固定 schema 生成和更新结构化 profile,让 agent 在 session 中不必每次 +做昂贵搜索也能读取 evolving information。同日 Memory Bank 支持 Gemini Embedding 2 +的 similarity-search configuration,并把 `IngestEvents` API 标为 GA;该 API 将事件 +ingestion 与 memory generation 解耦,支持 continuous stream、generation window +overlap、revision TTL/labels、以及 memory metadata merge。该更新加强了 Google +Memory Bank 的 structured profile + evented ingestion 形态,但仍是 Google 托管产品 +行为证据。 + ## 4. 决策相关性 / Decision relevance - **对照点**:Memory Bank 是 hyperscaler 级 scoped memory store 的典型样本。 diff --git a/products/openai-memory.md b/products/openai-memory.md index af42495..a6de4aa 100644 --- a/products/openai-memory.md +++ b/products/openai-memory.md @@ -10,7 +10,7 @@ memory_modules: - ingest-adapter - retriever-reranker status: full -last_revised: 2026-07-06 +last_revised: 2026-07-20 archive: archives/openai-memory-overview.md --- @@ -36,8 +36,9 @@ Memory 中查看 / 编辑 / 删除任意一条,或整体关闭;"Temporary Chat" "long-term memory layer",更像 stateful conversation store;跨 thread 的 语义记忆需要开发者自己在 vector store + retrieval 上搭。 -> 注:本次抓取时 OpenAI 帮助中心与 platform docs 均返回 403,以上为公开 -> 信息梳理。详见 archive 文件中的限制说明。 +> 注:OpenAI 帮助中心对 shell `curl` 抓取有 403 访问限制;2026-07-20 browser +> fetch 可访问当前 ChatGPT release notes。platform docs 仍可能随登录/网络策略变化。 +> 详见 archive 文件中的限制说明。 ## 3. 关键技术选择 @@ -65,6 +66,12 @@ vendor-managed memory 的 UX/治理升级,不是新的开放 memory kernel。 memory 不受该 ChatGPT 产品 memory 变更影响。该更新继续归类为 vendor-managed product behavior,不作为开放 agent-memory API 或独立质量证据。 +2026-07-20 复核时,当前 ChatGPT release notes / Help Center 继续暴露 memory +summary 的用户控制面:用户可从 summary 页面删除显示的 memories,也可用 "Delete +and turn off memory" 关闭 memory;summary 支持文本框编辑和高亮纠错,但关闭 memory +不会删除既有历史聊天。该更新只按当前官方帮助中心产品行为记录,不代表底层 schema、 +检索或 consolidation 机制公开。 + ## 4. 决策相关性 / Decision relevance - **对照点**:它定义了 "vendor-managed memory" 的对照基线 — 用户/开发者 @@ -96,8 +103,8 @@ product behavior,不作为开放 agent-memory API 或独立质量证据。 - **不透明**:写入策略、保留时长、整理逻辑均未公开,出问题难调试 - **行为变更风险**:产品 / 政策迭代会改变默认行为(2024 → 2025 已经发生 过一次扩展),依赖方需要持续跟踪 -- **抓取限制**:OpenAI 帮助中心 / platform docs 在本次抓取返回 403,以上 - 事实需要以官方页面为准 +- **抓取限制**:OpenAI 帮助中心对 shell `curl` 可能返回 403;本轮用 browser fetch + 复核 release notes,但事实仍需要以官方页面为准 ## 7. 进一步阅读 diff --git a/products/oracle-ai-agent-memory.md b/products/oracle-ai-agent-memory.md index 7f9c00d..9ece6c5 100644 --- a/products/oracle-ai-agent-memory.md +++ b/products/oracle-ai-agent-memory.md @@ -11,7 +11,7 @@ memory_modules: - retriever-reranker - policy-privacy status: seed -last_revised: 2026-06-29 +last_revised: 2026-07-20 archive: archives/oracle-ai-agent-memory-overview.md --- @@ -43,6 +43,16 @@ memory(`add` / `search` workflows),用于保存用户偏好、规则和跨会话 官方文档集存在。Oracle developer blog 的 Claude / Oracle / LangChain 组合文章在 本环境返回 403,因此不把该 blog 的架构定位升级为本仓强证据;后续可人工复核后再补。 +## 3.2 2026-07 refresh + +2026-07-20 复核时,Oracle 26.6 文档和官方产品页可访问。官方资料继续把 +AI Agent Memory 定位为 Oracle AI Database 上的 enterprise persistent memory +layer,并新增/强调 hybrid vector + keyword search、`add`/`search`/`delete`/`update` +workflow、record cascade delete、schema-level retention policy、chunked indexing、 +background/inline extraction、thread context cards、async APIs、metadata filtering +以及 custom extraction function。它们是 Oracle 产品行为和架构证据,不支持独立 +benchmark superiority claim。 + ## 4. 决策相关性 / Decision relevance - **对照点**:Oracle 代表 "enterprise database becomes memory substrate" 路线。 @@ -65,8 +75,11 @@ memory(`add` / `search` workflows),用于保存用户偏好、规则和跨会话 ## 7. 进一步阅读 - archive: [`archives/oracle-ai-agent-memory-overview.md`](archives/oracle-ai-agent-memory-overview.md) -- Docs:https://docs.oracle.com/en/database/oracle/agent-memory/26.4/agmea/about.html -- Docs index:https://docs.oracle.com/en/database/oracle/agent-memory/26.4/agmea/index.html +- Docs 26.4:https://docs.oracle.com/en/database/oracle/agent-memory/26.4/agmea/about.html +- Docs 26.6 index:https://docs.oracle.com/en/database/oracle/agent-memory/26.6/index.html +- What's new 26.6:https://docs.oracle.com/en/database/oracle/agent-memory/26.6/guide/whats-new.html +- API reference:https://docs.oracle.com/en/database/oracle/agent-memory/26.6/guide/api/agentmemory.html +- Product page:https://www.oracle.com/database/ai-agent-memory/ ---