Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolio
Paper · Dataset · MA-Retriever · Generation Results · Evaluation Results · Baseline Settings
A meta-analysis statistically combines results from independent studies that answer the same research question. It is usually conducted within a systematic review, where researchers retrieve candidate studies and screen them against predefined eligibility criteria. These criteria are often organized as PI/ECO: Population, Intervention or Exposure, Comparison, and Outcome.
The workflow links several distinct decisions: retrieving candidate articles, selecting evidence under the review protocol, extracting and analyzing study results, and reporting the final synthesis. Published protocols, included-study lists, and conclusions make retrieval, selection, and written synthesis especially suitable for reproducible evaluation.
MetaSyn turns 422 Nature Portfolio source reviews, mainly systematic reviews and meta-analyses, into linked benchmark tasks. Each task provides a research question, structured PI/ECO fields, search and eligibility information, the published included-study list, and reference synthesis labels. Stable article IDs connect the included-study lists to one shared corpus of 140,585 PubMed records.
The benchmark evaluates retrieval, protocol-based evidence selection, and report synthesis. This repository provides the benchmark code, evaluator prompts, and outputs for the 12 main configurations.
| Statistic | Value |
|---|---|
| Source reviews | 422 |
| Train / test reviews | 336 / 86 |
| Shared PubMed corpus | 140,585 articles |
| Review/article links | 7,374 |
| PI/ECO structure available | 100% |
The source-review split is disjoint. Included-study titles are linked to the shared corpus by stable IDs, and the complete title lists are retained for future matching through other literature databases.
We evaluate sparse and dense retrieval, one-pass RAG, explicit screening, and deep-research agents. MA-Retriever reaches 91.7% Recall@200, while the highest end-to-end inclusion recall under standard retrieval is 51.2%. Oracle and candidate-level diagnostics show that systems still omit many reference articles after retrieval, which makes protocol-based selection a separate and important evaluation target.
| Path | Contents |
|---|---|
src/metasyn |
Retrieval, RAG, agent adapters, and evaluation library |
scripts |
Runnable indexing, retrieval, generation, evaluation, and training commands |
configs |
Settings for the reported baseline configurations |
generation_results |
Reports and compact result records for all 12 configurations |
evaluation_results |
Per-review metrics for all 12 configurations |
tests |
Public interface and evaluator contract tests |
The release is validated with Python 3.11 and supports Python 3.10 and 3.11.
python -m venv .venv
source .venv/bin/activate
pip install -e .
cp .env.example .envSet an OpenAI-compatible endpoint in .env:
OPENAI_API_KEY=
OPENAI_BASE_URL=https://api.openai.com/v1
OPENAI_MODEL=Set OPENAI_API_KEY and OPENAI_MODEL for the endpoint you use.
Every LLM call uses the standard Python client:
from openai import OpenAI
client = OpenAI(api_key=api_key, base_url=base_url)
response = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": prompt}],
)python scripts/build_index.py --output artifacts/ma_retriever_indexThe command downloads the public corpus and MA-Retriever, encodes article titles and abstracts, and builds the FAISS HNSW index used by the benchmark.
python scripts/run_retrieval.py \
--retriever ma_retriever \
--review-id 63 \
--index artifacts/ma_retriever_index \
--top-k 200 \
--output outputs/retrieval_63.jsonThe retrieval query contains the research question and PI/ECO fields. It does not contain the source-review title. Any corpus record corresponding to the source review is removed before top-K truncation.
python scripts/run_rag.py \
--review-id 63 \
--index artifacts/ma_retriever_index \
--top-k 200 \
--mode actual \
--output outputs/rag_63Use --mode oracle to place all linked included articles first. Reference
articles are ordered by their MA-Retriever score for the task query, and
ordinary candidates fill the remaining positions in their normal retrieval
order. This is a label-informed, non-deployable diagnostic that uses the
benchmark reference set. The labels affect retrieval order but are not shown
to the LLM.
To reproduce an actual-retrieval BM25 RAG row, omit --index and run:
python scripts/run_rag.py \
--review-id 63 \
--retriever bm25 \
--top-k 200 \
--mode actual \
--output outputs/rag_bm25_63For raw BGE, build its index with scripts/build_index.py, then pass
--retriever dense_raw and that index. The raw BGE model is selected by
default; --retriever-model can override it.
Each run writes report.md and results.json. The latter records the review ID,
retrieved articles, parsed final included-article list, unmatched stated citations,
model name, and API token usage when the provider returns it.
By default, the evaluator reads report.md and parses its dedicated final
included-article list against the retrieved_articles stored in results.json.
Exact candidate titles and explicit Corpus ID values are accepted. An
unrecognized final list is an evaluation error rather than an empty prediction.
ID-based metrics do not require an API:
python scripts/evaluate.py \
--result outputs/rag_63/results.json \
--report outputs/rag_63/report.md \
--output outputs/rag_63/evaluation.json \
--id-onlyOmit --id-only to also compute the evaluator-dependent criteria, direction,
insight, and structure metrics. These calls use the same OpenAI-compatible
environment variables shown above. The benchmark reports inclusion recall,
precision, F1, screening accuracy, inclusion/exclusion criteria consistency,
conclusion-direction accuracy, insight consistency, and structure quality.
Unmatched IDs stated in the final included-article list remain predicted
inclusions and therefore count against inclusion precision. The evaluator
deduplicates these entries and derives their count from the parsed list.
The same ID-only evaluation also reports four retrieval diagnostics. Let G
be the linked reference set, P the deduplicated retrieved pool, and L the
final mapped inclusion list:
retrieval_recallislen(P & G) / len(G).retrieval_precisionislen(P & G) / len(P).conditional_retentionislen(P & G & L) / len(P & G).post_retrieval_lossislen((P & G) - L) / len(G).
Conditional retention is null when no reference article was retrieved. It is
not the complement of post-retrieval loss because the two metrics use different
denominators. The nine main pipeline metrics are unchanged; these four values
make retrieval and downstream loss directly reproducible from the same result.
Systems that already resolve their own final citations can bypass report-list
parsing with --selection-source json. Their results.json must use this exact
schema:
{
"review_id": 63,
"retrieved_article_ids": [101, 102, 103],
"included_article_ids": [101, 103],
"unmatched_included_article_ids": [],
"unmatched_included_entries": []
}included_article_ids must be a subset of retrieved_article_ids. Put an
explicit but unmapped numeric ID in unmatched_included_article_ids, and put an
unmapped title or other citation string in unmatched_included_entries. Do not
supply unmatched_included_article_count; that field is rejected as input.
The evaluator deduplicates both lists, derives the precision denominator, and
reports unmatched_included_article_count as the derived number of unique
unmapped IDs and citation strings. A full evaluator-dependent run still
needs report.md for the criteria and synthesis metrics:
python scripts/evaluate.py \
--result outputs/my_system/review_63/results.json \
--report outputs/my_system/review_63/report.md \
--selection-source json \
--judge-model gpt-5.5 \
--output outputs/my_system/review_63/evaluation.jsonEach evaluation records its UTC timestamp, evaluator and prompt versions,
dataset ID and revision, judge model, report hash, and API token usage. Exhausted
API retries or invalid judge responses produce status: failed and null text
metrics. They are never converted into low system scores.
Place exactly one results.json and, unless using --id-only with strict JSON,
one report.md under a directory for each of the 86 test reviews:
outputs/my_system/
review_23/results.json
review_23/report.md
...
review_466/results.json
review_466/report.md
Then run:
python scripts/evaluate_all.py \
--results-dir outputs/my_system \
--selection-source report \
--judge-model gpt-5.5 \
--output outputs/my_system/evaluation.jsonThe command checks that every released test review appears exactly once before making evaluator calls. It writes all per-review records, macro averages, and the valid denominator for each metric. Any missing review or failed evaluator call makes the run fail and suppresses the aggregate score.
python scripts/train_retriever.py \
--base-model BAAI/bge-large-en-v1.5 \
--output artifacts/ma_retrieverThe training script uses only training-review links and removes every article linked to a test review from the positive training pairs. The released model was trained for one epoch with Multiple Negatives Ranking Loss.
AgentCorpusTools provides the task-scoped search and fetch interface used by
the agent experiments. It excludes source-review records from every search and
allows an article to be fetched only after that task has returned it:
from metasyn.agent_tools import AgentCorpusTools, build_agent_prompt
tools = AgentCorpusTools(retriever, review, k=20, max_distinct_articles=200)
search_results = tools.search("focused protocol query")
article = tools.fetch(search_results[0]["corpus_id"])
prompt = build_agent_prompt(review)This interface can be connected to GPT-Researcher, Open Deep Research, or another agent without copying those third-party projects into this repository. The experiments used GPT-Researcher 0.14.7 and Open Deep Research 0.0.16. Evaluate the explicit final article list with the same result schema.
See BASELINES.md and configs/final_experiments.json for the
exact BM25, raw-BGE, ProtoMA, GPT-Researcher, and OpenDR settings.
generation_results contains the original
generated reports plus compact per-review result records for the nine one-pass
RAG configurations, ProtoMA, GPT-Researcher, and OpenDR. All 12 configurations
cover the same 86 test reviews, for 1,032 report/result pairs.
The compact records omit full retrieved-article objects, so reevaluate them with
--selection-source json. Newly generated run_rag.py outputs also support
report-based list parsing.
evaluation_results contains the corresponding
1,032 per-review metric records and evaluator-extracted criteria, conclusion
directions, and key insights. The generated reports are included unchanged.
Code and project-authored annotations are MIT licensed. PubMed metadata, abstracts, article identifiers, source-review excerpts, and PMC-derived text retain their upstream terms; the MIT license does not relicense third-party content. See the dataset card for field-level details.
@misc{metasyn2026,
title = {Benchmarking {LLM} Agents on Meta-Analysis Articles from {Nature} Portfolio},
author = {Anzhe Xie and Weihang Su and Yujia Zhou and Yiqun Liu and Min Zhang and Qingyao Ai},
year = {2026},
eprint = {2606.17041},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2606.17041}
}