Motivation
maildb is designed as a retrieval substrate for agent/LLM work, not a RAG system itself. Agents like Claude compose find/search/get_thread/search_all tool calls to answer questions. That's strictly more powerful than "naive RAG," and it's the right default.
But there's a gap at the low end: callers who just want one MCP call that returns an LLM-ready context block end up re-implementing the same glue every time (rank, dedupe, truncate, annotate with provenance, budget tokens). This issue proposes a single opinionated helper that packages search_all output into a prompt-ready string + structured citations.
Not a replacement for the tool-composition pattern. An additive convenience for the simple RAG shape.
Use cases
-
External-agent quickstart. A dev building an "ask my email" slackbot with the Anthropic API can call get_rag_context("meeting notes with Alice") once and stuff the result into their prompt — instead of calling search_all, looping over hits, formatting blocks, token-counting, and handling truncation.
-
LangChain / LlamaIndex adapter. These frameworks expect a Retriever interface that returns Document[]. A 20-line adapter can map get_rag_context(..., structured=True) to their type — with no maildb-internal formatting leaking into the adapter.
-
Evaluation harnesses. Reproducible RAG benchmarks need a canonical chunk-packing strategy so that accuracy deltas across runs reflect retrieval quality, not packing differences. Having the packer in the server guarantees every benchmark uses the same bytes.
-
Non-agent clients. A Jupyter notebook, a shell script, or a one-off tool that wants to drop pre-packaged context into a prompt with a single tool call. No need to consume the raw SearchResult / AttachmentSearchResult models.
-
Agentic RAG hybrid. Even Claude-via-MCP sometimes wants a clean context block rather than composing tools itself. get_rag_context("tax receipts from 2024") is ergonomic when the agent has already decided to do naive RAG for one step of a multi-step plan.
-
Cross-task reuse. The same packer serves email summarization, newsletter generation, legal-doc review, etc. — consistent output across downstream consumers.
Proposed API
MCP tool signature (adding to src/maildb/server.py):
@mcp.tool()
@log_tool
def get_rag_context(
ctx: Context,
query: str,
*,
max_tokens: int = 4000,
limit: int = 20,
include_bodies: bool = True,
# Pass-throughs to search_all:
sender: str | None = None,
sender_domain: str | None = None,
recipient: str | None = None,
after: str | None = None,
before: str | None = None,
labels: list[str] | None = None,
account: str | None = None,
direct_only: bool = False,
) -> dict[str, Any]:
"""Run semantic search over emails + attachments and return a prompt-ready
context block plus structured hits for citation.
Returns:
{
"context": str, # ready to paste into an LLM prompt
"hits": [ { ... }, ... ], # structured per-hit metadata
"token_estimate": int, # tokens in `context`
"truncated": bool, # true if any hits were cut
}
"""
Output shape
context — formatted string, one block per hit, numbered for citation:
[1] Email — alice@acme.com → you — 2024-03-15 — subject: Contract renewal
Hi — attached the revised contract. Section 4.2 on termination…
[2] Attachment — contract.pdf (PDF) — section: "Termination clause"
Either party may terminate with 30 days written notice after…
[3] Email — bob@vendor.com → you — 2024-03-18 — subject: Re: Contract renewal
Confirmed receipt. We'll have legal review by end of week.
hits — machine-readable list:
[
{
"rank": 1,
"source": "email",
"message_id": "<abc@example.com>",
"similarity": 0.78,
"snippet": "Hi — attached the revised...",
"sender": "alice@acme.com",
"date": "2024-03-15T...",
"subject": "Contract renewal",
},
{
"rank": 2,
"source": "attachment",
"attachment_id": 8421,
"similarity": 0.72,
"chunk_index": 12,
"heading_path": "Termination clause",
"filename": "contract.pdf",
},
...
]
Implementation notes
- Reuse
MailDB.search_all(query, ...) for the ranking; no new SQL.
- Use
maildb.tokenizer.count_tokens (already the HF tokenizer for nomic-embed-text) for the token budget.
- Packing algorithm: greedy by similarity order; per-hit budget is
max_tokens / limit; if a hit's snippet exceeds its slice, truncate at sentence boundary with an ellipsis; tag truncated=True when anything was cut.
- Deduplicate email hits by
message_id before packing (in case search_all surfaces both an email body and its attachment for the same message — cite once, prefer the attachment chunk for specificity).
- Respect all the existing filter params unchanged (account scoping, date ranges, recipient filters, etc.).
Non-goals
- No LLM call. maildb stays pure retrieval. The LLM step happens in the caller.
- Not a replacement for structured tools.
find(sender=..., after=...) is still the right path for "show me unreplied conversations with X last month." This helper is for the naive-RAG shape specifically.
- No re-ranking. The hit order is whatever
search_all returns. Re-ranking is a separate concern (hybrid, cross-encoder, etc.) and can layer on top.
Out of scope for this issue
- Cross-encoder re-ranking
- Multi-hop / iterative RAG
- HyDE-style query expansion
- Graph traversal
Each of those can be its own issue if they earn their keep.
Motivation
maildb is designed as a retrieval substrate for agent/LLM work, not a RAG system itself. Agents like Claude compose
find/search/get_thread/search_alltool calls to answer questions. That's strictly more powerful than "naive RAG," and it's the right default.But there's a gap at the low end: callers who just want one MCP call that returns an LLM-ready context block end up re-implementing the same glue every time (rank, dedupe, truncate, annotate with provenance, budget tokens). This issue proposes a single opinionated helper that packages
search_alloutput into a prompt-ready string + structured citations.Not a replacement for the tool-composition pattern. An additive convenience for the simple RAG shape.
Use cases
External-agent quickstart. A dev building an "ask my email" slackbot with the Anthropic API can call
get_rag_context("meeting notes with Alice")once and stuff the result into their prompt — instead of callingsearch_all, looping over hits, formatting blocks, token-counting, and handling truncation.LangChain / LlamaIndex adapter. These frameworks expect a
Retrieverinterface that returnsDocument[]. A 20-line adapter can mapget_rag_context(..., structured=True)to their type — with no maildb-internal formatting leaking into the adapter.Evaluation harnesses. Reproducible RAG benchmarks need a canonical chunk-packing strategy so that accuracy deltas across runs reflect retrieval quality, not packing differences. Having the packer in the server guarantees every benchmark uses the same bytes.
Non-agent clients. A Jupyter notebook, a shell script, or a one-off tool that wants to drop pre-packaged context into a prompt with a single tool call. No need to consume the raw
SearchResult/AttachmentSearchResultmodels.Agentic RAG hybrid. Even Claude-via-MCP sometimes wants a clean context block rather than composing tools itself.
get_rag_context("tax receipts from 2024")is ergonomic when the agent has already decided to do naive RAG for one step of a multi-step plan.Cross-task reuse. The same packer serves email summarization, newsletter generation, legal-doc review, etc. — consistent output across downstream consumers.
Proposed API
MCP tool signature (adding to
src/maildb/server.py):Output shape
context— formatted string, one block per hit, numbered for citation:hits— machine-readable list:[ { "rank": 1, "source": "email", "message_id": "<abc@example.com>", "similarity": 0.78, "snippet": "Hi — attached the revised...", "sender": "alice@acme.com", "date": "2024-03-15T...", "subject": "Contract renewal", }, { "rank": 2, "source": "attachment", "attachment_id": 8421, "similarity": 0.72, "chunk_index": 12, "heading_path": "Termination clause", "filename": "contract.pdf", }, ... ]Implementation notes
MailDB.search_all(query, ...)for the ranking; no new SQL.maildb.tokenizer.count_tokens(already the HF tokenizer for nomic-embed-text) for the token budget.max_tokens / limit; if a hit's snippet exceeds its slice, truncate at sentence boundary with an ellipsis; tagtruncated=Truewhen anything was cut.message_idbefore packing (in casesearch_allsurfaces both an email body and its attachment for the same message — cite once, prefer the attachment chunk for specificity).Non-goals
find(sender=..., after=...)is still the right path for "show me unreplied conversations with X last month." This helper is for the naive-RAG shape specifically.search_allreturns. Re-ranking is a separate concern (hybrid, cross-encoder, etc.) and can layer on top.Out of scope for this issue
Each of those can be its own issue if they earn their keep.