Repositories list
22 repositories
agentpostmortem
PublicEvery AI agent failure, documented. Public case registry.Voiceeval
PublicEvaluation for voice agents. Catches what text evals cannot see: mis-hearing, missing confirmation, latency, barge-in. Everyone can demo a voice agent; this tel…Ctxtrim
PublicTrim what bloats your AI coding context — find the files ballooning your Claude Code / Cursor / Codex token cost and write ignore files to cut it. Zero-dep. npx…VaultRAG
PublicEvalgate
PublicPrompt and agent regression CI. The build fails when your prompt gets dumber. GitHub Action with PR delta comments.MCP-audit
PublicSecurity scanner and linter for MCP servers. Audits a live server over stdio or HTTP, or a static manifest. 18 rules, SARIF output for GitHub code scanning, zer…Agentrace
PublicObservability for Claude Code subagents. Reads session transcripts, flags the results you should not trust. Checks derived from real agent failures..github
PublicCtxlens
PublicAnswerproof
PublicVerifiable, tamper-evident receipts for RAG answers. Merkle inclusion proofs and Ed25519 signatures.tokencut
PublicSkill-audit
PublicSecurity scanner for agent skills — flags prompt-injection, dangerous shell, secret access, and exfiltration before you install a Claude/agent Skill. 31 rules, …Tenantq
PublicMulti-tenant hybrid-search reference on Qdrant: tenant-isolated retrieval, dense+sparse RRF fusion, HNSW tuning, Recall@K/p95 benchmarks, batch ingestion, Docke…RelayG
PublicA support ticket triage agent built as a LangGraph state machine. LLM classification, refund policy as pure Python, and a human-in-the-loop interrupt that pause…Injection-arena
PublicCasebook-Chat
PublicA streaming AI chat UI that investigates AI-agent failures. Searches the live AgentPostmortem case registry over MCP, pulls full case files, and answers with ci…Webhands
PublicA computer-use agent for the tools that have no usable API. Drives the real dashboard via Cloudflare Browser Rendering, returns clean structured data, and refus…Greenlite
PublicMobile command and approval cockpit for AI agents. Agents escalate a proposed action with its context; you approve or deny in one tap and it routes back to the …Resolvd
PublicAn end-to-end inbox operator. Triages, drafts, and acts within policy on inbound support messages: auto-resolves order lookups and refunds under the limit, esca…Casebook-MCP
PublicA remote MCP server that turns AgentPostmortem, a public registry of documented AI-agent failures, into tools any agent can query. Ships with a companion invest…Bridgekit
PublicA scoped MCP server exposing company tools (Shopify, Triple Whale, Postgres) to an AI stack with per-client permission boundaries and an append-only audit log. …Tracecase
PublicCI for AI agents. Record agent runs, replay them against prompt and model changes, and catch regressions and unsafe tool calls before they ship. Diffs each suit…
ProTip! When viewing an organization's repositories, you can use the
props. filter to filter by custom property.