The memory layer for agent orchestration. liquid-org composes a team of agents for a task, grades the result against a real test suite, and records which team worked — so good teams get retrieved and reused on similar future tasks. It's the cross-task learning loop that LangGraph, CrewAI, and Claude's own Agent Teams don't have: they compose a team per run; liquid-org remembers which one was worth keeping.
The same task, composed two ways: fresh assembly picks a 5-talent hub-and-spoke; with the evolved pool liquid-org retrieves a proven, leaner 4-talent team that scored well before. (Real
liquid-org planoutput — model-free, zero cost. Regenerate withscripts/demo/.)
/plugin marketplace add U0001F3A2/liquid-org
/plugin install liquid-org@liquid-org
Then invoke the compose-org skill on a task. Standalone CLI:
pip install "git+https://github.com/U0001F3A2/liquid-org.git"
(Python ≥3.12) → liquid-org -h (plan / compose / record).
liquid-org is the memory layer, not another orchestrator — plan (model-free)
hands you a topology + per-role personas; your runtime runs them; record
feeds the grade back. Same pattern everywhere: agent per step, wire per
edges, then record. Runnable, model-free examples for Claude Code,
LangGraph, CrewAI, AutoGen, the OpenAI Agents SDK, and any CLI consumer are in
docs/integrations/ (+ examples/).
| composes a team | grades the outcome | learns which team works across tasks | |
|---|---|---|---|
| LangGraph / CrewAI / AutoGen | ✓ (hand-defined) | — | — |
| Claude Agent Teams / SDK subagents | ✓ | — | — |
| liquid-org | ✓ | ✓ (ruff+pytest+AST+judge) | ✓ (retrieval + Thompson over graded orgs) |
No production framework learns team composition from past runs — re-verified mid-2026 across LangGraph, CrewAI, AG2, OpenAI Agents SDK, and Claude Agent Teams. That cross-task learning loop is what liquid-org adds on top of orchestration you may already use.
A team is overkill for trivial, single-step edits — a single agent is cheaper and just as good. liquid-org composes the simplest org that fits the task signature (often a plain linear chain) and escalates topology only when retrieved evidence supports it. On easy/uniform tasks the honest measured quality lift is ≈0; the real win there is fewer tokens at equal quality, and the harness is built to tell you which case you're in — a frozen frozen metric plus paired significance (Wilcoxon + bootstrap CI + MDE) and an ON−OFF cost axis, never a vanity number. See EVAL_METHODOLOGY.
New here? Start with ONBOARDING.md (module map + terminology + "how to add a new X" recipes). For per-version diffs see CHANGELOG.md.
| Layer | State |
|---|---|
| Talents (personas) | 14 seeds (7 refactor + 7 analysis); inventor proposes new ones from failure clusters |
| Orgs (topologies) | 5 seed shapes (LINEAR / HUB / DEBATE / PARALLEL_FAN_IN / ADVERSARIAL_TRIANGLE) + DAG compiler for ad-hoc shapes |
| Composer (v1.2) | CompanyTemplate-based extractor + LLM-as-template-picker over 7 hand-coded company templates |
| Mutator (v1.1) | 5 random-mutation operators on the composer's output |
| Reflector (v1.1-m3) | Post-trace topology-edit suggestions (add_edge / remove_edge / swap_edge_pattern) |
| Inventor (v1.1-m4) | Failure-cluster-driven new TalentSeed proposals |
| Persistence layer (v1.3) | Candidate buffer + lifecycle states + snapshot/rollback + judge-gated promoter |
| Eval | Multi-axis grader (ruff + pytest + AST + coverage); bug-fix grader for real-repo; LLM-as-judge (4 prompt versions, v4 default + v1–v3 replayable) |
| Real-repo arm | Pinned-SHA bootstrap + isolated git worktree grading + per-spec circuit breaker (v0.9 M6.2) |
| Continuous loop | AND-quorum gated (self_evolve_loop.py) — refuses self-mod if any external class's recency exceeds its budget |
| Security | clone_url allowlist + subprocess env scrub + proposer prompt sanitization (v0.9.1) |
| CLIs | liquid-org-rollback, liquid-org-lifecycle, liquid-org-promote (v1.3 buffer/promote/lifecycle CLIs) |
Test suite: 1,268 passing (fully hermetic — pytest -q injects fake
LLMs, no network/credentials needed).
# Setup
uv sync # install deps
cp .env.example .env # add ANTHROPIC_API_KEY or ANTHROPIC_AUTH_TOKEN
# Tests + lint (fully hermetic; no creds needed)
uv run pytest -q
uv run ruff check
# Real-Claude flows (require credentials in .env):
# 1. Calibrated dogfood (multi-trial harness on the canonical toy tasks)
uv run python scripts/tools/dogfood_calibrated.py 3
# 2. Real-repo bug-fix dogfood (M5 + M6 breaker + M6.2 per-spec)
uv run python scripts/real_repo_dogfood.py --list-tasks
# 3. Continuous self-evolution loop (AND-quorum gated)
uv run python scripts/self_evolve_loop.py --dry-run # preview
uv run python scripts/self_evolve_loop.py # for real
# 4. v1.2 composer + extractor demos
uv run python scripts/v1_2_picker_demo.py
uv run python scripts/v1_2_ablation_demo.py
# 5. v1.3 persistence layer CLIs
liquid-org-lifecycle --help # archive stale candidates, etc.
liquid-org-promote --help # gate candidate → live via judge
liquid-org-rollback --help # snapshot + restore primitiveANTHROPIC_API_KEY (sk-ant-api...) or ANTHROPIC_AUTH_TOKEN
(sk-ant-oat..., Claude Code OAuth token from macOS Keychain) in
.env or the environment is required for any real-Claude flow.
The test suite injects fake LLMCallers; pytest -q is fully
hermetic (no network).
For the full module map see ONBOARDING.md §1. Short version:
liquid_org/ importable package (~21K LOC)
talents/ talent personas + inventor (v1.1-m4)
orgs/ org graph schema + seed topologies + CompanyTemplate (24 real-world orgs)
composition/ picker + matcher + extractor + mutator + reflector + signature
runtime/ LangGraph compilers (LINEAR + HUB + DAG) + patch_apply
eval/ grader + bug_fix_grader + judge (4 prompt versions)
feedback/ recorder + promoter + candidate_buffer + significance + ablation + judge_runner
state/ lifecycle FSM + snapshot/rollback primitives (v1.3-m2 + m3)
store/ SQLite WAL engine + alembic + per-graph artifacts + FTS5
real_repo/ v0.9 real-OSS infra: spec / bootstrap / registry / dogfood
tasks/v1/ v1.0 auto-task-generator: sources + quality_classifier
cli/ v1.3 CLIs: rollback / lifecycle / promote
orchestrator.py end-to-end compose_and_run
dogfood.py toy-task per-trial logic
orchestration/self_mod.py self-modification loop (worktree → gate stack → merge)
security.py clone_url allowlist + env scrub + prompt sanitization
runtime/safety.py v0.9.2 timeout + budget cap (combined from watchdog.py + budget.py)
config/feature_flags.py env-var gates
scripts/ ~14 entry points after v1.4-m1 cleanup
references/ evaluation methodology + reference data
contracts/ 4 component contracts (schema, invariants)
tests/ 843-test pytest suite
liquid_artifacts/ gitignored; per-graph execution traces
- Talents — persistent agent personas with append-only performance history.
- Composition — given a task signature, the matcher LLM ranks prior org graphs by 4-axis similarity (decision_depth, output_modality, embedding, tool_requirements) + outcome z-score; Pareto-frontier + Thompson sampling select non-dominated + exploratory candidates. v1.2 added the CompanyTemplate-based extractor + LLM picker as an alternative composition path.
- Runtime — compiles a graph topology + talent slot assignment into a LangGraph executable; runs it; captures the proposer's diff. v1.2 added a generic DAG compiler for ad-hoc shapes.
- Eval — grader runs ruff + pytest + AST validation + coverage of changed lines; bug-fix grader additionally runs the spec's bug_repro_tests + regression subset in an isolated git worktree.
- Feedback — recorder normalizes raw axes to z-scores against task-class prior pool (with cold-start cross-class fallback), updates Thompson posteriors, writes per-trial JSONL artifacts. v1.3 added the candidate buffer + judge-gated promoter for the reflector/inventor outputs.
- State (v1.3) — pure-Python lifecycle FSM (active → stale → archived) + snapshot/rollback primitives. Runs before every promotion pass; rollback recoverable.
The system MUST run at least three distinct task streams
concurrently — self-mod, toy refactor (dogfood_calibrated.py),
and real-OSS bug-fix (real_repo_dogfood.py). The AND-quorum
guard in self_evolve_loop.py refuses to run self-mod if any
external class's recency exceeds its per-class budget
(code_refactor=24h, real_repo_bug_fix=72h — different because
real-repo trials take 10-30 min each). The per-class
--allow-stale-class CLASS override scopes the silencer per class
so one broken stream doesn't disable the invariant for all.
| Milestone | Status | One-liner |
|---|---|---|
| v1.4 cleanup | in progress | Repo hygiene — gitignore + archive + docs + renames + 4 big refactors |
| v1.3-m6.2 | deferred (post-v1.4) | Wire reflector + inventor into orchestrator so candidate buffer fills with real data |
| v1.3-m6.3 | deferred (post-v1.4) | Human-review promotion gate (user is the held-out signal until a frozen task suite exists) |
For the full risk view across v1.2 + v1.3 see design notes. For superseded design docs see design notes.
