Skip to content

Repository files navigation

liquid-org

The memory layer for agent orchestration. liquid-org composes a team of agents for a task, grades the result against a real test suite, and records which team worked — so good teams get retrieved and reused on similar future tasks. It's the cross-task learning loop that LangGraph, CrewAI, and Claude's own Agent Teams don't have: they compose a team per run; liquid-org remembers which one was worth keeping.

compose-org demo

The same task, composed two ways: fresh assembly picks a 5-talent hub-and-spoke; with the evolved pool liquid-org retrieves a proven, leaner 4-talent team that scored well before. (Real liquid-org plan output — model-free, zero cost. Regenerate with scripts/demo/.)

Install (Claude Code plugin)

/plugin marketplace add U0001F3A2/liquid-org
/plugin install liquid-org@liquid-org

Then invoke the compose-org skill on a task. Standalone CLI: pip install "git+https://github.com/U0001F3A2/liquid-org.git" (Python ≥3.12) → liquid-org -h (plan / compose / record).

Use it from your framework

liquid-org is the memory layer, not another orchestrator — plan (model-free) hands you a topology + per-role personas; your runtime runs them; record feeds the grade back. Same pattern everywhere: agent per step, wire per edges, then record. Runnable, model-free examples for Claude Code, LangGraph, CrewAI, AutoGen, the OpenAI Agents SDK, and any CLI consumer are in docs/integrations/ (+ examples/).

What's different — who learns the team?

composes a team grades the outcome learns which team works across tasks
LangGraph / CrewAI / AutoGen ✓ (hand-defined) — —
Claude Agent Teams / SDK subagents ✓ — —
liquid-org ✓ ✓ (ruff+pytest+AST+judge) ✓ (retrieval + Thompson over graded orgs)

No production framework learns team composition from past runs — re-verified mid-2026 across LangGraph, CrewAI, AG2, OpenAI Agents SDK, and Claude Agent Teams. That cross-task learning loop is what liquid-org adds on top of orchestration you may already use.

When not to use it

A team is overkill for trivial, single-step edits — a single agent is cheaper and just as good. liquid-org composes the simplest org that fits the task signature (often a plain linear chain) and escalates topology only when retrieved evidence supports it. On easy/uniform tasks the honest measured quality lift is ≈0; the real win there is fewer tokens at equal quality, and the harness is built to tell you which case you're in — a frozen frozen metric plus paired significance (Wilcoxon + bootstrap CI + MDE) and an ON−OFF cost axis, never a vanity number. See EVAL_METHODOLOGY.

New here? Start with ONBOARDING.md (module map + terminology + "how to add a new X" recipes). For per-version diffs see CHANGELOG.md.

Status — v1.3-m6.1 + v1.4 cleanup in progress (2026-05-27)

Layer State
Talents (personas) 14 seeds (7 refactor + 7 analysis); inventor proposes new ones from failure clusters
Orgs (topologies) 5 seed shapes (LINEAR / HUB / DEBATE / PARALLEL_FAN_IN / ADVERSARIAL_TRIANGLE) + DAG compiler for ad-hoc shapes
Composer (v1.2) CompanyTemplate-based extractor + LLM-as-template-picker over 7 hand-coded company templates
Mutator (v1.1) 5 random-mutation operators on the composer's output
Reflector (v1.1-m3) Post-trace topology-edit suggestions (add_edge / remove_edge / swap_edge_pattern)
Inventor (v1.1-m4) Failure-cluster-driven new TalentSeed proposals
Persistence layer (v1.3) Candidate buffer + lifecycle states + snapshot/rollback + judge-gated promoter
Eval Multi-axis grader (ruff + pytest + AST + coverage); bug-fix grader for real-repo; LLM-as-judge (4 prompt versions, v4 default + v1–v3 replayable)
Real-repo arm Pinned-SHA bootstrap + isolated git worktree grading + per-spec circuit breaker (v0.9 M6.2)
Continuous loop AND-quorum gated (self_evolve_loop.py) — refuses self-mod if any external class's recency exceeds its budget
Security clone_url allowlist + subprocess env scrub + proposer prompt sanitization (v0.9.1)
CLIs liquid-org-rollback, liquid-org-lifecycle, liquid-org-promote (v1.3 buffer/promote/lifecycle CLIs)

Test suite: 1,268 passing (fully hermetic — pytest -q injects fake LLMs, no network/credentials needed).

Quickstart

# Setup
uv sync                                                    # install deps
cp .env.example .env                                       # add ANTHROPIC_API_KEY or ANTHROPIC_AUTH_TOKEN

# Tests + lint (fully hermetic; no creds needed)
uv run pytest -q
uv run ruff check

# Real-Claude flows (require credentials in .env):

# 1. Calibrated dogfood (multi-trial harness on the canonical toy tasks)
uv run python scripts/tools/dogfood_calibrated.py 3

# 2. Real-repo bug-fix dogfood (M5 + M6 breaker + M6.2 per-spec)
uv run python scripts/real_repo_dogfood.py --list-tasks

# 3. Continuous self-evolution loop (AND-quorum gated)
uv run python scripts/self_evolve_loop.py --dry-run     # preview
uv run python scripts/self_evolve_loop.py               # for real

# 4. v1.2 composer + extractor demos
uv run python scripts/v1_2_picker_demo.py
uv run python scripts/v1_2_ablation_demo.py

# 5. v1.3 persistence layer CLIs
liquid-org-lifecycle --help              # archive stale candidates, etc.
liquid-org-promote --help                # gate candidate → live via judge
liquid-org-rollback --help               # snapshot + restore primitive

ANTHROPIC_API_KEY (sk-ant-api...) or ANTHROPIC_AUTH_TOKEN (sk-ant-oat..., Claude Code OAuth token from macOS Keychain) in .env or the environment is required for any real-Claude flow. The test suite injects fake LLMCallers; pytest -q is fully hermetic (no network).

Layout

For the full module map see ONBOARDING.md §1. Short version:

liquid_org/                  importable package (~21K LOC)
  talents/                   talent personas + inventor (v1.1-m4)
  orgs/                      org graph schema + seed topologies + CompanyTemplate (24 real-world orgs)
  composition/               picker + matcher + extractor + mutator + reflector + signature
  runtime/                   LangGraph compilers (LINEAR + HUB + DAG) + patch_apply
  eval/                      grader + bug_fix_grader + judge (4 prompt versions)
  feedback/                  recorder + promoter + candidate_buffer + significance + ablation + judge_runner
  state/                     lifecycle FSM + snapshot/rollback primitives (v1.3-m2 + m3)
  store/                     SQLite WAL engine + alembic + per-graph artifacts + FTS5
  real_repo/                 v0.9 real-OSS infra: spec / bootstrap / registry / dogfood
  tasks/v1/                  v1.0 auto-task-generator: sources + quality_classifier
  cli/                       v1.3 CLIs: rollback / lifecycle / promote
  orchestrator.py            end-to-end compose_and_run
  dogfood.py                 toy-task per-trial logic
  orchestration/self_mod.py  self-modification loop (worktree → gate stack → merge)
  security.py                clone_url allowlist + env scrub + prompt sanitization
  runtime/safety.py          v0.9.2 timeout + budget cap (combined from watchdog.py + budget.py)
  config/feature_flags.py    env-var gates

scripts/                     ~14 entry points after v1.4-m1 cleanup
references/                  evaluation methodology + reference data
contracts/                   4 component contracts (schema, invariants)
tests/                       843-test pytest suite
liquid_artifacts/            gitignored; per-graph execution traces

Architecture (one sentence per layer)

  • Talents — persistent agent personas with append-only performance history.
  • Composition — given a task signature, the matcher LLM ranks prior org graphs by 4-axis similarity (decision_depth, output_modality, embedding, tool_requirements) + outcome z-score; Pareto-frontier + Thompson sampling select non-dominated + exploratory candidates. v1.2 added the CompanyTemplate-based extractor + LLM picker as an alternative composition path.
  • Runtime — compiles a graph topology + talent slot assignment into a LangGraph executable; runs it; captures the proposer's diff. v1.2 added a generic DAG compiler for ad-hoc shapes.
  • Eval — grader runs ruff + pytest + AST validation + coverage of changed lines; bug-fix grader additionally runs the spec's bug_repro_tests + regression subset in an isolated git worktree.
  • Feedback — recorder normalizes raw axes to z-scores against task-class prior pool (with cold-start cross-class fallback), updates Thompson posteriors, writes per-trial JSONL artifacts. v1.3 added the candidate buffer + judge-gated promoter for the reflector/inventor outputs.
  • State (v1.3) — pure-Python lifecycle FSM (active → stale → archived) + snapshot/rollback primitives. Runs before every promotion pass; rollback recoverable.

Anti-overfit invariant

The system MUST run at least three distinct task streams concurrently — self-mod, toy refactor (dogfood_calibrated.py), and real-OSS bug-fix (real_repo_dogfood.py). The AND-quorum guard in self_evolve_loop.py refuses to run self-mod if any external class's recency exceeds its per-class budget (code_refactor=24h, real_repo_bug_fix=72h — different because real-repo trials take 10-30 min each). The per-class --allow-stale-class CLASS override scopes the silencer per class so one broken stream doesn't disable the invariant for all.

What's next

Milestone Status One-liner
v1.4 cleanup in progress Repo hygiene — gitignore + archive + docs + renames + 4 big refactors
v1.3-m6.2 deferred (post-v1.4) Wire reflector + inventor into orchestrator so candidate buffer fills with real data
v1.3-m6.3 deferred (post-v1.4) Human-review promotion gate (user is the held-out signal until a frozen task suite exists)

For the full risk view across v1.2 + v1.3 see design notes. For superseded design docs see design notes.

Releases

Packages

Contributors

Languages