Graceful degradation for concurrent LLM/agent terminals. When system resources get hairy, sessions back off voluntarily and in priority order — before the OS OOM-kills everything at once — so the work that matters survives and a returning human comes back to an autopsy, not a smoking crater.
Run several agent terminals on one box and the failure mode is brutal:
nothing backs off, memory craters, and the OS kills the whole VM — every
session dies at once, ungracefully, mid-thought. llm-grace makes
that failure legible, bounded, and recoverable.
A tiny monitor reads cheap signals the kernel already emits (/proc
counters — memory, swap-out rate, iowait, disk saturation, loadavg) and
reads them as signatures, not thresholds: the classic “disk-thrash
low CPU” pattern is the pre-OOM tell. Sessions coordinate through one
shared crash-safe LMDB ledger. Under pressure, the least valuable work
yields first — value measured as motion, not size, so a thrashing
runaway is shed before productive work — climbing a per-session ladder:
silent PAUSE (auto-resumes) → CHECKPOINT (a human-in-the-loop latch,
loud) → clean TERMINATE (state already saved, never auto-relaunched).
Every transition is timestamped; terminated sessions get a per-session
autopsy, and the whole episode aggregates into one timeline so you can
see how it went wobbly.
The autopsy and death-spiral timeline are fallbacks you hope never to read, not features to admire. Success is the tool’s own invisibility: thousands of silent cheap pauses, near-zero loud events. If you read autopsies often, the tool has failed at its real job. No warranty — the promise is not “it won’t go wrong” but “when it does, you’ll know exactly what killed you, and nothing is lost.”
The project has tested sampler/classifier and LMDB primitives plus a fixture-tested, observe-only monitor API. There is no resident daemon, session startup integration, production hook, or deployable application yet; issue #4 remains open. See ADR-0006 and the compliance review for tested evidence and remaining gates. The architecture is recorded as ADRs:
Build is tracked as native sub-issues under the requirements epic. First
adapter targets Claude Code hooks (PreToolUse, UserPromptSubmit,
SessionStart); the core idea is adapter-agnostic. It will be built
test-first against a controlled memory balloon in an isolated session
before anything goes global.
scripts/run.sh help
scripts/run.sh check
scripts/run.sh test Debug
scripts/run.sh test ReleaseSafe
scripts/run.sh observe --once
just --listThis is a development dispatcher, not an application launcher. observe --once
performs one explicit read-only host sample; build, install, service start, and
system-integration requests fail explicitly until a real application runtime
exists. See entrypoint boundaries and AI setup.
MPL-2.0 (see LICENSE) — legally effective today.
This is a deliberate, owner-ruled standalone exception to the estate
canonical MPL-2.0, not drift. It is consistent with the estate
model: MPL-2.0 (= Palimpsest-MPL) has MPL-2.0 as its automatic legal
base, so an MPL-2.0 repo is a clean subset. There is no “PMPL
relicensing debt” — that framing was a corrected false premise.
Canonical policy: hyperpolymath/standards LICENCE-POLICY.adoc.
Rationale:
ADR-0005.
Scaffolded from rsr-template-repo (the neutral RSR skeleton). The
full RSR placeholder bootstrap and the per-file SPDX relicense are
tracked build sub-issues rather than rushed inline — foundation-first.
There is no production installer yet. For a safe development checkout, see AI-assisted development setup. Local structural checks are not proof of runtime safety or release readiness.