Skip to content

Latest commit

 

History

91 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

License: MPL-2.0 Status: design

Graceful degradation for concurrent LLM/agent terminals. When system resources get hairy, sessions back off voluntarily and in priority order — before the OS OOM-kills everything at once — so the work that matters survives and a returning human comes back to an autopsy, not a smoking crater.

Why

Run several agent terminals on one box and the failure mode is brutal: nothing backs off, memory craters, and the OS kills the whole VM — every session dies at once, ungracefully, mid-thought. llm-grace makes that failure legible, bounded, and recoverable.

How (one paragraph)

A tiny monitor reads cheap signals the kernel already emits (/proc counters — memory, swap-out rate, iowait, disk saturation, loadavg) and reads them as signatures, not thresholds: the classic “disk-thrash
low CPU” pattern is the pre-OOM tell. Sessions coordinate through one shared crash-safe LMDB ledger. Under pressure, the least valuable work yields first — value measured as motion, not size, so a thrashing runaway is shed before productive work — climbing a per-session ladder: silent PAUSE (auto-resumes) → CHECKPOINT (a human-in-the-loop latch, loud) → clean TERMINATE (state already saved, never auto-relaunched). Every transition is timestamped; terminated sessions get a per-session autopsy, and the whole episode aggregates into one timeline so you can see how it went wobbly.

Ethos (anti-failure-theatre)

The autopsy and death-spiral timeline are fallbacks you hope never to read, not features to admire. Success is the tool’s own invisibility: thousands of silent cheap pauses, near-zero loud events. If you read autopsies often, the tool has failed at its real job. No warranty — the promise is not “it won’t go wrong” but “when it does, you’ll know exactly what killed you, and nothing is lost.”

Status

The project has tested sampler/classifier and LMDB primitives plus a fixture-tested, observe-only monitor API. There is no resident daemon, session startup integration, production hook, or deployable application yet; issue #4 remains open. See ADR-0006 and the compliance review for tested evidence and remaining gates. The architecture is recorded as ADRs:

Build is tracked as native sub-issues under the requirements epic. First adapter targets Claude Code hooks (PreToolUse, UserPromptSubmit, SessionStart); the core idea is adapter-agnostic. It will be built test-first against a controlled memory balloon in an isolated session before anything goes global.

Developer Entry Points

scripts/run.sh help
scripts/run.sh check
scripts/run.sh test Debug
scripts/run.sh test ReleaseSafe
scripts/run.sh observe --once
just --list

This is a development dispatcher, not an application launcher. observe --once performs one explicit read-only host sample; build, install, service start, and system-integration requests fail explicitly until a real application runtime exists. See entrypoint boundaries and AI setup.

Licence

MPL-2.0 (see LICENSE) — legally effective today.

This is a deliberate, owner-ruled standalone exception to the estate canonical MPL-2.0, not drift. It is consistent with the estate model: MPL-2.0 (= Palimpsest-MPL) has MPL-2.0 as its automatic legal base, so an MPL-2.0 repo is a clean subset. There is no “PMPL relicensing debt” — that framing was a corrected false premise. Canonical policy: hyperpolymath/standards LICENCE-POLICY.adoc. Rationale: ADR-0005.

Provenance / housekeeping

Scaffolded from rsr-template-repo (the neutral RSR skeleton). The full RSR placeholder bootstrap and the per-file SPDX relicense are tracked build sub-issues rather than rushed inline — foundation-first.

AI-Assisted Installation

There is no production installer yet. For a safe development checkout, see AI-assisted development setup. Local structural checks are not proof of runtime safety or release readiness.

About

Graceful degradation for concurrent LLM/agent terminals: cheap /proc-signature load-shedding, a shared crash-safe ledger, and legible post-mortems. Fail legibly, not catastrophically.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages