CraftKit is a cross-agent toolkit for creating, improving, and operationalizing prompts and skills for coding agents such as Claude Code and Codex.
Prompt assets and agent skills often become fragmented, provider-specific, and hard to reuse. CraftKit exists to keep them file-first, portable, reviewable, and easy to improve over time.
CraftKit covers two wedges. The first is artifact quality: author prompts and carry sessions forward (craft-*). The second is the repo spec axis: spec-charter lands direction and system shape; spec-grill is optional. Those files are reference contracts that other tools can consume — most directly dev-backlog, which measures sprints and triage against them.
CraftKit is not a general coding-agent workflow suite, project-management layer, deployment system, or runtime framework. The spec axis defines what good looks like — it does not manage tasks, sprints, or backlog priority; that stays with dev-backlog. When a workflow needs those things, CraftKit should produce clear files, specs, or handoffs that another tool can use rather than becoming the tool itself.
Start with the smallest skill that does the job:
| If you need to... | Use |
|---|---|
| write a new prompt or reusable prompt template | craft-prompt |
Reach for the other skills when the job gets more specific:
craft-handoff— end a long session with a durable doc plus a resume prompt.spec-charter— land a brownfield repo's direction and system shape.spec-grill— optional capability contracts.
The skills are the delivery vehicle; the durable part is two review disciplines that hold up as models get more capable — a smarter agent is exactly what finds the loophole in a loose contract or talks a loop past "good enough." Both are written up as standalone references you can apply without adopting any skill:
docs/methodology/predicate-test.md— the 3-axis test (Authority / Distributional / Manipulability) for deciding whether a written contract is safe for an agent to optimize against. Applied byspec-grill.docs/methodology/loop-stop-conditions.md— falsifiable exit conditions (Self-LGTM / persistent fixpoint / no-op / hard cap) for agent improvement loops. Skill-independent: apply it to any loop where an agent grades its own output.
CraftKit installs as Agent Skills for Claude Code, and each skill is also a plain SKILL.md file that can be used by Codex and other compatible agents.
npx skills add sungjunlee/craftkitAdd -g -y for global install without prompts:
npx skills add sungjunlee/craftkit -g -y/plugin marketplace add https://github.com/sungjunlee/craftkit.git
/plugin install craftkit@craftkit
Install from a local clone
git clone https://github.com/sungjunlee/craftkit.git
cd craftkit
npx skills add . -g -yFor Codex or any other agent, see Use in other agents below.
| Skill | Use when | Side effect |
|---|---|---|
craft-prompt |
a new prompt is needed from scratch for any LLM or agent interface | returns copy-pasteable text |
craft-handoff |
a session is ending and the next session needs a copy-paste-ready continuation prompt | writes handoff files and may copy to clipboard |
spec-charter |
a repo needs direction, Objectives, Decisions, system shape, or stale-spec reassessment | creates or amends spec/charter.md and spec/system-map.md |
spec-grill |
optional: keep only when a consumer exists, the contract is cross-tree, or a 3-axis audit is needed — not a required follow-on to charter | creates or refines spec/capabilities.md after evidence review |
When two skills could trigger, choose the least invasive one that answers the request: new or reworked prompt text goes to craft-prompt; session wrap-up goes to craft-handoff. Reviewing or improving an existing artifact needs no dedicated skill — ask for it directly.
The spec-* skills land a spec axis: spec-charter (charter + system map) is the default; spec-grill is optional. Use them when a brownfield repo needs a compact spec axis grounded in real repo evidence instead of a generic architecture document. Monorepo vs sibling-git vs type-1 topology is stated once in skills/spec-charter/references/spec-axis.md.
The spec axis supersedes dev-backlog's retired backlog-charter skill (dev-backlog split that surface into the spec-series in its 0.6.0): spec/charter.md is the successor home for the project reference axis, and spec-charter's amend mode reads a legacy root CHARTER.md as a fallback and migrates it deliberately rather than silently. dev-backlog consumes the axis — it measures sprints and triage against spec/charter.md — but does not own it.
Each skill lives at skills/<skill-name>/SKILL.md — plain markdown with YAML frontmatter, loadable as a Claude Code skill or copy-pasteable into any other agent.
craft-prompt and craft-handoff were optimized through maintainer-local eval passes before the eval-loop skill that ran them was retired. The spec-* skills have maintainer-local or repo-local contract evidence. Publicly reproducible status and local-maintainer evidence boundaries are tracked in docs/status.md.
- generating new prompts from scratch (task, research, templates)
- prompt design and restructuring
- repo spec-axis creation for charter, system map, and optional capability contracts
- durable session continuity between agent sessions
- copy-pasteable outputs for agent workflows
- File-first and diff-friendly
- Small composable units
- Explicit inputs and outputs
- Cross-agent portability (core skill spines stay provider-neutral; platform-specific detail stays in templates or reference files)
- Copy-pasteable results over fancy abstractions
- Weight follows durability — as models improve, move each skill's center of gravity from "tell the model how to think" toward "give the model durable state and direction it cannot hold on its own"
- Durable state and review discipline stay; procedure goes. spec-* ship what a model cannot hold on its own — direction, non-goals with reasons, standing decisions, hard constraints, learnings — plus the 3-axis predicate test; how-to-interview procedure trends to zero as models improve.
- Living spec files hold only the current position; git is the archive. Rewriting a decision beats appending to a ledger; a retired objective keeps only its ID.
Principle 6 is the axis CraftKit is actively re-sized against: machinery (deterministic paths, clipboard, archiving), time-sensitive curated knowledge, and direction-setting judgment contracts are model-independent and stay; raw prescription erodes as models improve and gets cut. Principles 7-8 are its 2026-09 spec-* application (epic #255): the two rules that decide what a spec-* file keeps versus what it lets the interview procedure drop.
AGENTS.md keeps an absolute 500-line format ceiling for each SKILL.md, but CraftKit's release gate is stricter: npm run verify fails when a skill spine exceeds 220 lines or a frontmatter description exceeds 50 words.
- The spine is a router: what the skill delivers, the rules that must hold, and pointers to
references/. Current spines run about 50-70 lines. - A spine growing past ~100 lines is usually restating a rule in a second section or carrying procedure a current model chooses better itself; move examples, platform notes, maintenance commands, or edge-case catalogs into
references/, and cut the rest.
The spine should still be usable alone for the common case: what the skill delivers, its output contract, the rules that must hold, and links to on-demand references. Inputs, steps, examples, and limitations appear only when a capable model would otherwise get them wrong. References carry depth; the spine carries the operating path. Mirrored references are allowed only when the verifier guards them against drift.
See docs/skill-anatomy.md for the canonical per-family section contract each skill is normalized against.
Most CraftKit skills are explicit workflow selectors, not always-on background guidance. Use implicit invocation only when a skill is low-risk and broadly helpful when matched, such as direct prompt drafting.
For explicit-only workflows, pair both platform controls:
# SKILL.md frontmatter, used by Claude Code
disable-model-invocation: true# agents/openai.yaml, used by Codex
policy:
allow_implicit_invocation: falseUse explicit-only policy for skills that edit files, write artifacts, mutate clipboard state, create spec files, or otherwise turn a broad user request into a higher-ceremony workflow. Keep the description concise and useful for manual skill lists even when it is not injected for implicit routing.
Use these lightweight checks after editing skill descriptions or routing boundaries. They are manual contract checks, not a new runtime.
| Prompt | Expected skill | Failure signal |
|---|---|---|
| "write a prompt for GPT" | craft-prompt |
refuses to deliver a copy-pasteable prompt |
| "create a system map for this repo" | spec-charter (map) |
looks for a dedicated map skill instead of routing here |
CraftKit skills are plain markdown with YAML frontmatter, so they port easily:
- Open the relevant
SKILL.md. - Paste the body (everything after the frontmatter) into the target agent's system prompt or instructions.
- For implicit skills, keep the frontmatter
descriptionline as context so the agent knows when to apply the skill. - For explicit-only skills, keep the description in the file for menus and manual selection, but preserve the invocation policy fields above when the target agent supports them.
Run the repo-local smoke check before release or packaging changes:
npm run verifyIt checks JSON syntax, package boundaries, skill frontmatter, SKILL.md line budgets, terminology leaks, required README/status paths, and npm pack --dry-run.
.worknode.yml provides an optional foreground recipe for the existing Node
verifier tests (test/*.test.mjs). This example assumes Node 24.18.1 is already
installed at mise's default location,
$HOME/.local/share/mise/installs/node/24.18.1/bin/node. Adjust the recipe's
node_bin and task PATH for a different installation before running it. The
recipe does not install Node or change the target's global PATH. Use an
inventoried Linux target with that toolchain.
source_args=(--source-include .worknode.yml)
while IFS= read -r file; do
source_args+=(--source-include "$file")
done < <(git ls-files -- package.json scripts test test-support \
skills/spec-grill docs/skill-anatomy.md)
worknode run verify-tests --target YOUR_LINUX_TARGET \
--source "$PWD" "${source_args[@]}" --plan --jsonThe source list includes the real skill used by the verifier fixtures and
excludes the repository's CLAUDE.md symlink, which worknode does not transfer.
Remove --plan to execute. The task writes TAP output to out/verify-tests.tap, which the foreground
runner verifies and returns as its declared artifact. This focused task does
not replace npm test, which also covers the spec-grill extract-signals suite.
sungjunlee/prompt-builder— predecessor project. Its templates and prompt-authoring lessons were absorbed intocraft-prompt; the original 5-step, 6-block method has since been replaced by a lean outcome-and-boundary approach. Kept on GitHub for reference; new work happens here.karpathy/autoresearch— Andrej Karpathy's ML training-loop project that introduced the autoresearch methodology (give an agent a baseline, let it experiment overnight, keep what improves, discard what doesn't). CraftKit's retiredcraft-autoresearchskill adapted that loop discipline to prompt and skill artifacts instead of model training code; the exit-condition methodology it used survives indocs/methodology/loop-stop-conditions.md.byungjunjang/jangpm-meta-skills— four-skill meta toolkit for Claude Code and Codex (blueprint,deep-dive,reflect,autoresearch). Itsautoresearchskill contributed implementation patterns — experiment contract shape, the three-eval-type taxonomy (binary / comparative / fidelity), deletion discipline — to CraftKit's retiredcraft-autoresearchskill.
MIT — see LICENSE.