parse is permissive; validate is the explicit step that enforces the
format. You choose when to run it:
import { validateLesson, validateManifest } from "learn-content-engine";
const result = validateLesson(JSON.parse(rawLessonJson));
// result: { valid, errors[], warnings[] }
// each issue: { path, message, id, severity: "error" | "warning", docAnchor }
if (!result.valid) console.error(result.errors);
if (result.warnings.length) console.warn(result.warnings);Neither function throws; both return a ValidationResult. valid is
errors-only: warnings never block. Every issue carries a stable id and a
docAnchor; the complete list of ids is the
rule catalog. For an offline author workflow
that surfaces both errors and warnings, use the
learn-content-engine lint CLI.
Validation runs in two layers (field checks before cross-field checks):
The input is checked against the bundled JSON-Schema
(schema/lesson.schema.json /
schema/content-manifest.schema.json,
draft 2020-12) with ajv. This enforces required fields,
field types and lengths, enum values (type, cloze_mode, direction,
media_type, ...), and nested shapes (Pair, PictureImage, ClozeBlank).
The schema is strict: additionalProperties: false everywhere, so unknown
fields are rejected.
Why strict? This engine is a format reference: content authored against it must load in any consumer: for example adaptive-learner, the reference consumer, whose schema is itself strict. Tolerating unknown fields here would let content pass the engine yet fail a strict consumer, the opposite of a reliable reference. Strict rejection keeps parity: if it validates here, it is shape-valid there.
If any structural error is found, validation stops and returns those errors (the semantic layer assumes a well-formed shape).
These rules cannot be expressed in JSON-Schema; they mirror the reference
consumer's (adaptive-learner) Pydantic model_validators one-for-one:
| Rule | Message contains |
|---|---|
Theory step requires body, no exercise; exercise step requires exercise, no body |
THEORY step ... / EXERCISE step ... |
matching requires non-empty pairs (unless from_cards) |
MATCHING exercise requires non-empty 'pairs' |
matching with from_cards requires non-empty card_ids and must not also list explicit pairs |
MATCHING with 'from_cards' |
multiple_choice requires >= 2 options with unique texts |
MULTIPLE_CHOICE requires at least 2 'options' / option texts must be unique |
multiple_choice (single) requires exactly one option marked correct; with multiple at least one |
exactly one option marked 'correct' / at least one option marked 'correct' |
picture_choice requires >= 2 images, exactly one is_correct: "true" |
exactly one image marked |
free_text requires non-empty accept |
FREE_TEXT exercise requires non-empty 'accept' |
word_tiles requires >= 2 tiles; each accept_orderings entry is a permutation |
permutation of [0..n-1] |
cloze (type/select) requires sentence + blanks with markers == blanks.length; select also needs distractors |
CLOZE marker count mismatch |
cloze (multiselect) requires sentence, non-empty accept + distractors, and the two must be disjoint |
must be disjoint |
Every card_ids entry must resolve to a card in the lesson |
references unknown card |
A stable_id is unique within the lesson (exercises and cards share one namespace) |
is used more than once in this lesson |
validateLesson sees ONE lesson, so it can only enforce that a stable_id
is unique inside that document. The set-wide half of the promise (schema v1.9,
engine#90) is exported as a helper the caller drives over the lessons of a set:
import { collectStableIds } from "learn-content-engine";
const report = collectStableIds(lessonsOfOneSet);
// report.total -> how many stable_ids were checked (the checked quantity)
// report.duplicates -> [{ stableId, locations: [{ lessonId, kind, elementId }] }]Version STABILITY (an id still pointing at the same element after an update) is not checkable here at all: it needs the previous version. That is the content repos' stability gate, which diffs the set against its last published state. See the scope and limit of the stage.
Warnings never affect valid or appear in errors; they live in warnings and
flag likely authoring mistakes: an unused card (W-CARD-UNUSED), an ambiguous
matching (W-MATCH-AMBIG), duplicate word tiles without accept_orderings
(W-TILES-DUP), a distractor equal to the answer (W-DISTRACTOR-ANSWER), a
distractor image sharing the correct label (W-PIC-DUP-LABEL), or a hint that
reveals the answer length (W-HINT-LENGTH). Full list + descriptions:
rule catalog.
Each issue is { path, message }:
pathis a JSON-pointer-ish location, e.g./steps/2/exerciseor/steps/2/exercise/card_ids, or/for a root-level problem.messageis a human-readable reason. For a rejected unknown field the offending key is named, e.g.must NOT have additional properties (surprise).
A cloze whose marker count does not match its blanks (invalid input):
An exercise referencing a card that does not exist (invalid input):
// INVALID: no card with id "keopi" in the lesson's cards
{ "type": "word_tiles", "id": "w1", "prompt": "...",
"card_ids": ["keopi"], "tiles": ["a", "b"] }
// -> /steps/0/exercise/card_ids:
// exercise references unknown card 'keopi'An unknown field (strict rejection):
// INVALID: "surprise" is not a known lesson field
{ "id": "x", "title": "t", "steps": [ ... ], "surprise": true }
// -> /: must NOT have additional properties (surprise)validateManifest normalizes the legacy language alias to target_language
before structural checking (the pre-v1.2 alias, see
concepts.md), then applies the strict
manifest schema. A set missing a required field (id, title,
target_language, level, version, lesson_count) is rejected; a set with
neither language nor target_language fails on the missing target_language.
Since 0.20.0 validateManifest also carries two author lints in warnings
(never blocking, like the lesson lints in Layer 3): W-DOMAIN-UNKNOWN for a
domain outside the known vocabulary (KNOWN_CONTENT_DOMAINS), and
W-LEVEL-UNKNOWN for a level that is neither a CEFR band (A1..C2,
case-insensitive) nor, for a non-language set, the explicit none sentinel.
The contract lives in
content domains.
Since 0.21.0 a manifest may declare metadata.retired_ids (deliberate
retirement, unlocked in engine#131). The validator checks what one manifest
can prove: E-RETIRED-IDS-TYPE rejects a list that is not made of strings,
and W-RETIRED-IDS-DUP flags duplicate entries (never blocks). Everything
that needs the lesson inventory - an undeclared disappearance, un-declaring
a published retirement, retired-yet-alive - lives in the stability gate
(V1/V5/V6, see
stable identity).
Besides the two JSON-Schemas the package ships
schema/quality-rules.json: the shared
quality minimums (minExercisesPerLesson, minExerciseTypes,
minFreeTextAccepts, minMatchingPairs, minTheorySteps). Like the two
schemas, its canonical home is this engine (since the v0.6.0 authority flip).
The engine does not evaluate these rules itself (there is no
validateQuality API); the artifact exists so consumer validators
(adaptive-learner, the content repos) can mirror the numbers from the
pinned engine release instead of owning a repo-local copy, the same
engine → consumers channel as the schemas. Consume it via
import qualityRules from "learn-content-engine/schema/quality-rules.json"
or read it from the installed package.
make conformance-real runs the full pipeline (parse + validate) over both
public content repos on demand. It is not part of the mandatory CI: the CI
truth is the vendored fixtures and the doc examples, which run offline. See
architecture.md.