Skip to content

feat: Add managed construction and inspectable experiment workflows - #177

Draft
ibro45 wants to merge 38 commits into
mainfrom
codex/reliability-and-documentation
Draft

ibro45 wants to merge 38 commits into
mainfrom
codex/reliability-and-documentation

Conversation

@ibro45

@ibro45 ibro45 commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Preserve native Lightning execution while making complete experiments easier to inspect, run and continue. This accumulated PR fixes optimizer ownership, scientific loss/metric reporting, identifiable prediction output and freezing/checkpoint behavior; adds local attempt records and two complete workflows; and replaces the documentation's fragmented examples with an ordinary Python/YAML learning path.

For example, a resumed fit now makes saved LR/momentum and restored progress visible alongside the requested recipe. CSV export preserves IDs such as 00001, NA, Unicode and embedded newlines. Invalid literal record options fail before project/component construction, while opaque or dynamic inputs remain deferred.

Prepared from e7ac6d361e70c66034514db744b9f74e38bd6a76 against origin/main 360fdd9980510eab79b05c373fec886b4220bb5e (merge base 360fdd9980510eab79b05c373fec886b4220bb5e): 110 changed files, 38 commits.

Paired dependency

Review with Sparkwheel #7. Lighter requires its retained-definition/scoped-construction APIs and the structured blocked-path exception; an older 0.0.x release or earlier incompatible development snapshot is insufficient. Both development packages can be installed explicitly from public immutable Git revisions in one resolver invocation. They are not published stable releases, and paired-source PR CI does not establish registry delivery. Start with the pinned public installation and compatibility guide. The committed PR workflow currently pins Sparkwheel 9a67c0796fce60f9c47a4afefe9da9009eb9be38.

Suggested review order

Slice Review focus Files
1. Native execution and managed construction Input authority, initialization timing and native ownership 14
2. Measurements, prediction, freezing and accelerator controls Scientific observations, parameter identity and real updates 19
3. Static inspection and local attempt records Requested source versus observed/restored runtime facts 4
4. Complete workflows and corrected example populations A diagnostic, a native comparison and honest specialist status 35
5. Packaging, paired CI, coverage and release preparation Installable source pair, real artifacts and guarded publication 17
6. Documentation, first use and reference rendering One readable ordinary-Python path for people and agents 21

1. Native execution and managed construction

Stage overrides and seed validation occur before avoidable construction. Ordinary LighterModule optimizers/schedulers are built at native setup; custom and prebuilt ownership stays native. Stage results are returned, optional stages remain optional, and ordinary imports avoid installing global process/logging policy. Typed blocked-path provenance gives advice only for actual managed optimizer/scheduler timing errors.

Start with src/lighter/engine/runner.py, src/lighter/engine/construction.py, src/lighter/utils/dynamic_imports.py.

Tests: tests/integration/test_execution_inputs.py, tests/integration/test_managed_construction.py, tests/integration/test_native_interfaces.py.

Review risk: Construction timing, import side effects and ownership are compatibility boundaries. The final Spark companion must be present; no silent fallback to an older API is provided.

2. Measurements, prediction, freezing and accelerator controls

Observe the final scientific loss before accumulation normalization, keep metrics without an external logger, isolate evaluation populations, preserve native prediction behavior and stream CSV fields without type inference. Freezer tracks only its selected parameters and restored flags. Tiny accelerator tests count actual optimizer calls and compare state/gradients on the same profile.

Start with src/lighter/model.py, src/lighter/callbacks/csv_writer.py, src/lighter/callbacks/freezer.py.

Tests: tests/integration/test_measurement_contract.py, tests/integration/test_prediction_contract.py, tests/integration/test_freezer_contract.py, tests/integration/test_accelerator_contract.py.

Review risk: Check denominators, tail batches, string IDs, parameter identity and restoration—not only scalar logs or global_step. Host CPU/MPS results do not certify skipped CUDA paths.

3. Static inspection and local attempt records

Add static inspect and JSON runs list/show/diff. Records retain requested configuration, observed optimizers/progress, metrics, artifacts and source/import/environment provenance. Literal record errors fail before discovery/construction; dynamic values and opaque mappings/keys defer to normal resolution/final validation.

Start with src/lighter/engine/inspection.py, src/lighter/engine/records.py.

Tests: tests/integration/test_run_records.py, tests/unit/test_inspection.py.

Review risk: Inspection must not evaluate configured Python. Records can be absent before setup or incomplete after interruption; they are not proof of scientific correctness, isolation or process liveness.

4. Complete workflows and corrected example populations

Add the download-free regression lifecycle and Compare and Continue with explicit validation-selected versus last checkpoints, independent native Lightning code and retained results. Correct CIFAR/LoRA held-out splits and vision-language text/image-group inputs. Expose device/precision selection while preserving the historical CPU defaults. Label other integrations as references instead of implying qualification.

Start with projects/tabular_regression/workflow.py, projects/experiment_comparison/workflow.py, projects/experiment_comparison/native.py, projects/vision_language/dataset.py.

Tests: tests/integration/test_reference_workflow.py, tests/integration/test_showcase_experiment_comparison.py, tests/integration/test_example_protocols.py.

Review risk: The retained research results are revision-scoped. Do not infer a rerun on every device or full specialist validity from changed README labels or a passing miniature diagnostic.

5. Packaging, paired CI, coverage and release preparation

Define development versions/reference build profiles, verify wheels outside checkout with child import origins, combine subprocess coverage, supply exact companion source in PR CI and reject unavailable registry dependencies in ordinary setup. Pin repaired Codecov verification, include LICENSE in artifacts and require stable version-matching tags on main ancestry; manual dispatch builds only.

Start with pyproject.toml, scripts/check_paired_install.py, .github/actions/setup/action.yml, .github/workflows/publish.yml, .github/scripts/check_release_tag.py.

Tests: tests/unit/test_paired_install.py, tests/test_release_tag_contract.py.

Review risk: Deleting the old lock is intentional but registry delivery remains open. Manual builds are not publication; release auth is unchanged and actual upload/tag execution remains untested here.

6. Documentation, first use and reference rendering

Rewrite the manual around install → inspect → change → fit → evaluate/export → continue/compare. Explain native versus managed ownership, checkpoint meaning, concrete errors and source-documentation versions. Generate the research page from its README, repair reference rendering and use durable contributor links.

Start with README.md, docs/quickstart.md, docs/guides/compatibility.md, docs/guides/lighter-module.md, docs/gen_examples.py.

Tests: tests/integration/test_documented_interfaces.py.

Review risk: Install pins identify a reviewed snapshot rather than following a branch. Source inspection and executable examples must stay distinct; rendered docs and agent walkthroughs do not establish human usability.

Validation and current-head status

Current-head Ubuntu/Python 3.12 CI passes. Local checks below retain their separate source/environment scopes.

Check Recorded result and scope
Full reference source suite 508 passed, 8 CUDA-only skips, 3 unchanged slow exclusions; statement/branch coverage 95.52%, with Ruff/format/mypy passing. Python 3.11.14, Torch 2.7.1, Lightning 2.5.1. Source/tests remained fixed during the run; only license-file metadata changed. This predates the later Spark pickle repair and final opaque-key inspection refinement.
Later focused regressions 61 record checks passed after deferring nested opaque keys; the Spark exception repair is covered by its full 930-test suite and installed old/new artifact controls. Eight documented-interface checks passed after the final documentation corrections.
Numerical and accelerator controls CPU float32, CPU bfloat16 and MPS float32 exercised native/plain-Torch comparisons, actual optimizer steps, unscaled gradients/clipping, partial accumulation windows and checkpoint restoration. Wrong state/scaling/trace controls are rejected. CUDA tests are explicit skips on this host.
Public installation and artifacts Literal public pip/Git installation, matching example checkout and complete CPU lifecycle passed. Copy-only stable promotions built and installed wheels and sdist-derived wheels; all six archives carry exact LICENSE text, and installed exception pickle behavior is checked against the failing old wheel. These stable rehearsal receipts bind their recorded snapshots. A subsequent fresh wheel build/install of the exact final committed pair passed all CPU fit/test/predict/continue/inspect/records checks. The final public installation snapshot12c09fc/b73e786 also passed its literal onboarding commands; it differs from final heads only in documentation/CI pointers.
Earlier scientific walkthrough At the recorded 1643aa9 pilot snapshot, native/Lighter CIFAR-subset runs matched independent replay across 1,920 updates, selected-checkpoint evaluation and two continuations. This is historical CPU evidence, not a rerun of the final accelerator-enabled example.
Documentation/tooling Strict site builds and the final documented-interface checks passed. Release guards passed 20 network-free tests per repository and actionlint; no tag, dispatch or publication was used to test them.

Final-head CI run 35469075693 passes at e7ac6d361e70c66034514db744b9f74e38bd6a76: 506 passed, 10 hardware skips (8 CUDA, 2 MPS), 3 unchanged slow exclusions. Formatting, lint, mypy, the existing 95% overall coverage gate, PR-title and dependency checks pass. The workflow tests the exact Sparkwheel9a67c079 companion. CI Full and CUDA are not qualified; the existing trusted-main Codecov upload guard is unchanged. Both PRs remain drafts. Previous CodeRabbit findings retain their documented disposition; no fresh bot approval is claimed. Independent source and actual GLM reviews of the current implementation are complete.

Compatibility and remaining limits

  • Managed optimizers are constructed at native setup against actual parameters. Eager consumers should share the requested scalar or use the appropriate runtime hook; custom constructors/hooks and prebuilt objects retain native ownership. An LR override alone does not replace restored optimizer state.
  • Automatic-training loss dictionaries require a scalar loss plus optional loss_terms. Per-loader metric isolation and stricter inputs can expose previously accepted ambiguous configurations. Stateful metrics with persistent checkpoint state require explicit native ownership.
  • Local records are observations, not scientific proof or process-liveness checks; run: false opts out, and pre-recorder failures can leave no record. Opaque objects remain application-owned. Freezer does not invent missing optimizer groups or checkpoint history.
  • The incompatible old registry uv.lock is deleted deliberately. Compatible Sparkwheel must be published first, then Lighter's genuine registry lock and stable release must be qualified. Public source pins/local wheel hash requirements are not registry locks. No merge, tag or publication is included here.
  • CUDA/fp16, multi-node and arbitrary sharding/distributed continuation are not qualified. Linux image acquisition was blocked in the local environment. Fresh agent walkthroughs are engineering checks, not unfamiliar-human studies; specialist examples and company workloads still need their own validation.

Type of change

  • Runtime fixes and explicit compatibility changes
  • New functionality with regression coverage
  • Documentation, examples and development tooling
  • Security fix

Checklist

  • Changed behavior has targeted regression coverage and migration notes
  • Local checks are reported with their source/profile scope
  • Full accumulated diff is included in the review map below
  • Final-head remote CI verified at the linked runs
  • Stable release and deployment readiness (separate work)
Complete changed-file inventory

1. Native execution and managed construction (14 files)

2. Measurements, prediction, freezing and accelerator controls (19 files)

3. Static inspection and local attempt records (4 files)

4. Complete workflows and corrected example populations (35 files)

5. Packaging, paired CI, coverage and release preparation (17 files)

6. Documentation, first use and reference rendering (21 files)

@coderabbitai

coderabbitai Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant