Repository navigation
Conversation
Draft
7 of 9 tasks
Contributor
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueThanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Preserve native Lightning execution while making complete experiments easier to inspect, run and continue. This accumulated PR fixes optimizer ownership, scientific loss/metric reporting, identifiable prediction output and freezing/checkpoint behavior; adds local attempt records and two complete workflows; and replaces the documentation's fragmented examples with an ordinary Python/YAML learning path.
For example, a resumed fit now makes saved LR/momentum and restored progress visible alongside the requested recipe. CSV export preserves IDs such as
00001,NA, Unicode and embedded newlines. Invalid literal record options fail before project/component construction, while opaque or dynamic inputs remain deferred.Prepared from
e7ac6d361e70c66034514db744b9f74e38bd6a76againstorigin/main360fdd9980510eab79b05c373fec886b4220bb5e(merge base360fdd9980510eab79b05c373fec886b4220bb5e): 110 changed files, 38 commits.Paired dependency
Review with Sparkwheel #7. Lighter requires its retained-definition/scoped-construction APIs and the structured blocked-path exception; an older 0.0.x release or earlier incompatible development snapshot is insufficient. Both development packages can be installed explicitly from public immutable Git revisions in one resolver invocation. They are not published stable releases, and paired-source PR CI does not establish registry delivery. Start with the pinned public installation and compatibility guide. The committed PR workflow currently pins Sparkwheel
9a67c0796fce60f9c47a4afefe9da9009eb9be38.Suggested review order
1. Native execution and managed construction
Stage overrides and seed validation occur before avoidable construction. Ordinary LighterModule optimizers/schedulers are built at native setup; custom and prebuilt ownership stays native. Stage results are returned, optional stages remain optional, and ordinary imports avoid installing global process/logging policy. Typed blocked-path provenance gives advice only for actual managed optimizer/scheduler timing errors.
Start with src/lighter/engine/runner.py, src/lighter/engine/construction.py, src/lighter/utils/dynamic_imports.py.
Tests: tests/integration/test_execution_inputs.py, tests/integration/test_managed_construction.py, tests/integration/test_native_interfaces.py.
Review risk: Construction timing, import side effects and ownership are compatibility boundaries. The final Spark companion must be present; no silent fallback to an older API is provided.
2. Measurements, prediction, freezing and accelerator controls
Observe the final scientific loss before accumulation normalization, keep metrics without an external logger, isolate evaluation populations, preserve native prediction behavior and stream CSV fields without type inference. Freezer tracks only its selected parameters and restored flags. Tiny accelerator tests count actual optimizer calls and compare state/gradients on the same profile.
Start with src/lighter/model.py, src/lighter/callbacks/csv_writer.py, src/lighter/callbacks/freezer.py.
Tests: tests/integration/test_measurement_contract.py, tests/integration/test_prediction_contract.py, tests/integration/test_freezer_contract.py, tests/integration/test_accelerator_contract.py.
Review risk: Check denominators, tail batches, string IDs, parameter identity and restoration—not only scalar logs or global_step. Host CPU/MPS results do not certify skipped CUDA paths.
3. Static inspection and local attempt records
Add static inspect and JSON runs list/show/diff. Records retain requested configuration, observed optimizers/progress, metrics, artifacts and source/import/environment provenance. Literal record errors fail before discovery/construction; dynamic values and opaque mappings/keys defer to normal resolution/final validation.
Start with src/lighter/engine/inspection.py, src/lighter/engine/records.py.
Tests: tests/integration/test_run_records.py, tests/unit/test_inspection.py.
Review risk: Inspection must not evaluate configured Python. Records can be absent before setup or incomplete after interruption; they are not proof of scientific correctness, isolation or process liveness.
4. Complete workflows and corrected example populations
Add the download-free regression lifecycle and Compare and Continue with explicit validation-selected versus last checkpoints, independent native Lightning code and retained results. Correct CIFAR/LoRA held-out splits and vision-language text/image-group inputs. Expose device/precision selection while preserving the historical CPU defaults. Label other integrations as references instead of implying qualification.
Start with projects/tabular_regression/workflow.py, projects/experiment_comparison/workflow.py, projects/experiment_comparison/native.py, projects/vision_language/dataset.py.
Tests: tests/integration/test_reference_workflow.py, tests/integration/test_showcase_experiment_comparison.py, tests/integration/test_example_protocols.py.
Review risk: The retained research results are revision-scoped. Do not infer a rerun on every device or full specialist validity from changed README labels or a passing miniature diagnostic.
5. Packaging, paired CI, coverage and release preparation
Define development versions/reference build profiles, verify wheels outside checkout with child import origins, combine subprocess coverage, supply exact companion source in PR CI and reject unavailable registry dependencies in ordinary setup. Pin repaired Codecov verification, include LICENSE in artifacts and require stable version-matching tags on main ancestry; manual dispatch builds only.
Start with pyproject.toml, scripts/check_paired_install.py, .github/actions/setup/action.yml, .github/workflows/publish.yml, .github/scripts/check_release_tag.py.
Tests: tests/unit/test_paired_install.py, tests/test_release_tag_contract.py.
Review risk: Deleting the old lock is intentional but registry delivery remains open. Manual builds are not publication; release auth is unchanged and actual upload/tag execution remains untested here.
6. Documentation, first use and reference rendering
Rewrite the manual around install → inspect → change → fit → evaluate/export → continue/compare. Explain native versus managed ownership, checkpoint meaning, concrete errors and source-documentation versions. Generate the research page from its README, repair reference rendering and use durable contributor links.
Start with README.md, docs/quickstart.md, docs/guides/compatibility.md, docs/guides/lighter-module.md, docs/gen_examples.py.
Tests: tests/integration/test_documented_interfaces.py.
Review risk: Install pins identify a reviewed snapshot rather than following a branch. Source inspection and executable examples must stay distinct; rendered docs and agent walkthroughs do not establish human usability.
Validation and current-head status
Current-head Ubuntu/Python 3.12 CI passes. Local checks below retain their separate source/environment scopes.
1643aa9pilot snapshot, native/Lighter CIFAR-subset runs matched independent replay across 1,920 updates, selected-checkpoint evaluation and two continuations. This is historical CPU evidence, not a rerun of the final accelerator-enabled example.Final-head CI run 35469075693 passes at
e7ac6d361e70c66034514db744b9f74e38bd6a76: 506 passed, 10 hardware skips (8 CUDA, 2 MPS), 3 unchanged slow exclusions. Formatting, lint, mypy, the existing 95% overall coverage gate, PR-title and dependency checks pass. The workflow tests the exact Sparkwheel9a67c079 companion. CI Full and CUDA are not qualified; the existing trusted-main Codecov upload guard is unchanged. Both PRs remain drafts. Previous CodeRabbit findings retain their documented disposition; no fresh bot approval is claimed. Independent source and actual GLM reviews of the current implementation are complete.Compatibility and remaining limits
lossplus optionalloss_terms. Per-loader metric isolation and stricter inputs can expose previously accepted ambiguous configurations. Stateful metrics with persistent checkpoint state require explicit native ownership.run: falseopts out, and pre-recorder failures can leave no record. Opaque objects remain application-owned. Freezer does not invent missing optimizer groups or checkpoint history.uv.lockis deleted deliberately. Compatible Sparkwheel must be published first, then Lighter's genuine registry lock and stable release must be qualified. Public source pins/local wheel hash requirements are not registry locks. No merge, tag or publication is included here.Type of change
Checklist
Complete changed-file inventory
1. Native execution and managed construction (14 files)
Msrc/lighter/init.pyAsrc/lighter/main.pyAsrc/lighter/engine/construction.pyMsrc/lighter/engine/runner.pyMsrc/lighter/utils/dynamic_imports.pyAtests/integration/test_execution_inputs.pyAtests/integration/test_import_boundaries.pyAtests/integration/test_managed_construction.pyAtests/integration/test_native_interfaces.pyAtests/integration/test_runner_results.pyMtests/unit/test_engine_cli.pyMtests/unit/test_engine_runner.pyMtests/unit/test_engine_runner_errors.pyMtests/unit/test_utils_dynamic_imports.py2. Measurements, prediction, freezing and accelerator controls (19 files)
Msrc/lighter/callbacks/base_writer.pyMsrc/lighter/callbacks/csv_writer.pyMsrc/lighter/callbacks/file_writer.pyMsrc/lighter/callbacks/freezer.pyMsrc/lighter/data.pyMsrc/lighter/model.pyMsrc/lighter/utils/logging.pyAtests/fixtures/distributed_prediction.pyAtests/fixtures/distributed_training.pyAtests/integration/test_accelerator_contract.pyAtests/integration/test_distributed_contract.pyAtests/integration/test_freezer_contract.pyAtests/integration/test_measurement_contract.pyAtests/integration/test_prediction_contract.pyMtests/unit/test_callbacks_freezer.pyAtests/unit/test_csv_fidelity.pyMtests/unit/test_data.pyMtests/unit/test_model.pyMtests/unit/test_utils_logging.py3. Static inspection and local attempt records (4 files)
Asrc/lighter/engine/inspection.pyAsrc/lighter/engine/records.pyAtests/integration/test_run_records.pyAtests/unit/test_inspection.py4. Complete workflows and corrected example populations (35 files)
Mprojects/README.mdMprojects/cifar10/README.mdMprojects/cifar10/configs/example.yamlAprojects/cifar10/dataset.pyMprojects/eeg/README.mdAprojects/experiment_comparison/README.mdAprojects/experiment_comparison/init.pyAprojects/experiment_comparison/lighter.pyAprojects/experiment_comparison/artifacts.pyAprojects/experiment_comparison/config.yamlAprojects/experiment_comparison/data.pyAprojects/experiment_comparison/high_lr.yamlAprojects/experiment_comparison/native.pyAprojects/experiment_comparison/results.jsonAprojects/experiment_comparison/task.pyAprojects/experiment_comparison/workflow.pyMprojects/huggingface_llm/README.mdMprojects/lora/README.mdMprojects/lora/configs/lora.yamlMprojects/lora/dataset.pyMprojects/medical_segmentation/README.mdMprojects/self_supervised/README.mdAprojects/tabular_regression/README.mdAprojects/tabular_regression/init.pyAprojects/tabular_regression/lighter.pyAprojects/tabular_regression/config.yamlAprojects/tabular_regression/task.pyAprojects/tabular_regression/workflow.pyMprojects/video_recognition/README.mdMprojects/vision_language/README.mdMprojects/vision_language/configs/clip.yamlMprojects/vision_language/dataset.pyAtests/integration/test_example_protocols.pyAtests/integration/test_reference_workflow.pyAtests/integration/test_showcase_experiment_comparison.py5. Packaging, paired CI, coverage and release preparation (17 files)
M.github/actions/setup/action.ymlA.github/scripts/check_release_dependencies.pyA.github/scripts/check_release_tag.pyM.github/workflows/ci-full.ymlM.github/workflows/ci.ymlM.github/workflows/publish.ymlM.github/workflows/release.ymlM.gitignoreMjustfileMpyproject.tomlArequirements/profiles/build.constraintsArequirements/profiles/numpy2.constraintsArequirements/profiles/reference.constraintsAscripts/check_paired_install.pyAtests/test_release_tag_contract.pyAtests/unit/test_paired_install.pyDuv.lock6. Documentation, first use and reference rendering (21 files)
MCONTRIBUTING.mdMREADME.mdMdocs/examples/index.mdMdocs/faq.mdAdocs/gen_examples.pyMdocs/gen_ref_pages.pyMdocs/guides/best-practices.mdAdocs/guides/compatibility.mdMdocs/guides/configuration.mdMdocs/guides/custom-code.mdAdocs/guides/experiment-records.mdAdocs/guides/freezing.mdMdocs/guides/lighter-module.mdMdocs/guides/lightning-module.mdAdocs/guides/predictions.mdMdocs/guides/training.mdMdocs/index.mdMdocs/quickstart.mdMdocs/reference/cli.mdMmkdocs.ymlAtests/integration/test_documented_interfaces.py