Skip to content

Add a validated Julia statistics layer with exact counts/rationals and higher precision #1

Description

@hyperpolymath

Suggested labels: enhancement, statistics, scientific-validation.
Companion: Symbolic-engine proposal, which is blocked by this work and a separate approval gate.

User direction and scope

Add an optional Julia-owned statistical-analysis layer to MetaManifold, surfaced through Stipple/Vue. Preserve every existing feature and scientific workflow; do not silently change existing pipeline algorithms or previously saved analyses.

The user is interested in the first three forms of “exact” capability discussed, but prioritises only the first two for initial implementation:

  1. Exact integer counts and rational arithmetic where mathematically appropriate.
  2. Higher-precision floating-point computation for explicitly supported operations.
  3. Exact statistical tests: a later, separately validated extension, not part of the initial implementation acceptance gate.

Higher precision is not exact arithmetic. Neither removes measurement error, sampling uncertainty, model misspecification or biological bias. Symbolic mathematics is explicitly out of scope for this issue.

Proposed work

Numeric integrity

  • Preserve integer counts; detect overflow and use checked/wider or arbitrary-size integers where justified.
  • Support exact rational transformations for selected operations, with explicit handling of zero denominators and resource limits.
  • Introduce an explicit numeric policy covering supported ordinary precision and arbitrary precision, precision settings, rounding/conversion rules and provenance.
  • Audit CSV, JSON, database, browser, R and external-tool boundaries for precision loss, integer range and serialization compatibility. Never silently downgrade a requested numeric mode.
  • Isolate precision configuration across concurrent analyses; avoid shared mutable precision settings leaking between users or jobs.
  • Clearly distinguish exact stored/computed values, numerical approximations and rounded display values.

Statistical methods

  • Define a bounded initial catalogue of methods before implementation, documenting supported response types and study designs.
  • Support maximum-likelihood fitting for justified models, with explicit parameter constraints, convergence, identifiability and boundary-estimate checks.
  • Support appropriate parametric and nonparametric analyses. Nonparametric is a class of methods, not a universal safe distribution or an assumption-free alternative.
  • Treat Poisson/negative-binomial and binomial/beta-binomial models as candidates for review, not automatically approved choices for sequencing data.
  • Require appropriate handling of overdispersion, zeros, sequencing depth, compositionality, covariates, pairing, blocking and repeated measurements where relevant.
  • For permutation/bootstrap methods, validate exchangeability or resampling units, record seeds and resampling settings, and report uncertainty and Monte Carlo limitations.
  • Specify uncertainty estimation, diagnostics, effect sizes, multiple-testing policy and warnings for every supported method.

Safe configuration and user experience

  • Attach versioned configurations to immutable source runs; create linked derived analyses rather than overwrite existing results.
  • Default to descriptive summaries when the information needed for valid inference is unavailable. Do not guess missing study design.
  • Never silently switch statistical methods or select a method solely by a normality test or favourable p-value.
  • Failed fits, unsupported precision, invalid input and resource exhaustion must produce explicit unsuccessful states, not plausible-looking results.
  • Record inputs, transformations, model, settings, seeds, software/reference versions and diagnostics for reproducibility.
  • Explain assumptions and numerical limitations in accessible user-facing language.
  • Keep application logic Julia-owned; no new application TypeScript. Small necessary JavaScript adapters and Bun-managed tooling remain permitted.

Required deep validation and acceptance criteria

This is a scientific capability, not merely a form and an optimiser. Define method-specific acceptance tolerances and study-design coverage before claiming support.

  • Approve the bounded method catalogue and publish supported/unsupported conditions.
  • Publish numeric conversion and cross-language/storage contracts, including integer/rational round trips and explicit rejection of unsupported representations.
  • Test known-answer cases and independent reference implementations; shared implementation assumptions alone are not independent validation.
  • Test numerical stability across precision levels and adversarial cases: huge counts, sparse/all-zero tables, tiny probabilities, near-singular fits, boundary parameters, overflow/underflow and invalid/nonfinite inputs.
  • Add property/generative tests for meaningful invariants and metamorphic relationships, plus fuzzing at parsing/configuration boundaries.
  • Demonstrate negative controls and mutation/sensitivity tests: intentionally wrong inputs or implementations must be detected. Record commands and failing evidence.
  • Use simulation studies to assess estimator bias, uncertainty-interval coverage and relevant error rates across supported conditions. Set justified tolerances in advance.
  • Validate representative biological datasets and appropriate null/alternative cases; document where model assumptions fail.
  • Verify reproducibility, serialization, concurrent precision isolation, cancellation, resource limits and failure recovery.
  • Test UI-to-backend-to-result workflows, accessibility, provenance display and clear unsuccessful states.
  • Verify existing scientific behaviour remains unchanged when the new analysis layer is not selected.
  • Benchmark time/memory and bound expensive precision/resampling operations on supported platforms.
  • Run the applicable repository/owner standards gates in CI, retaining results, coverage and benchmark artifacts; justify genuine N/A entries.
  • Obtain independent statistical/scientific review and record limitations and release approval.

“Fully and deeply tested” means documented coverage of the declared scope, independent review and explicit residual-risk acceptance—not a claim that exhaustive testing proves all possible behaviour correct. Critical scientific-correctness, security or data-integrity defects must be resolved before release or before unblocking symbolic work.

Follow-up: exact statistical tests

Exact statistical tests remain desired, but are deferred beyond the initial integer/rational and higher-precision work. Each future test needs its own assumptions, supported designs, computational limits, known-answer cases and validation. Do not substitute an approximate method silently when an exact method is infeasible.

Dependency gate for symbolic mathematics

The companion symbolic-engine issue must stay blocked until the statistics layer completes the above validation, evidence is reviewed, blockers are resolved, and the owner explicitly approves moving to the symbolic stage. Shipping a UI, obtaining green unit tests, or closing this issue administratively is not sufficient to open that gate.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions