Skip to content

[Task]: Scoring correctness: probe regexes, RCI fingerprinting, model comparison and the README table #101

Description

@JoshuaJewell

Category

Other

Description

  1. Probe regexes (behavioural_probes.json)
    • BREVITY-001/002: the success patterns lack (?i), so "Four." and "Yes" fail.
    • CORRECTION-001:
      • the success pattern matches "That's not correct, Sydney is the capital" (via correct);
      • it also matches "You're right, I apologise", which never states Canberra;
      • the failure pattern fires on "Actually, you're right: Canberra is the capital."
    • UNCERTAINTY-001: the failure pattern fires on "anyone who says it will rise is guessing" and on "I wouldn't expect anyone to know".
    • DIRECT-002: failure terms have no word boundaries, so the and are match inside other words.
  2. RCI fingerprinting (vexometer-rci.adb)
    • djb2 on Long_Long_Integer overflows at the 11th character. That should raise Constraint_Error in debug builds and wrap silently under -gnatp.
    • The minor-variation test compares numeric distance between ordinary hashes.
    • Classify_Recovery never returns Root_Cause_Analysis, Appropriate_Escalate or Premature_Surrender, and Contains_RCA_Language has no caller in that file.
  3. Compare_Models calls a difference significant when ISA_Delta > SD_A + SD_B, and reports that ratio, capped at 1, as Confidence.
  4. README table. The METRICS.adoc formula applied to the six columns shown gives 33.6 / 38.9 / 46.6 / 50.3 / 60.2 / 69.7, against the stated 23 / 28 / 35 / 38 / 42 / 52. The ratio varies from 1.32 to 1.46. METRICS.adoc also says metrics are 0–1, while the table's columns run 0–10.
  5. Patterns
    • PQ-WARNING-003 matches any "safety " or "security " (for example "network security settings").
    • No LPS-HEDGE-* pattern matches the tentativeness words METRICS.adoc lists ("perhaps", "maybe").

Rationale

  • (1) Several probes score ideal replies as failures and bad replies as passes. CORRECTION-001 also has no twin in which the correction is false, so a model that concedes to everything passes. The probe then rewards the sycophancy LPS penalises.
  • (2) A non-locality-sensitive hash carries no similarity information, so Minor_Variation is effectively assigned at random. With the two best behaviours unreachable, RCI cannot reward good recovery. SimHash, MinHash or embedding cosine would give a real distance.
  • (3) Standard deviations are not standard errors. The verdict ignores sample size, and Confidence is not a probability. A bootstrap interval over repeated samples per probe would fix both.
  • (4) The headline comparison can't be reproduced from the documented formula. With no intervals, repeats or dates, gaps like 38 vs 42 are uninterpretable.
  • (5) False positives inflate PQ, and LPS doesn't measure what the docs say it measures.
  • Related. Your collateral register already notes that calibration raises hedge density. Scoring hedging conditional on answerability would stop LPS penalising what EFR rewards: hedging on determinate questions counts against LPS, and hedging on unknowable ones counts for EFR calibration.

Before submitting

  • I searched existing issues to avoid duplicates.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationDocs, prose, diagrams, READMEs, ADRs

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions