Skip to content

Program health refresh (run 5) - #16

Open
github-actions[bot] wants to merge 96 commits into
mainfrom
pipeline/program-health-5
Open

github-actions[bot] wants to merge 96 commits into
mainfrom
pipeline/program-health-5

Conversation

@github-actions

Copy link
Copy Markdown

Program health refresh

Provenance: data exported by dashboard/fetch_jira.py this run;
every figure recomputes from dashboard/data/jira_issues.json.
Review the numbers, then merge — merging publishes the refresh.

Ashritha Gonuguntla and others added 30 commits April 19, 2026 16:51
Complete Software Engineering System for eParts capstone:
- Multi-agent architecture: 12 domain agents (requirements, architecture,
  coach memory, knowledge, coding, project management)
- Pipeline executor with SharedMemory wiki + EventBus for cross-pipeline
  communication (Karpathy wiki pattern)
- Central FastAPI orchestrator with 20+ endpoints
- MCP clients: GitHub (live), Jira, Bitbucket, Vector Store (ChromaDB)
- Prompt registry with version control and peer review workflow
- Risk register auto-populated from architecture + coach sessions
- ETVX process model + GQIM measurement plan
- Interactive dashboards: metrics, intelligence graph, agent flow
- Processed transcripts: 5 client meetings + 4 coach sessions
- Full documentation: SES assessment, SDLC choice, practice areas,
  why-everything justification, presentation guide

Made-with: Cursor
Atlassian deprecated /rest/api/3/search — migrated to /search/jql endpoint.
Added GitHub MCP client with Contents API for file commits, branch creation,
and PR management. Added orchestrator endpoints for both integrations.

Made-with: Cursor
- ticket_creator: create_ticket → create_issue, read pipeline_context
- wbs_updater: get_sprint_state → get_board_status, deposit to wiki
- alert_agent: real Jira queries, wiki drift check, unlinked reqs check
- weekly_digest: pull from wiki + Jira + EventBus instead of empty lists
- req_extractor: read from pipeline_context, deposit to wiki
- stale_detector: query wiki + Jira for stale requirements
- traceability_builder: build matrix from wiki + Jira data
- context_packager: gather from wiki + Jira + EventBus
- test_generator: actually commit generated tests to repo
- decision_logger: read pipeline_context, deposit to wiki
- commitment_tracker: search Jira + wiki for delivery evidence
- registry: add GitHub MCP to all agents that commit files
- README: full architecture docs, trigger flow, extensibility guide

Made-with: Cursor
…risk chains

New TraceabilityStore (SQLite) with artifacts and directed links:
- 172 artifacts seeded from 5 client meetings, 4 coach sessions, Jira,
  risk register, and architecture decisions
- 134 links (RAISED_IN, MITIGATES, IMPLEMENTS, ADDRESSES)
- 77.3% coverage — can trace any concern to what addresses it
- Forward/backward chain traversal (follow any artifact to its origin
  or to what it became)
- Gap detection: finds unmitigated risks and unaddressed concerns
- API endpoints: /traceability, /traceability/{id}, /traceability/gaps/*
- traceability_builder agent now generates full matrix from the store

Made-with: Cursor
- Add 12 requirement artifacts derived from meetings + architecture report
- Replace naive keyword matching with domain-aware thematic linking
- Use Jira labels (architecture, ML, requirements, etc.) for auto-linking
- Add explicit curated mappings for all 50 Jira tickets
- New link types: BECAME (concern→requirement), DECIDED_BY, TRIGGERED
- Full lifecycle chains: concern → decision → requirement → architecture → risk → Jira
- Regenerate traceability.md with 7 sections: overview, lifecycle paths,
  req→jira coverage, architecture chains, risk mitigation, concern traceability,
  commitment fulfillment

Before: 134 links, 3 link types, 37 orphan Jira tickets
After: 760 links, 7 link types, 0 orphan artifacts
Made-with: Cursor
New traceability tab shows:
- Summary cards (184 artifacts, 760 links, 0 orphans)
- Artifact type and relationship type breakdowns with bars
- Top 15 relationship patterns
- Full lifecycle chain explorer (3 chains showing 8 types)
- Requirement → Jira ticket coverage matrix
- Risk mitigation status (all 16 mitigated)
- Concern traceability (addressed by + became)

Made-with: Cursor
Full system architecture showing all 24 agents, 7 pipelines, 8 MCP servers,
shared infrastructure (SharedMemory, EventBus, TraceabilityStore, PromptRegistry,
RiskRegister, MetricsCollector), cross-pipeline event flows, storage layer
(7 SQLite DBs + ChromaDB), and output artifacts. Includes live status indicators.

Made-with: Cursor
Full explanation for someone with no context: what the system is, end-to-end
walkthrough (meeting → Jira ticket), all 7 pipelines, how agents communicate
(SharedMemory + EventBus), traceability store with example chains, and complete
mapping to the CMU meta-model framework (artifacts, processes, resources,
measurements). Includes ASCII diagrams throughout.

Made-with: Cursor
arjunnai and others added 27 commits July 29, 2026 00:30
…ching

ETIM class/feature/value matching is performed by the ML service AFTER
attribute matching. It is not done by the Intermediate Structured Layer
during normalization. The v1.1 integration edit put it in the wrong
component, and HLR-2, section 2.1 and SCEN-1 all read that way.

An ETIM assignment is a matching decision with a confidence attached, so it
needs a model and a route to human review. Normalization has neither, and
putting the assignment there would also collapse evidence and interpretation
into one step, which is what ADR-014 exists to prevent.

Spec v1.3:
- HLR-2 is now mechanical cleanup only, with ETIM assignment explicitly
  downstream in the ML service (HLR-6, FR-9).
- Section 2.1 Intermediate Structured Layer no longer assigns ETIM
  identifiers; Attribute Prediction Service (ML) now owns all matching in
  two named phases.
- FR-9 attributed to the ML service and ordered after attribute matching.
- SCEN-1 step 3 no longer assigns ETIM; step 4 cites FR-3 and FR-9.

ADR-016 retitled from "Decompose Attribute Matching into Staged ETIM ..." to
"Add ETIM Matching as a Second ML Phase After Attribute Matching". Phase 1 is
ADR-003's matcher, unchanged and no longer described as narrowed. Phase 2 is
the ETIM pipeline. ETIM-keying-during-normalization is recorded as an
alternative considered and rejected.

Diagram v6.0 now shows both phases inside the PredictionServiceInterface
boundary: a solid "ML attribute matching" band over a dashed "ML ETIM
matching" group, and matched_product_attribute is labelled ML-owned.

Swept adr-index, the requirements-to-ADR matrix, the change record, the
changelog, studio-req+arch, the slide deck and the talking script.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Implements the deterministic-tooling and evals recommendations from the
AI-tools coaching session with Cory Gwin (GitHub Copilot), 2026-07-24.

evals/ - scenario-based eval harness. 20 scenarios across two suites:
orchestrator routing (executes the real routing table) and defect-triage
skill contracts. Runs stdlib-only with no API key in under a second, so it
gates every PR. Detects lost capabilities: removing an agent from the
routing table fails the suite naming the missing agent, which is the point
Cory made about agent output being non-deterministic and ordinary tests
being unable to catch a regression.

tools/lint_ses.py - custom linter for our own agent conventions. 5 rules,
stdlib-only so it runs without the ML dependencies installed. SES003
(agent not registered in the orchestrator) and SES005 (commit_file without
an explicit branch, which defaults to main and bypasses the human approval
gate) both found real problems on first run.

.github/workflows/quality-gates.yml - three jobs, cheapest first:
deterministic checks, Trivy security scan, and the offline eval tier. The
model-dependent eval tier is opt-in via workflow_dispatch since it costs
tokens.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
More of the Cory Gwin session (2026-07-24). His framing was that each agent
is a stage in the workflow with one defined responsibility, and that a
sub-agent costs almost nothing to add.

RefactorAgent - cleanup and reorganisation of code that already works. It
refuses to run when the trigger says tests are failing, and it filters test
files out entirely so it cannot rewrite a test to accommodate a refactor.
Static heuristics run with no API key.

TestReviewAgent - reviews tests rather than implementation: whether sad
paths are covered, whether assertions are meaningful, and whether a
coverage number is real or gamed. Cross-checks reported coverage against
assertion density.

PlanGeneratorAgent - translates a spec into a reviewable implementation plan
before any code is written, including the "grill me" step where the agent
surfaces every clarifying question up front rather than discovering the
ambiguity mid-build. docs/plan_template.md is the standard format: files to
change, code structures, class breakdown, and required tests with what each
one proves.

All three are comment-only and set requires_human_review. Registering them
in the orchestrator also clears the linter's SES003 findings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Project Health (was "Program Health"):

Stream grouping was walking one parent link and then truncating at the em
dash, so parent tasks appeared as phantom streams: "OCR-8 — OCR testing on
Azure" rendered as a stream called "OCR-8", and the weekly Ingestion
sub-epics as "W3".."W8". It now walks the parent chain up to a real Epic.
The fold verifies: OCR 9/13 + OCR-8 9/9 + OCR-9 4/4 = 22/26.

Work with no owning Epic is split into two honest buckets instead of one
unexplained group: work created before the Epics existed (so it could not
have been filed under one) and work that is genuinely unfiled.

Forecast is now calendar aware. It counts only working weeks and maps them
onto the academic calendar, so it does not spend effort in the summer-to-fall
break or the fall break week, and it reports "after 18 Dec" rather than
inventing a 2027 date. Velocity samples from the week after the mid-June
tool migration: that week shows 128 created and 109 resolved, which is
already-completed work being entered, not throughput. Those issues still
count toward percent complete.

Zero-activity weeks are excluded from the velocity sample, because five
consecutive semester-break weeks were dragging the median to 5/wk when the
team's actual delivery rate is 13/wk.

SES dashboard: added Risk Register Generation as an eighth practice area
with its three real sources, and refreshed the codebase size tile by
counting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ndary

qa/ - a small purpose-built interface for exercising pipeline/vtt_processor.py
in isolation, per Cory's point that code is cheap enough now to build
throwaway QA tools rather than only reading code. Stdlib only, binds to
localhost, read only. It found four real defects: .cc.vtt files parse to zero
turns silently, question detection has a dead branch that can never see a
question mark, unanchored email matching resolves strangers to a teammate,
and out-of-order cues produce a negative duration.

docs/ses_product_repo_integration.md - the harness stays in its own repo and
acts on the product repo through the Bitbucket API client rather than being
merged into it, with the alternatives we rejected and why. Rollout is staged
by trust: observe, advise, then propose. Granting write access is gated on
fixing SES005 first, because commit_file defaults to main and seven agents
call it without naming a branch, which would let an agent write to the team's
main with no pull request.

.env.example now points at the product repo and documents the token scope,
starting at PR comment only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…health

SES hardening from the Cory Gwin session, plus Project Health fixes
…abase

ADR-020 states plainly that we are "deliberately accepting that the catalog
will go stale relative to ETIM" and that the ADR "should be revisited before
any such transition". That accepted risk was never in the register, so the
register understated what we know. RISK-ARCH-09 records it: pinned to ETIM 10.0
EI under constraint C-4, so if suppliers move to 11.0 the new classes and
features are unavailable and affected products fall to ETIM Other handling or
review. Status is mitigating rather than open, because the pin is deliberate and
carries controls: etim_release_id travels through the reference tables, the
interpretation table and the PIMS writeback key, and the loader rejects an
archive whose release does not match the declared one.

Register now 20 risks: 4 critical, 7 high, 9 medium; 3 mitigating, 17 open.

pipeline/render_risk_register.py is new. docs/risk_register.md carried the
footer "Auto-generated from risk_register.db" while nothing in the repo
generated it, so the markdown and the database could drift apart silently and
had. The register is now rendered from the database, every header count is
computed, and two sections were added: a traceability table showing which
requirements and architecture artifacts each risk threatens, and an
exit-condition check that names any risk lacking an owner or a mitigation
rather than publishing it quietly.

Traceability graph gains the risk plus MITIGATES edges to ADR-020, ADR-013,
ADR-014 and ADR-017. Graph totals recomputed rather than hand-edited: 189
artifacts, 764 links, 324 MITIGATES.

Also in this commit:

- Risk Register added as an eighth practice area on the interactive
  architecture dashboard, with the four-step end-to-end panel and the
  counterfactual block the other pipelines have. Marked green because the
  pipeline has run and produced the register.
- The two SES dashboards contradicted each other on pipeline count and MCP
  integrations. Reconciled to 8 pipelines and 6 MCP clients wired (7 exist;
  Drive is built but not registered in the orchestrator).
- Burnup chart now projects forward at the measured throughput across the
  remaining working weeks, with an 18 December deadline marker, so schedule
  health is visible on the same axes as the history instead of only in a
  separate histogram.
- Stale traceability figures corrected across eight docs and three dashboards.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…health

Log the ETIM release-pin risk, and generate the register from the dat…
…et-field coverage

Slide 5 of the Software System deck, presented in the closing slot.

Lesson 1: 3-day ticks did not hold. Too much changed per tick, one small
blocker consumed a cycle, and reviewing an agent PR took longer than the agent
took to write it. Agents moved the constraint from writing code to reviewing
it and the cycle was sized for the old constraint. Now 7-day cycles on Jira.

Lesson 2 (AI in SE): agents own ticket drafting, docs and PR comments. Speed
is 3 min -> 15 s, about 11 hours over 234 summer tickets, which is under an
hour a week and is deliberately not the headline. The finding that matters is
field completeness: of 56 hand-written spring tickets, 0% had story points and
0% had an epic parent; of 234 agent-drafted tickets, 90% have points and 94%
have the right epic. That is what makes Monte Carlo forecasting possible at
all - it needs points on every ticket.

Numbers derived from dashboard/data/jira_issues.json (290 issues; JQL and
fetch timestamp recorded in the file).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Reflection is presented in the closing slot at the end of the team talk, not as
part of Software System, so it should not be slide 5 of that deck.

build_section3_deck.py now builds both files from one script, so the palette,
grid and 20 pt projection floor cannot drift between them:
  eParts_Section3_SoftwareSystem.pptx  4 slides
  eParts_Reflection_Closing.pptx       1 slide

Factored out new_deck() and save(); save() asserts the expected slide count and
repaints theme hyperlink colours, which python-pptx cannot set and which the
default Office theme forces to blue.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…imings

Slide 4's link now goes to ADR-018 (routing with class-review-first) rather than
the index, because that is the ADR we open live. It is the right exemplar: it is
where the class-error argument from slide 3 becomes a concrete design rule, so it
pays off something the audience has already heard.

Added a 31-second walkthrough script for it - context, rejected alternatives,
the signal table and its governing rule, consequences, traceability. Cue is
scroll-don't-read: the point is the shape of the artifact.

Timing correction. Earlier word counts in this file were estimated from the whole
document rather than the spoken lines, and were wrong - the script was ~1,600
spoken words, about 10 minutes, against a 5-minute slot. Trimmed to 786 spoken
words (5:14) and every per-section heading now carries its measured time. The
over-time table gives measured savings per cut instead of guesses.

Reflection script measured too: 334 words, 2:13.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Slide 1 was doing four things badly. It is now the comparison it should have
been: a v1.0 -> v1.3 table (high-level 5->6, functional 8->10, derived 3->4,
constraints 3->4, total 19->24) beside the two numbers that matter - 89% of the
baseline unchanged, 37% churn.

Counts verified from the LaTeX sources rather than asserted: v1.1 and v1.3 carry
HLR-6, FR-10, DR-4, so v1.0 is those minus the four IDs v1.1 introduced; C-4
arrived in v1.2. Five IDs added, two rewritten (HLR-1, HLR-2), 17 of 19
untouched.

Palette collapsed to black, white and grey. GREEN/AMBER are gone, replaced by
ACCENT (white), DIM (grey) and RULE (hairline), so no slide carries a hue. The
theme's stock Office accents are greyed on save as well, otherwise a themed
shape inserted later would reintroduce orange. Verified: zero non-grey colours
across both decks' slides and themes.

Diagram thumbnails on slide 3 are greyscale copies so the slide stays mono; the
full-colour v6 PNG remains the artifact opened live, where its legend needs hue.

Also fixed slide 2's banner, which still read 1.0 -> 1.1 -> 1.2.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…locked-on-client card

Slide changes driven by review:
- Artifact links move from the header to the footer, where the audience looks for
  them: evidence strip on line one, 'Label: url' on line two, hyperlinked to the
  full target but showing a short navigable form.
- Five deltas reworded out of jargon. 'Frozen ML handoff' meant nothing to a
  reader; it is now 'Fixed handoff format to ML'. Likewise 'ETIM reference layer'
  -> 'ETIM dictionary loaded', 'Evidence vs. interpretation' -> 'Raw values kept
  separately', 'PIMS re-keyed on ETIM' -> 'PIMS keyed by ETIM IDs'.
- Slide 2's 'Blocked on the client' card removed; the slide is now one full-width
  card describing what we actually did.
- Slide 4 rewritten as plain tradeoffs instead of ADR bookkeeping: 'ML does the
  matching', 'Raw supplier values kept in their own table', 'New ADRs, not edits
  to the old ones', 'Stay on ETIM 10.0'.
- Slide 3 thumbnails bottom-aligned so the two captions line up; row grid on
  slide 4 tightened so the last row no longer collides with the banner.

Script synced to match.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every slide lost a block. Crowding was the problem, not missing content - the
sentences belong in the script, not on the screen.

- Slide 1: banner removed. Title, the v1.0/v1.3 table, and the two numbers.
- Slide 3: banner removed; the solid/dashed key moved into the footer, so the
  two diagrams and the change list get the space. 'Five deltas' -> 'Five
  changes' - delta is not a word this audience needs to decode.
- Slide 4: trailing 'five client decisions still open' banner removed. Four
  rows, nothing else; the open decisions are spoken instead.
- Slide 5: more air between the two cards.
- Row grids and card heights retuned so nothing sits within 20 pt of the footer.

Script synced: the blocked-on-client paragraph is gone with its card, and the
open client decisions moved into slide 4's spoken lines where they now live.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The title promised what we chose against, and the slide gave benefits instead -
the right column restated the decision rather than naming the road not taken.
It is now a real two-column comparison, headed 'We chose' and 'Instead of',
with a small multiply glyph between them:

  ML does the matching          x  ETIM keys at normalization
  Raw values in own table       x  one table holding both
  New ADRs for new decisions    x  editing the April ones
  Stay on ETIM 10.0             x  building an upgrade path

A rejected alternative is more informative than a benefit, and it is what the
rubric means by justifying a decision. Every row is one line at 20 pt, and the
grid clears the footer.

Script synced; the open client decisions fold into row 4's spoken lines.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The old script was written in fragments and rhetorical beats - 'Requirements
first.', 'Five things changed. Two matter.', setup-then-payoff pairs, stage
directions wedged inside sentences. It read as performance rather than
explanation, and it did not sound like anyone talking.

Rewritten as full sentences someone would actually say, with cues moved onto
their own lines outside the quotes. Verified mechanically: every spoken line
ends in terminal punctuation, so there are no half-lines left.

Renamed to talking_script.md and measured from the blockquotes: 825 words,
about 5:30. The cut table gives measured savings per paragraph to get under
five minutes.

Content also re-synced to the current four slides, which changed since the last
script: slide 1 is now the v1.0/v1.3 comparison, slide 2 lost the
blocked-on-client card, and slide 4 pairs each decision with the alternative
rejected rather than with its benefit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The section covered requirements and architecture but never mentioned
quality-attribute scenarios, tests, or the risk register — three artifacts
the Software System rubric names explicitly — and it asserted that trace
links survived without ever walking one.

Spec v1.4 (QAS-3, VAL-4, VAL-5). The v1.1 ETIM integration never revised
the Quality Attribute Scenarios or Validation Requirements sections, so for
a month those two described a pre-ETIM system. QAS-3 holds ADR-019's
policy-as-configuration decision; VAL-4 documents the ETIM dictionary load
that automated tests already cover; VAL-5 specifies class-review-first and
is marked not-yet-executable because the matching stages are not built.

Also fixed three descriptions the v1.3 correction left behind — the
Canonical Table glossary entry and §1.2 both still implied ETIM keying
during normalization, and §1.2 did not mention ETIM matching at all — plus
a stale SCEN-1 and HLR-2 reference inside ADR-016. The §2 figure now uses
the current v6.0 diagram, so the spec and the slides show one architecture.

QAS-3 collides by number with the v2.0 lineage's QAS-3 (Accuracy). It is
the next free number in this document; the collision is recorded in the
changelog and the mapping matrix rather than dodged by skipping a number.

Slide 1 now compares all six requirement classes (24 → 32, 92% of the April
set untouched) instead of four, and slide 2 carries the full trace chain
HLR-6 → FR-9 → ADR-16 → migration 0005 → 11 tests.

Talking script rewritten: the list-reading and the unmotivated
open-client-decisions line are gone, and the space went to the trace thread,
the test evidence, and RISK-ARCH-09.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ran the suite. tests/unit/test_etim_loader.py is 10 passed in 3m43s.
tests/integration/test_etim_real_files.py skips unconditionally on a clean
checkout — it guards on .tmp_etim_csv, which .gitignore excludes — so the
earlier "eleven tests" would not have survived anyone running it.

Slide 2's footer, the script, spec §5.2, ADR-016 and both trace matrices now
say ten passing plus one that skips, and the script's delivery notes name the
expected "10 passed, 1 skipped" so the skip gets volunteered rather than
discovered. Script re-measured at 892 words / 5:56; the cut table needs three
cuts rather than two to reach five minutes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three changes Arjun asked for.

Slide 1 loses the Quality scenarios and Validation tests rows, back to the
four requirement classes and 19 -> 24, 89% / 37%. QAS-3, VAL-4 and VAL-5 stay
in the spec and the trace matrix; he narrates them instead of adding rows.

Slide 2's footer said "migration 0005", which means nothing to an assessor and
nothing to the presenter either. It now reads "ETIM tables in code", and the
script says the same in words.

New slide 4 fixes the unreadable diagram, and not by scaling it. The v6 drawing
is portrait 1600x1898; on a 720x405 slide at full height it is 253 pt wide and
its stage labels land near 1.4 pt. Cropping doesn't help — reaching ~16 pt
labels means fitting only ~40% of the diagram's width. So the changed region is
redrawn as native shapes at 20 pt: the existing attribute-matching pass, then
the five new ETIM stages, dashed because none of them are built. Slide 3 keeps
the thumbnails, whose only job is silhouette comparison, and the script now
says out loud that they aren't meant to be readable.

Deck goes 4 -> 5 slides. Script restructured to match and re-measured at 866
words / 5:46, with the first two cuts landing on 5:00.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Keeps his voice — shorter declaratives, fewer hedges ("I'll walk one", "The
order matters") — and fixes the four things that moved underneath it:

- slide 1 lost the six-classes phrasing and 92% goes back to 89%, since the
  Quality scenarios and Validation tests rows came off that table
- "migration 0005" becomes "exist in the code as the ten ETIM reference tables"
- slide 3 splits: comparison and the staging split stay, change 2 moves to the
  new slide 4, and the solid/dashed note goes with it because that is the slide
  where the dashes are actually visible
- decisions slide renumbered to 5, and the stability answer now says 89%

846 words / 5:38 measured, no fragments.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The v6 diagram had become a rewrite rather than an addition. It had dropped the
staging/canonical store, the whole rejection column (rejected items log, audit
trail, catalog team alert, correction store), and replaced the single matching
filter with five sub-boxes, a client-feature-policy box and an evidence/
interpretation pair. That buries the actual story, which is that ETIM changed one
thing structurally.

docs/build_pipe_filter_v6.py transcribes v5.0 from its React source and adds
exactly one filter — "ETIM matching" — behind "ML / AI attribute matching",
inside the same PredictionServiceInterface boundary. Every other element keeps
v5's own coordinates; a Y() helper applies the 92 px shift the new box needs so
nothing downstream was retyped. The removed boxes are gone and the ADR
annotations with them.

Knock-on fixes: the two thumbnails on slide 3 are now near-identical in aspect
(1.09 vs 1.01) so they sit at the same height and read as one drawing; the card
becomes "What ETIM changed" with the new phase as item 1; the footer says "one
new box, nothing else moved"; slide 4 is retitled "Inside the new box"; and the
spec's §2 figure is regenerated from the new PNG.

Script slide 3 rewritten to the one-box story — five changes, only the first
visible on the drawing. 843 words / 5:37.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Dropping a .md into the Confluence editor attaches it as a download card, which
is what the blank page was showing. Real page content comes from pasting the
text — the editor converts Markdown on paste.

docs/build_confluence_adrs.py emits docs/confluence/: one ADRs-all.md holding
every ADR under an index table, an ADRs-index.md with just the table, and 21
per-ADR files for child pages. Two transforms are needed or the paste breaks:
headings are demoted one level so each ADR nests under the page title instead of
competing with it, and relative links to sibling .md files are rewritten to
GitHub URLs, since those targets do not exist on Confluence.

Every page says the repo is the source of truth, so the copies are read-only by
convention rather than a second place to edit.

Not uploaded: the Atlassian connector only grants epartsmse.atlassian.net, and
the crit page lives on cmu-mse.atlassian.net.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
"Get the class wrong and everything under it is wrong too" asserted a conclusion
without the mechanism, so it read as a slogan. The mechanism is that ETIM
features belong to the class — etim_class_feature is the list of features a class
is allowed to have — so a wrong class means every attribute is matched against
the wrong candidate list.

Slide line becomes "Wrong class, wrong feature list. Every attribute under it is
wrong", which names the cause rather than only the effect. The script gets the
concrete case: a ball valve classified as a butterfly valve matches its torque
against a butterfly-valve feature, so the number is right and the feature is
wrong, on every attribute. And the reason that is dangerous rather than merely
wrong — each match scores high, because it was the best match in the list it was
given, so routing auto-accepts instead of flagging.

Slide 4 grows to 216 words / 86 s and the section is now 894 / 5:57. Left long on
purpose; the cut table needs all four cuts to reach 5:02 and says so.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
v4 measures exactly as its table claims — 868 words / 5:47, no fragments — and
the IDs, percentages and footers all match the built slides. Adopted as-is except
for one factual conflict: slide 3 called staging unchanged and then listed the
evidence/interpretation split, which is a staging change, among the four changes
that don't show on the drawing. The audit trail genuinely is untouched, so the
claim now stops there. 866 words / 5:46.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four claims become four artifact references, which is what "architectural
descriptions and decisions" asks for. ML owns matching is ADR-016, raw values in
their own table is ADR-014, staying on 10.0 is ADR-020. Row 3 has no single ADR —
the artifacts are 16 through 21 — so it cites the range instead of inventing a
number.

Two things had to give to make room. "New ADRs for new decisions" is now "New
ADRs, not edits": the column lost 44 pt to the ADR column and the longer phrase
wrapped. And a wrap=False option was added to text(), though LibreOffice ignores
it and substitutes a wider font for Aptos, so the columns are sized for the
substituted metrics rather than trusting the flag.

The ADR-018 footer now points at Confluence rather than GitHub, since the ADRs are
published there now — the room lands on a rendered page instead of raw Markdown.
Shown URL trimmed to the space root so it matches slides 1 and 3 and stops
overflowing the strip.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants