diff --git a/README.md b/README.md
index 2d0b936..c6a48bf 100644
--- a/README.md
+++ b/README.md
@@ -139,7 +139,7 @@ Our team's operational playbook for using Bob — prompt patterns, export workfl
| **Frontend** | Next.js 14 (App Router) + TypeScript + Tailwind + shadcn/ui + Recharts | Modern React, fast static + SSR, great DX for visualizations |
| **Backend** | FastAPI + Python 3.11 + Pydantic v2 + SQLAlchemy 2 + asyncpg | Async by default, type-safe, fits a Python AI pipeline natively |
| **AI Pipeline** | spaCy 3.7 (`en_core_web_lg`) + sentence-transformers (`all-mpnet-base-v2`) + scikit-learn + umap-learn | Industry-standard English NLP; strong 768-dim semantic embeddings; reproducible |
-| **LLM** | IBM Watsonx (Llama 3.3 70B + Granite 3 8B) | End-to-end IBM stack; full-IBM points; strong creative English generation |
+| **LLM** | IBM Watsonx (`meta-llama/llama-3-3-70b-instruct`) — sole model in use; `ibm/granite-4-h-small` declared as an unexercised fallback | Every generation runs on IBM Watsonx; strong creative English generation |
| **Database** | PostgreSQL 16 + pgvector (Supabase) | Single DB for relational + vector; HNSW index for fast RAG |
| **Crypto** | python-jose, ES256 (ECDSA P-256) | Standard JWS; small signatures; native browser verification |
| **Hosting** | Vercel (frontend) + Railway (backend) + Supabase (DB) | Zero-ops, free or near-free tiers, push-to-deploy |
diff --git a/backend/tests/test_watsonx_client.py b/backend/tests/test_watsonx_client.py
index 3464f09..ef0efeb 100644
--- a/backend/tests/test_watsonx_client.py
+++ b/backend/tests/test_watsonx_client.py
@@ -96,7 +96,7 @@ def test_generate_retries_with_backoff_then_succeeds(monkeypatch):
patch("app.services.watsonx_client.Credentials"),
patch("app.services.watsonx_client.time.sleep", side_effect=sleeps.append),
):
- text = generate("prompt", None, "ibm/granite-3-8b-instruct")
+ text = generate("prompt", None, "meta-llama/llama-3-3-70b-instruct")
assert text == "ok"
assert sleeps == [1.0]
@@ -118,7 +118,7 @@ def test_generate_exhausts_retries(monkeypatch):
patch("app.services.watsonx_client.time.sleep", side_effect=sleeps.append),
pytest.raises(WatsonxError, match="failed after"),
):
- generate("prompt", None, "ibm/granite-3-8b-instruct")
+ generate("prompt", None, "meta-llama/llama-3-3-70b-instruct")
assert sleeps == list(_RETRY_DELAYS_SECONDS)
assert mock_model.chat.call_count == 1 + len(_RETRY_DELAYS_SECONDS)
@@ -142,7 +142,7 @@ def test_generate_hard_timeout(monkeypatch):
future = MagicMock()
future.result.side_effect = FuturesTimeoutError()
executor.submit.return_value = future
- generate("prompt", None, "ibm/granite-3-8b-instruct")
+ generate("prompt", None, "meta-llama/llama-3-3-70b-instruct")
assert sleeps == list(_RETRY_DELAYS_SECONDS)
assert future.result.call_count == 1 + len(_RETRY_DELAYS_SECONDS)
@@ -161,7 +161,7 @@ def test_generate_live_watsonx():
text = generate(
prompt="Reply with exactly one word: pong",
system_prompt="You are a terse assistant. Answer with a single word only.",
- model_id="ibm/granite-3-8b-instruct",
+ model_id="meta-llama/llama-3-3-70b-instruct",
params={"max_tokens": 16, "temperature": 0},
)
assert isinstance(text, str)
diff --git a/docs/MVP.md b/docs/MVP.md
index d79d333..d83fd24 100644
--- a/docs/MVP.md
+++ b/docs/MVP.md
@@ -51,8 +51,9 @@ These three decisions shape everything else. They are locked; changing one is a
- **LLM: IBM Watsonx, end-to-end.**
- Primary (creative writing): `meta-llama/llama-3-3-70b-instruct`
- - Auxiliary (structured/classification tasks): `ibm/granite-3-8b-instruct`
- - Baseline ("vanilla"): the **same** `llama-3-3-70b` but **without** style conditioning.
+ - Auxiliary: **none**. `ibm/granite-3-8b-instruct` was originally locked here for structured/classification tasks, but it was never wired into any code path and it is **not available to this project**: it is absent from the `eu-de` catalogue and, where it was still listed (`us-south`), its lifecycle is `withdrawn` since 2026-03-31. Revised 2026-07-28 — see `docs/decision_log.md`. Every call the product makes goes to the primary model.
+ - Escalation backup (R1), **not in use**: `ibm/granite-4-h-small` — confirmed `available` in `eu-de` on 2026-07-28. Nothing calls it today; it exists only as the named fallback if the Sprint 1 voice-match gate fails (see §11 R1). Adopting it would be a separate decision.
+ - Baseline ("vanilla"): the **same** `llama-3-3-70b` but **without** style conditioning. The escalation backup never touches this comparison.
- *Why same model on both sides:* it's an honest comparison. The only variable is our style conditioning — judges notice and respect that.
- **Language: everything in English — UI, docs, video, and the generated text.**
@@ -324,8 +325,9 @@ Each is justified so we don't reopen the discussion mid-sprint.
### LLM
- `ibm-watsonx-ai` SDK
-- Creative: `meta-llama/llama-3-3-70b-instruct`
-- Auxiliary: `ibm/granite-3-8b-instruct`
+- Creative: `meta-llama/llama-3-3-70b-instruct` — the only model the product calls
+- Auxiliary: none (the originally locked `ibm/granite-3-8b-instruct` is unavailable; see §2)
+- Escalation backup, **not wired in**: `ibm/granite-4-h-small` (available in `eu-de`, ratified 2026-07-28). No code path, no env var, no call — only R1's named fallback
- **SPRINT 1 TASK:** validate voice-matching quality with 5 real prompts (does the conditioned output read like the target author?); if Llama < 6/10 human eval, escalate (see R1).
### Database
@@ -509,9 +511,9 @@ The project is "done" when **all** of these are true:
| # | Risk | Mitigation |
|---|---|---|
-| **R1** | Llama-3.3-70b doesn't convincingly match a target voice | Validate Sprint 1 day 2. If <6/10, plan B: stronger conditioning / more RAG passages, `llama-3-1-405b`, or `granite-3-8b`; last resort Mistral Large via Watsonx. |
+| **R1** | Llama-3.3-70b doesn't convincingly match a target voice | Validate Sprint 1 day 2. If <6/10, plan B: stronger conditioning / more RAG passages first. **Model-swap escalation: `ibm/granite-4-h-small`** — ratified 2026-07-28, confirmed `available` in `eu-de` that day. The three backups this row used to name (`llama-3-1-405b`, `granite-3-8b`, Mistral Large) are all unreachable from `eu-de` and were removed. The escalation is **declared, not exercised**: nothing calls Granite 4 today, and it never becomes the vanilla baseline — the A/B stays `llama-3-3-70b` vs itself. |
| **R2** | Stylistic metrics don't distinguish authors | Validate Sprint 1 with a blind A/B. If Austen ≈ Dickens (both 19thc British), add bigram/trigram distinctive features. (Poe is the easy separation.) |
-| **R3** | Generation >15s, demo doesn't flow | P3 measures latency Sprint 2 day 1. If >10s: switch parallel→sequential with optimistic loading, and/or use `granite-8b` for the baseline (faster). |
+| **R3** | Generation >15s, demo doesn't flow | P3 measures latency Sprint 2 day 1. If >10s: switch parallel→sequential with optimistic loading, shorten `max_tokens`, or fall back to the pre-recorded generations of R5. **R3 has no model-swap plan B and no replacement for one.** The old "use a faster `granite-8b` for the baseline" is void twice over: that model does not exist for us (see §2), *and* swapping only the baseline would break the A/B whatever model replaced it — the comparison is only honest while both columns run the same model. R1's `ibm/granite-4-h-small` is **not** an R3 mitigation for the same reason. The three levers above are the whole of R3. |
| **R4** | Live demo fails during recording | Sprint 3: record each step separately as an editable fallback. Final video may mix live + pre-recorded. |
| **R5** | Watsonx rate-limit / load spike | Exponential backoff (Sprint 2). For the demo, pre-record two famous generations as a local cache fallback. |
| **R6** | A teammate is sick for a week | Each role has a designated backup; pair on critical pieces; the async daily keeps the backup current. |
diff --git a/docs/ONBOARDING.md b/docs/ONBOARDING.md
index fcfc646..f46cc98 100644
--- a/docs/ONBOARDING.md
+++ b/docs/ONBOARDING.md
@@ -229,7 +229,7 @@ Lives at **`docs/decision_log.md`**. 4-column table:
Read them in `docs/decision_log.md`. The key ones to keep in your head:
- Idea = **AutorIA** (auditable authorship layer)
-- LLM = **IBM Watsonx end-to-end** (Llama 3.3 70B + Granite 3 8B auxiliary)
+- LLM = **IBM Watsonx end-to-end** (`meta-llama/llama-3-3-70b-instruct` is the only model we call; the Granite 3 8B auxiliary was dropped 2026-07-28 — withdrawn upstream, absent from `eu-de`. `ibm/granite-4-h-small` is the *declared, unused* R1 escalation backup, not part of the active path)
- **English everywhere** — app UI, repo, README, video, commits, and the generated text. Only the **internal team chat** stays Spanish.
- Multi-author with **3 preloaded**: Jane Austen, Charles Dickens, Edgar Allan Poe
- Stack: Next.js + FastAPI + Postgres+pgvector (Supabase) + JWS ES256
@@ -326,7 +326,7 @@ Full document at **`docs/MVP.md`**. Executive summary:
| Decision | Value | Why |
|---|---|---|
-| LLM | **IBM Watsonx end-to-end**. Primary: `meta-llama/llama-3-3-70b-instruct`. Auxiliary: `ibm/granite-3-8b-instruct` | IBM challenge = points for full-IBM stack |
+| LLM | **IBM Watsonx end-to-end**. Primary and only model in use: `meta-llama/llama-3-3-70b-instruct`. No auxiliary model (`ibm/granite-3-8b-instruct` was revised out on 2026-07-28 — unavailable in `eu-de`, withdrawn upstream). `ibm/granite-4-h-small` is the ratified R1 escalation backup — declared, never called | IBM challenge = points for running the whole product on Watsonx |
| Comparison baseline | **Same model WITHOUT style conditioning** | Honest comparison = judges notice |
| App UI + video + generated text | **English** | International IBM jury; judges can read the voice match directly |
| Authors | **Austen, Dickens, Poe** (English, public domain) | Iconic, maximally distinct voices; instantly recognizable |
@@ -370,7 +370,8 @@ tiktoken (approx chunking)
```
ibm-watsonx-ai SDK
Creative model: meta-llama/llama-3-3-70b-instruct
-Auxiliary model: ibm/granite-3-8b-instruct
+Auxiliary model: none (Granite 3 8B revised out 2026-07-28 — unavailable in eu-de)
+R1 escalation backup (declared, not called): ibm/granite-4-h-small
Sprint 1 task: validate voice-matching quality with 5 prompts (gate)
```
@@ -961,7 +962,7 @@ On Sunday, unfinished issues either move to "Sprint N+1" or back to "Backlog" if
Your role has a backup designated (P1→P3, P2→P3, P3→P2). The daily async ensures your backup knows the state. Pair sessions ensure >1 person has touched critical pieces. Not ideal but the project doesn't collapse.
### "What if Watsonx doesn't match the author's voice well?"
-Plan B in MVP: stronger conditioning / more RAG passages, then `llama-3-1-405b-instruct` or `granite-3-8b-instruct` via Watsonx. If all fail, fallback to Mistral Large via Watsonx. **But**: we validate this Sprint 1 day 2, not Sprint 4.
+Plan B in MVP: stronger conditioning / more RAG passages first, then swap the model to **`ibm/granite-4-h-small`** (ratified 2026-07-28, confirmed available in `eu-de`). The escalation chain this FAQ used to name — `llama-3-1-405b-instruct`, `granite-3-8b-instruct`, Mistral Large — was fiction: none of the three is in the `eu-de` catalogue we run against, and Granite 3 8B is withdrawn upstream everywhere. Granite 4 H Small is a *declared* backup: we do not call it today, and it never replaces the vanilla baseline (that would break the A/B). **But**: we validate this Sprint 1 day 2, not Sprint 4.
### "What if a seemingly essential feature emerges that's not in the MVP?"
5-min call with all 3, 2/3 vote. If it passes → Decision Log entry + add. If not → `docs/roadmap.md` for "post-July".
diff --git a/docs/architecture.md b/docs/architecture.md
index e254f6c..7f0553d 100644
--- a/docs/architecture.md
+++ b/docs/architecture.md
@@ -36,7 +36,7 @@ graph TB
sys[["AutorIA
Style DNA · Conditioned Generation ·
Authorship Passport"]]
end
- watsonx[("IBM Watsonx
Llama 3.3 70B + Granite 3 8B
[external LLM]")]
+ watsonx[("IBM Watsonx
Llama 3.3 70B
[external LLM]")]
gutenberg[("Project Gutenberg
public-domain corpus
[external source]")]
creator -->|"selects author, prompts,
views side-by-side, downloads Passport"| sys
@@ -55,8 +55,11 @@ graph TB
audience). Only needs the token + public JWKS.
**External dependencies**
-- **IBM Watsonx** — the LLM provider (creative + auxiliary models). The only
- hard runtime external dependency for generation.
+- **IBM Watsonx** — the LLM provider. A single model, `meta-llama/llama-3-3-70b-instruct`,
+ serves both branches of the side-by-side (conditioned and vanilla); there is no
+ auxiliary model. `ibm/granite-4-h-small` is the declared R1 escalation backup and is
+ never called on the active path (MVP §2 / §11). The only hard runtime external
+ dependency for generation.
- **Project Gutenberg** — source of the three preloaded corpora; used **offline**
during seeding, not at request time.
diff --git a/docs/decision_log.md b/docs/decision_log.md
index 97232c0..6d372df 100644
--- a/docs/decision_log.md
+++ b/docs/decision_log.md
@@ -50,3 +50,4 @@ Every decision that affects the product, the process, or the team lives here. Ap
| 2026-07-28 | **`distinctive_vocab` excludes proper nouns, and its corpus is sampled instead of truncated** (#100, WO-18). (a) The 800k-lemma-char memory cap in `style_profile.py` no longer keeps a *prefix* of the corpus: chunks are fed to spaCy in a deterministic bisection order, so the capped sample spans every document of every author. Before the fix the cap fell at 21% of Dickens' chunks (34% Austen, 89% Poe), so "each author's full corpus is one document" (`docs/style_features.md` §4.1) was really *Great Expectations* alone. (b) Tokens tagged `PROPN` by spaCy are dropped **inside** the lemmatization pass (`_lemmas_from_docs`), not by a blacklist after the TF-IDF. `docs/style_features.md` §4.1 updated to match. Re-seeding is required (#86): every existing `style_profiles` row is invalidated. | Sergi | Character names are statistically distinctive but they are plot, not style — they identify the novel, not the author's hand. §4.1's own purpose is that "the jury can read and feel the difference", and its own example (`countenance`, `physiognomy`, `presently`) is a style marker, not a cast list; the demo is judged by what a non-technical human perceives. Filtering at the POS level is free (the tag is already computed) and stays correct for authors added later, which a hand-maintained word blacklist would not. |
| 2026-07-28 | **Exception accepted: #100 ships with a top-10 `distinctive_vocab` overlap of 5, not the ≤3 its Definition of Done requires.** The DoD of #100 states verbatim: *"top-10 overlap across the three authors stays ≤3 (#89 already brought it down from 6 to 3; it must not get worse)"*. **Measured after the fix: 5** — `know`, `little`, `make`, `say`, `time` are common to all three top-10 lists (pairwise: austen∩dickens 8, austen∩poe 6, dickens∩poe 6). The criterion is **not met and is worse than what #89 left**. The full measured top-10 per author is recorded in `docs/style_features.md` §4.1 ("Measured output"), together with the honest reading: the feature returns ordinary high-frequency verbs and nouns, not signature vocabulary, and none of `countenance` / `physiognomy` / `presently` reaches any author's top-30. The `PROPN` filter itself is confirmed working (no proper noun survives in any top-30) and is now covered by a regression test (`ai_pipeline/tests/test_style_profile_compute.py::test_lemmatize_corpus_drops_proper_nouns`, verified to fail when the filter is removed). Ratification: entered by the executor of #100 on measured evidence and accepted by the project owner; **still needs the 2/3 vote** the decision policy of 2026-06-24 requires | Sergi (project owner) — pending 2/3 ratification | With only three documents in the collection, every term present in all three has `df = 3/3` and therefore identical idf (1.000), so TF-IDF degenerates into raw frequency: `countenance` (29 occurrences in Dickens) and `say` (1 866) are idf-tied and the common word always wins. Proper nouns were the only terms with a discriminating `df`, so the overlap of 3 that #89 measured was an artefact of the cast lists being counted as style. Dropping them (the correct call — character names are plot, not the author's hand) necessarily raised the overlap to 5. Two alternatives were considered and rejected for this issue: (a) redefining the scoring (log-odds, or idf against a general-English reference corpus) would change the algorithm `docs/style_features.md` §4.1 declares as closed, which is a scope decision and not a fix to #100; (b) reverting the `PROPN` filter would restore the ≤3 number by putting `pip`, `havisham` and `wemmick` back in front of the jury as "style", i.e. buying the metric with the defect the issue exists to remove. A dull-but-honest list is preferred to a varied-but-wrong one; the gap is documented rather than hidden |
| 2026-07-28 | **`api_contract.yaml` descriptions corrected for `dialogue_ratio` and `first_person_ratio` (prose only — NO contract amendment).** Both were described as *"Fraction of sentences…"*. Neither is. `dialogue_ratio` is the fraction of non-punctuation **tokens** inside straight ASCII double quotes (`docs/style_features.md` §3.3, `dialogue_tokens / total_tokens`; implemented that way in `ai_pipeline/autoria_ai/extractor/stylistic.py`). `first_person_ratio` is first-person singular pronouns **per 1,000 tokens** (§3.4), so it is **not bounded by 1** — Poe's expected range is 18–30. The declared shape (`type: number`) is unchanged in both cases, so no client or server contract breaks; only the prose changed | Sergi (repo owner) | Found while implementing #92: three fixture `first_person_ratio` values were written as fractions ([0,1]) because the contract described them that way, making them wrong by a factor of ~45 against the ranges in §7. The LOCKED policy (2026-06-24) requires a log entry for any change to this file; recorded here as a documentation correction, not a scope change — the contract never matched the implementation or the feature spec, and a reader building a UI against it would size the widget for 0–1 |
+| 2026-07-28 | **Granite 3 8B dropped from the stack, and `ibm/granite-4-h-small` named as the R1 escalation model — RATIFIED** (supersedes the LLM stack entry of 2026-06-24, which locked `ibm/granite-3-8b-instruct` as the auxiliary model in `docs/MVP.md` §2). Two linked calls, one vote. **(a) The active LLM path is single-model:** IBM Watsonx with `meta-llama/llama-3-3-70b-instruct` as the only model the product calls; there is no auxiliary model. Granite 3 8B was never wired into any code path (zero occurrences in `backend/app`, `ai_pipeline`, `frontend/src`, `scripts`, `.env.example`); its only caller was the live integration test `backend/tests/test_watsonx_client.py::test_generate_live_watsonx`, which now targets the model the product actually uses. **(b) The R1 escalation model is `ibm/granite-4-h-small`** (resolves the open question of #102, which found all three previously advertised backups unreachable). It is a **declared but unexercised backup**: no code calls it, no configuration points at it, and it does **not** enter the side-by-side — the A/B compares `meta-llama/llama-3-3-70b-instruct` against itself, with and without style conditioning, and that does not change. **R3 gets no replacement plan B, and that is stated rather than papered over:** its old mitigation was "use the faster `granite-8b` for the baseline", which is void twice over — the model does not exist, and swapping only the baseline would break the A/B honesty rule no matter which model replaced it. What remains for R3 is what §11 now says and nothing more: parallel→sequential generation with optimistic loading, a shorter `max_tokens`, and the pre-recorded generations of R5. Reconciled in the same pass: `docs/MVP.md` §2/§6/§11, `README.md`, `docs/architecture.md`, `docs/ONBOARDING.md` | Sergi Torres (P1, repo owner) — 2/3 vote cast 2026-07-28, ratifying the entry drafted by the executor of #98 | A stack claim we cannot demonstrate live is a liability in front of the IBM jury, and an escalation chain where every exit is closed is worse than an honestly declared single route. Measured 2026-07-28 against the public Watsonx catalogue, no credentials required — `GET https://eu-de.ml.cloud.ibm.com/ml/v1/foundation_model_specs?version=2024-05-01&limit=200`, HTTP 200, 18 models listed for `eu-de`: `ibm/granite-4-h-small` → `available` (lifecycle `available` since 2025-10-02, provider IBM); `meta-llama/llama-3-3-70b-instruct` → `available`; `ibm/granite-3-8b-instruct` → **not present in `eu-de` at all**. Granite 3 8B is withdrawn upstream since 2026-03-31, so no region change recovers it; matching runtime evidence 2026-07-27 in `eu-de`: `meta-llama/llama-3-3-70b-instruct` → `pong`, `ibm/granite-3-8b-instruct` → `404 model_not_supported`. Granite 4 H Small was chosen over the other confirmed-reachable candidate (`meta-llama/llama-4-maverick-17b-128e-instruct-fp8`) because the Watsonx catalogue points to it as the successor of the withdrawn Granite, and because it keeps the full-IBM angle the challenge rewards. It is **named, not adopted**: declaring a verified-reachable escape route costs nothing and answers the jury's "what if the model fails?" honestly, whereas wiring a second model into the demo week would add untested surface for a risk that has not materialised |