Paste a message you are not sure about. Unhook shows you the exact words being used to manipulate you — and tells you what to do next.
Older adults lose more money to fraud than any other group.
These messages do not arrive as spam. They come as a text from a real phone number, a WhatsApp from someone who seems to know you, a call from a number that matches your bank's. No filter ever sees them. The person is the filter.
And every tool built to help answers the same question — is this a scam, yes or no. You get a verdict. It may save you this once, but you learn nothing about how it worked on you, so the next scam catches you just the same.
The question a frightened person actually has is "why should I not trust this?" — and after that, "what do I do now?"
Unhook answers those two. It marks the specific words doing the manipulating, names the technique, and gives one instruction. The goal is that the reader starts spotting it themselves.
Paste a text, email, or message. Ten seconds later:
| A verdict in plain words | "This looks like a scam" — not "Risk: 87/100" |
| The tells, marked in your own message | Tap any highlighted phrase to see why it is there |
| One instruction | "Call your grandchild on their usual number, and tell their parents." |
| What happens next | The scam's next three moves, projected from the tactics actually found |
| Read aloud | For the reader whose eyesight is why they were unsure in the first place |
| Show someone you trust | Reveals the exact text before you copy it — nothing is sent by Unhook |
| Answers in the message's language | Fraud targets people in their first language; a warning you have to translate is one you won't act on |
Five real examples are one click away and answer instantly.
┌──────────────────────────────────────┐
pasted message → │ 1. normalise │
│ strip zero-width, NFKC, homograph │
└──────────────┬───────────────────────┘
│
┌────────────────────┴────────────────────┐
▼ ▼
┌──────────────────────┐ ┌──────────────────────────┐
│ 2a. deterministic │ │ 2b. model, in two passes │
│ ~2 ms, free │ (these two │ ① evidence: which of 15 │
│ 16 signal types: │ layers │ tactics, exact quote │
│ lookalike domains, │ never see │ ② narration: say it in │
│ crypto addresses, │ each │ plain words │
│ gift-card requests, │ other) │ │
│ urgency markers… │ │ never returns a score │
└──────────┬───────────┘ └────────────┬─────────────┘
│ │
│ ┌────────────▼─────────────┐
│ │ 3. span resolver │
│ │ quote → exact offsets │
│ │ 3 tiers, testable │
│ └────────────┬─────────────┘
└────────────────────┬────────────────────┘
▼
┌──────────────────────────────────────┐
│ 4. risk rubric — published constants │
│ superlinear severity, no double │
│ counting, combination bonuses │
└──────────────┬───────────────────────┘
▼
the analysis
The model never returns a score. It returns evidence; the number comes
from a rubric in lib/risk/rubric.ts that you can read
and disagree with. Model self-reported confidence is uncalibrated and moves when
you reword the prompt — and swapping the model would silently move every score.
The model never returns character offsets. Models quote reliably and count
characters unreliably. It returns the verbatim quote; a three-tier resolver
(lib/spans/resolve.ts) computes the offsets where it
can be unit-tested. That is what makes the highlighting trustworthy — measured
at 1.000 match quality across 166 findings.
The two detection layers are blind to each other. The deterministic detectors do not see the model's output and the model does not see theirs. That independence is what makes their disagreement informative — and the UI says so when they disagree, rather than hiding it.
Six pre-registered experiments over a hand-labelled set of 81 messages —
43 scams and 38 genuine, of which 27 are genuine messages deliberately chosen
to look like fraud (real 2FA codes, a real bank fraud alert, a hospital asking
you to call urgently). Full harness in lib/eval/, raw results in
experiments/runs/.
Accuracy is not the headline. The dataset is 53% scam, so answering "scam" every time scores 53% and is worse than useless. The numbers that decide whether this is shippable are the false-positive rate and the hard-case recall.
| pattern checks only | + language model | |
|---|---|---|
| recall | 18.6% | 97.7% |
| false positives | 0.0% | 0.0% |
| F1 | 0.314 | 0.988 |
| recall on hard cases | 9.1% | 100% |
The deterministic layer is perfectly precise and nearly blind. It scores 0/100 on the romance scam, the business-email compromise, and the "safe account" bank fraud — messages with no link, no payment rail and no urgency vocabulary. Those are pure pretext, and no pattern can reach them. This is the empirical case for the architecture, rather than an assertion about it.
| v1 taxonomy | v2 minimal | v3 few-shot | |
|---|---|---|---|
| recall | 90.7% | 95.3% | 88.4% |
| false positives | 0.0% | 19.4% | 2.7% |
| F1 | 0.951 | 0.901 | 0.927 |
Strip the taxonomy and the false-positive rules and recall goes up while the same model raises seven false alarms where it had raised none — among them a real Barclays fraud alert, scored 89/100. The prompt is not teaching the model to find scams; the schema does that. It is teaching it the discipline not to flag.
| one pass | two passes | |
|---|---|---|
| next-action reading grade | 9.3 (22 words) | 4.5 (12 words) |
| summaries at grade ≤6 | 51% | 60% |
| recall | 90.7% | 97.7% |
| cost | $0.20 | $0.24 |
Asking one call to find evidence and write for a frightened 74-year-old does both jobs worse. Splitting them moved the single most important sentence in the product from high-school reading level to primary-school, for +18% cost.
| Haiku 4.5 | Opus 5 | GPT-5.5 | |
|---|---|---|---|
| F1 | 0.977 | 0.988 | 0.938 |
| false positives | 2.6% | 0.0% | 0.0% |
| hard-case recall | 100% | 100% | 90.9% |
| tactic F1 | 0.805 | 0.875 | 0.840 |
| unresolved spans | 9.8% | 2.4% | 0.0% |
| latency p50 | 3.7 s | 10.5 s | 7.9 s |
| prompt cache hit | 0% | 87.8% | 0% |
| summaries at grade ≤6 | 60% | 94% | 98% |
Opus 5 was adopted, and the cache result is why it is affordable. Its 512-token cache minimum clears our ~1,500-token system prompt; Haiku's 4096-token minimum never engages. Same code, same prompt — the model tier decides whether the optimisation exists at all, which made the highest list-price model the cheapest arm to run.
The production model was chosen by this experiment, not before it.
Flesch–Kincaid on every line the reader sees, via npm run e5.
| before | after | |
|---|---|---|
| hand-written taxonomy copy | grade 5.8, 8 of 15 above target | grade 2.9, 0 of 15 above target |
| model next-action line | grade 9.3 | grade 4.1 |
Both numbers moved because they were measured. The taxonomy rewrite was driven by the per-tactic report; the action line was fixed by the two-pass split in E2. The lever in both cases turned out to be sentence length rather than word length.
Monotonic across all five buckets: ECE 0.145, Brier 0.072. The score is systematically under-confident — the 40–60 band was 100% scams. That is the safe direction to be wrong in for a product whose worst failure is crying wolf, but it misses the conventional 0.10 ECE bar and is reported rather than tuned away.
Not a checklist item — it is the brief. The reader is often in their seventies, on a phone, and frightened.
- Atkinson Hyperlegible, drawn by
the Braille Institute for low-vision readers, where
I l 1andO 0are deliberately differentiated - 19px base, nothing below 17px, 48px minimum touch targets
npm run contrastverifies every colour pair — all text ≥7:1 (WCAG AAA), graphical elements ≥3:1 (WCAG 1.4.11). It caught two failures that would otherwise have shipped.- Colour is never the only signal: each verdict carries a distinct icon shape
- Read-aloud, full keyboard operation,
prefers-reduced-motionrespected, zoom not capped
npm install
cp .env.example .env.local # add one API key
npm run check:keys # verifies without printing them
npm run build:demo # MUST run before build — see below
npm run build && npm start
npm run build:demoprecomputes the five demo analyses intolib/demo/cache.json. Building without it bakes an empty cache into the bundle and every example silently falls through to a live API call. There is a test guarding this, because it happened once.
| command | |
|---|---|
npm test |
122 tests |
npm run dataset |
dataset composition and tactic coverage |
npm run eval -- --id x --provider anthropic |
one experiment arm |
npm run compare -- run-a run-b |
arms side by side |
npm run e5 / npm run e6 |
readability / calibration reports |
npm run contrast |
WCAG verification |
Stated plainly, because a fraud tool that oversells itself is its own hazard.
- It can be wrong. 97.7% recall means roughly one scam in forty is missed. When money or an account is involved, the advice is always to check with the company on a number you already have.
- The interface chrome is English. The analysis — summary, action, and every explanation — comes back in whatever language the message was written in, and span highlighting works unchanged because quotes are copied verbatim from the source. But the tactic labels come from our own taxonomy and stay English, and the measured numbers below were gathered on an English test set.
- 81 cases is a small set, written from published fraud patterns rather than collected from victims. Two runs of the same configuration differed by ~2 points of recall, so single-run differences of that size are noise.
- Two tactics are under-measured —
escalating_commitmenthas only 2 labelled examples. Its per-tactic score is not meaningful. - The rate limiter is per-instance and in-memory. On serverless it is a speed bump, not a security control.
- No screenshot support, and most people receive scams as images.
MIT — see LICENSE.
The manipulation taxonomy draws on Cialdini's Influence (1984), Stajano & Wilson, Understanding scam victims: seven principles for systems security (CACM 54:3, 2011), and the FTC Consumer Sentinel and FBI IC3 category reports.