Skip to content

Repository files navigation

Unhook

Read. Reveal. Unhook.

unhook-one.vercel.app

Paste a message you are not sure about. Unhook shows you the exact words being used to manipulate you — and tells you what to do next.


The problem

Older adults lose more money to fraud than any other group.

These messages do not arrive as spam. They come as a text from a real phone number, a WhatsApp from someone who seems to know you, a call from a number that matches your bank's. No filter ever sees them. The person is the filter.

And every tool built to help answers the same question — is this a scam, yes or no. You get a verdict. It may save you this once, but you learn nothing about how it worked on you, so the next scam catches you just the same.

The question a frightened person actually has is "why should I not trust this?" — and after that, "what do I do now?"

Unhook answers those two. It marks the specific words doing the manipulating, names the technique, and gives one instruction. The goal is that the reader starts spotting it themselves.

What it does

Paste a text, email, or message. Ten seconds later:

A verdict in plain words "This looks like a scam" — not "Risk: 87/100"
The tells, marked in your own message Tap any highlighted phrase to see why it is there
One instruction "Call your grandchild on their usual number, and tell their parents."
What happens next The scam's next three moves, projected from the tactics actually found
Read aloud For the reader whose eyesight is why they were unsure in the first place
Show someone you trust Reveals the exact text before you copy it — nothing is sent by Unhook
Answers in the message's language Fraud targets people in their first language; a warning you have to translate is one you won't act on

Five real examples are one click away and answer instantly.

How it works

                    ┌──────────────────────────────────────┐
   pasted message → │ 1. normalise                         │
                    │    strip zero-width, NFKC, homograph │
                    └──────────────┬───────────────────────┘
                                   │
              ┌────────────────────┴────────────────────┐
              ▼                                         ▼
   ┌──────────────────────┐                ┌──────────────────────────┐
   │ 2a. deterministic    │                │ 2b. model, in two passes │
   │     ~2 ms, free      │   (these two   │  ① evidence: which of 15 │
   │  16 signal types:    │    layers      │     tactics, exact quote │
   │  lookalike domains,  │    never see   │  ② narration: say it in  │
   │  crypto addresses,   │    each        │     plain words          │
   │  gift-card requests, │    other)      │                          │
   │  urgency markers…    │                │  never returns a score   │
   └──────────┬───────────┘                └────────────┬─────────────┘
              │                                         │
              │                            ┌────────────▼─────────────┐
              │                            │ 3. span resolver         │
              │                            │    quote → exact offsets │
              │                            │    3 tiers, testable     │
              │                            └────────────┬─────────────┘
              └────────────────────┬────────────────────┘
                                   ▼
                    ┌──────────────────────────────────────┐
                    │ 4. risk rubric — published constants │
                    │    superlinear severity, no double   │
                    │    counting, combination bonuses     │
                    └──────────────┬───────────────────────┘
                                   ▼
                              the analysis

Three decisions that shape everything

The model never returns a score. It returns evidence; the number comes from a rubric in lib/risk/rubric.ts that you can read and disagree with. Model self-reported confidence is uncalibrated and moves when you reword the prompt — and swapping the model would silently move every score.

The model never returns character offsets. Models quote reliably and count characters unreliably. It returns the verbatim quote; a three-tier resolver (lib/spans/resolve.ts) computes the offsets where it can be unit-tested. That is what makes the highlighting trustworthy — measured at 1.000 match quality across 166 findings.

The two detection layers are blind to each other. The deterministic detectors do not see the model's output and the model does not see theirs. That independence is what makes their disagreement informative — and the UI says so when they disagree, rather than hiding it.

What the numbers say

Six pre-registered experiments over a hand-labelled set of 81 messages — 43 scams and 38 genuine, of which 27 are genuine messages deliberately chosen to look like fraud (real 2FA codes, a real bank fraud alert, a hospital asking you to call urgently). Full harness in lib/eval/, raw results in experiments/runs/.

Accuracy is not the headline. The dataset is 53% scam, so answering "scam" every time scores 53% and is worse than useless. The numbers that decide whether this is shippable are the false-positive rate and the hard-case recall.

E4 — is the hybrid architecture worth it?

pattern checks only + language model
recall 18.6% 97.7%
false positives 0.0% 0.0%
F1 0.314 0.988
recall on hard cases 9.1% 100%

The deterministic layer is perfectly precise and nearly blind. It scores 0/100 on the romance scam, the business-email compromise, and the "safe account" bank fraud — messages with no link, no payment rail and no urgency vocabulary. Those are pure pretext, and no pattern can reach them. This is the empirical case for the architecture, rather than an assertion about it.

E1 — does the prompt earn its place?

v1 taxonomy v2 minimal v3 few-shot
recall 90.7% 95.3% 88.4%
false positives 0.0% 19.4% 2.7%
F1 0.951 0.901 0.927

Strip the taxonomy and the false-positive rules and recall goes up while the same model raises seven false alarms where it had raised none — among them a real Barclays fraud alert, scored 89/100. The prompt is not teaching the model to find scams; the schema does that. It is teaching it the discipline not to flag.

E2 — one model call, or two?

one pass two passes
next-action reading grade 9.3 (22 words) 4.5 (12 words)
summaries at grade ≤6 51% 60%
recall 90.7% 97.7%
cost $0.20 $0.24

Asking one call to find evidence and write for a frightened 74-year-old does both jobs worse. Splitting them moved the single most important sentence in the product from high-school reading level to primary-school, for +18% cost.

E3 — which of three models?

Haiku 4.5 Opus 5 GPT-5.5
F1 0.977 0.988 0.938
false positives 2.6% 0.0% 0.0%
hard-case recall 100% 100% 90.9%
tactic F1 0.805 0.875 0.840
unresolved spans 9.8% 2.4% 0.0%
latency p50 3.7 s 10.5 s 7.9 s
prompt cache hit 0% 87.8% 0%
summaries at grade ≤6 60% 94% 98%

Opus 5 was adopted, and the cache result is why it is affordable. Its 512-token cache minimum clears our ~1,500-token system prompt; Haiku's 4096-token minimum never engages. Same code, same prompt — the model tier decides whether the optimisation exists at all, which made the highest list-price model the cheapest arm to run.

The production model was chosen by this experiment, not before it.

E5 — can the audience actually read it?

Flesch–Kincaid on every line the reader sees, via npm run e5.

before after
hand-written taxonomy copy grade 5.8, 8 of 15 above target grade 2.9, 0 of 15 above target
model next-action line grade 9.3 grade 4.1

Both numbers moved because they were measured. The taxonomy rewrite was driven by the per-tactic report; the action line was fixed by the two-pass split in E2. The lever in both cases turned out to be sentence length rather than word length.

E6 — does the score mean anything?

Monotonic across all five buckets: ECE 0.145, Brier 0.072. The score is systematically under-confident — the 40–60 band was 100% scams. That is the safe direction to be wrong in for a product whose worst failure is crying wolf, but it misses the conventional 0.10 ECE bar and is reported rather than tuned away.

Accessibility

Not a checklist item — it is the brief. The reader is often in their seventies, on a phone, and frightened.

  • Atkinson Hyperlegible, drawn by the Braille Institute for low-vision readers, where I l 1 and O 0 are deliberately differentiated
  • 19px base, nothing below 17px, 48px minimum touch targets
  • npm run contrast verifies every colour pair — all text ≥7:1 (WCAG AAA), graphical elements ≥3:1 (WCAG 1.4.11). It caught two failures that would otherwise have shipped.
  • Colour is never the only signal: each verdict carries a distinct icon shape
  • Read-aloud, full keyboard operation, prefers-reduced-motion respected, zoom not capped

Running it

npm install
cp .env.example .env.local     # add one API key
npm run check:keys             # verifies without printing them

npm run build:demo             # MUST run before build — see below
npm run build && npm start

npm run build:demo precomputes the five demo analyses into lib/demo/cache.json. Building without it bakes an empty cache into the bundle and every example silently falls through to a live API call. There is a test guarding this, because it happened once.

command
npm test 122 tests
npm run dataset dataset composition and tactic coverage
npm run eval -- --id x --provider anthropic one experiment arm
npm run compare -- run-a run-b arms side by side
npm run e5 / npm run e6 readability / calibration reports
npm run contrast WCAG verification

Limitations

Stated plainly, because a fraud tool that oversells itself is its own hazard.

  • It can be wrong. 97.7% recall means roughly one scam in forty is missed. When money or an account is involved, the advice is always to check with the company on a number you already have.
  • The interface chrome is English. The analysis — summary, action, and every explanation — comes back in whatever language the message was written in, and span highlighting works unchanged because quotes are copied verbatim from the source. But the tactic labels come from our own taxonomy and stay English, and the measured numbers below were gathered on an English test set.
  • 81 cases is a small set, written from published fraud patterns rather than collected from victims. Two runs of the same configuration differed by ~2 points of recall, so single-run differences of that size are noise.
  • Two tactics are under-measured — escalating_commitment has only 2 labelled examples. Its per-tactic score is not meaningful.
  • The rate limiter is per-instance and in-memory. On serverless it is a speed bump, not a security control.
  • No screenshot support, and most people receive scams as images.

Licence

MIT — see LICENSE.

The manipulation taxonomy draws on Cialdini's Influence (1984), Stajano & Wilson, Understanding scam victims: seven principles for systems security (CACM 54:3, 2011), and the FTC Consumer Sentinel and FBI IC3 category reports.

About

Shows you the exact words a scam message is using to manipulate you, and what to do next. Built for older adults. Hybrid pattern-matching + LLM, span-level evidence, published risk rubric, six measured experiments.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages