Copy Detection Pattern (CDP) authentication for anti-counterfeiting — built during my machine-learning internship at Checko, and shared with their permission.
A CDP is a square of random dots printed inside a product's QR code at the limit of the printer's resolution. Every print loses a little detail, so a genuine label (printed once) is sharper than a counterfeit (photographed or scanned, then printed again). The job: decide genuine or fake from a single customer photo, taken on an ordinary phone, where each dot of the pattern is only about 3 pixels wide.
On 3,309 phone photos (1,749 counterfeit, 1,560 genuine; 96 designs, 6 phones, 3 printer families), with 5-fold cross-validation held out by design and every number averaged over 3 training runs:
| Before | After | |
|---|---|---|
| Counterfeits accepted as genuine | 3.47 per 100 | 1.05 per 100 (0.43 with the double check) |
| Genuine products wrongly rejected | 7.65 per 100 | 2.98 per 100 |
| Crop placement error (median) | 2.9–3.5 px | 0.5–0.9 px |
| Separation (ROC AUC) | 0.988 | 0.998 |
What made the difference, in order:
- Precise cropping. Straightening the photo from the QR corners left the crop about one dot off. Reading the edges of the CDP off the white strips that surround it in the label design cut that error 4–6×, and turned the strips into a check: if they are not where they should be, the crop is refused instead of guessed. Genuine products wrongly rejected fell from 7.65 to 4.25 per 100.
- Comparing with the original design. The QR code says which design was printed, and a copy is always one print step further from it. With crops now accurate to a third of a dot, the aligned design fed as a second input channel cut counterfeits accepted 3.4× and genuine rejections by a third, both at once.
- An honest evaluation. Pattern-held-out folds, 3 seeds per setup (single-run differences under ~1.5 points are noise here), comparisons at equal genuine-rejection rates, pass lines set only from genuine data, and refused photos counted as retakes rather than dropped.
flowchart TD
A[Phone photo] --> B[Cut out the QR code]
B --> C[Straighten it: ArUco markers,<br/>else the QR corners]
C --> D{White strips where<br/>the design puts them?}
D -- no --> F[Next candidate,<br/>then a bigger cut-out]
F --> D
F -- nothing passes --> R[Ask for a retake]
D -- yes --> E[Read the CDP edges off the strips;<br/>crop at native scale, greyscale]
E --> H[Tile CNN on the photo<br/>+ its aligned design]
E --> J[Correlation with<br/>the design]
H --> K{Both say genuine?}
J --> K
K -- yes --> G[Genuine]
K -- no --> X[Counterfeit]
- Localisation (
cdp/extract.py): calibrate once on the clean digital template, then fit a homography from ArUco marker corners or the QR's finder patterns and warp the CDP into template space. - White-strip refinement and check (
cdp/refine.py): in the straightened label, four clean white bands per axis bracket the CDP at fixed, known positions. A linear fit to the bands gives the CDP's true edges; failing to find them flags a misplaced crop. - Model (
cdp/model.py): a ~62k-parameter CNN on 32×32 tiles. A constrained high-pass first layer (every kernel a difference operator, a trick from image forensics) makes it learn print texture rather than image content; a second branch sees the photo and its aligned design together. - Design comparison (
cdp/design.py): the pattern's own design, slid over the crop in all 8 orientations. Needs no training, and works the same for every printer.
A new ruler made this measurable: every CDP is a random pattern, so its digital design lines up with the photo at exactly one position, and the offset of the best match is the crop error. The previous proxy correlated the whole QR code (the CDP is 5% of it) and had been used to justify a pipeline change it could not actually measure.
Every crop the check refused was genuinely in the wrong place (29 of 29, confirmed by matching the design in all 8 orientations and by eye). The first version asked 4.6% of genuine users for a retake; tracing those to one phone model showed the capture app records the QR box in the wrong place on that phone (right size, shifted by up to a third of the QR). Retrying with a larger cut-out rescued 318 photos, every one of which matches its design, and brought retakes down to 0.26%.
| Per 100 photos | Genuine wrongly rejected | Counterfeits accepted |
|---|---|---|
| Photo-only model | 4.26 | 3.53 |
| Photo + design model | 2.98 | 1.05 |
| … plus design-match line, 3% | 4.35 | 0.75 |
| … plus design-match line, 5% | 5.53 | 0.43 |
| … plus design-match line, 8% | 8.55 | 0.30 |
The design-match line is a pass mark on the raw correlation, set from genuine crops only (and cross-fitted, so it is never set on the crops it is tested on). How strict to be is a product decision; the chart above shows the whole trade-off. In September, feeding the design to the model lost to the photo alone — with crops that were about a dot off. Fixing the cropping is what made it work.
Hiding one printer's counterfeits from training shows the photo-only model partly learns
which printer made this, not was this copied: copies from the brand's own printer model get
through 39–74% of the time. Comparing with the design needs no training, so it does not learn
that shortcut and cuts this to 6–16%. Training only on matched genuine/copy pairs did not
help (a negative result, kept in docs/findings.md). The remaining fix is
data: copies made on every printer model the brand uses.
cdp/ the library
extract.py QR / ArUco localisation, homography, CDP crop
refine.py white-strip refinement, candidate ladder, retake rule
design.py matching and aligning a crop with its digital design
model.py tile CNN (constrained high-pass front, optional design branch)
data.py, config.py capture-set access, label rules, paths (all via env vars)
scripts/
build_dataset.py photos -> crops + manifest + refused list
train.py 5-fold pattern-held-out training; options for every experiment
score_design_match.py training-free design correlation
align_designs.py the aligned design channel
measure_crop_error.py the crop-placement ruler; check_crops.py sanity check
eval_*.py the read-out for each experiment; make_figures.py
experiments/ the exact command sequence behind each result
results/ the printed outputs every number above comes from
docs/findings.md the full experiment log, including what did not work
The capture data, digital designs and label photos belong to Checko and are not in this repository. With access to them:
pip install -r requirements.txt
export CDP_FULL_DATA=/path/to/full_data # the part-zips
export CDP_TEMPLATES=/path/to/designs # VIKRAM00.png ... VIKRAM99.png
export CDP_SESSIONS=/path/to/sessions.csv # per-session metadata
export CDP_WORK=./work # outputs
python scripts/build_dataset.py --out work/nfhd_dataset_gate2
python scripts/train.py --manifest work/nfhd_dataset_gate2/manifest.csv \
--patch 32 --stride 16 --epochs 10 --seed 42 --tag gate2_s42
bash experiments/run_design_model.shA full training run (5 folds, 10 epochs) takes about 6 minutes on a laptop RTX 4050.
Stack: Python, PyTorch, OpenCV (ArUco, homographies, template matching), scikit-learn, pandas, matplotlib.
Skand Agarwal · B.Tech Computer & Communication Engineering, Manipal University Jaipur (2026) · Machine Learning Intern, Checko