A tiny byte-level multi-head content classifier — ~1.5M params, ~9MB single-file ONNX (FP32), ~18ms CPU inference.
Classifies any content from raw bytes: coarse type · modality · subtype · code language · text language · file MIME · risk flags
- No tokenizer — operates directly on raw UTF-8 bytes (supports all languages, no preprocessing)
- 7 heads, one forward pass — coarse type, modality, subtype, code language, text language, file MIME, risk flags
- 4 Matryoshka tiers — tiny (16d) → small (64d) → base (192d) → pro (576d) — same trunk, accuracy scales with dim
- ~9MB single-file ONNX (FP32) — deploy on edge devices, serverless, browser (WebAssembly/ONNX Runtime Web)
- ~18ms inference on CPU via ONNX Runtime
- CLI, Python API, Gradio Space, MCP server — ready to use
| Head | Classes | Accuracy | Dataset |
|---|---|---|---|
| coarse | 12 | 100% | Synthetic eval |
| modality | 8 | 100% | Synthetic eval |
| subtype | 24 | 93.8% | Synthetic eval |
| code_lang | 62 | 60.3% | The Heap — 24 real-world langs, 1,200 samples |
| text_lang | 30 | 98.3% | Wikipedia — 30 langs, 1,500 samples |
| file_mime | 90 | 100% | Synthetic eval |
| risk (mAP) | 6 | 100% | Synthetic eval |
v0.1 baseline (synthetic-only): code_lang 3%, text_lang 19%. Real-data training in v2 improves code by 57pp and text by 79pp.
| Excellent (90%+) | Good (70–89%) | Needs Work (<50%) |
|---|---|---|
| cpp 96%, dart 98%, erlang 98%, rust 98%, r 94%, swift 92%, python 88%, lua 88% | go 86%, ruby 86%, ocaml 84%, php 78%, csharp 76%, java 76%, kotlin 76%, c 62% | perl 50%, haskell 24%, scala 4%, javascript 2%, clojure 0%, elixir 0%, julia 0%, sql 0% |
Note: Low-accuracy languages have fewer real training samples. More data will improve them.
pip install picotype# Classify from stdin
echo "def hello(name):\n return f'Hi {name}'" | picotype --pretty
# Classify a file
picotype --file document.txt
# Classify clipboard content
picotype --clip
# All 4 tiers available
echo "..." | picotype --tier profrom picotype import load_onnx_model, run_onnx
session = load_onnx_model("base")
result = run_onnx(session, "def hello(): pass")
print(result)
# {
# "coarse": "code",
# "code_language": "python",
# "modality": "textual",
# "confidence": 0.98,
# ...
# }pip install picotype
PICOTYPE_MODEL_DIR=./checkpoints python -m model.pico_type.mcp_serverThen add to your MCP config:
{
"mcpServers": {
"pico-type": {
"command": "python",
"args": ["-m", "model.pico_type.mcp_server"],
"env": { "PICOTYPE_MODEL_DIR": "./checkpoints" }
}
}
}Try it live: huggingface.co/spaces/eulogik/pico-type
Bytes ─▶ ByteEmbed(256→96d) ─▶ 3×Conv1D(k=3,5,7) ─▶ 2×BiAttention(RoPE) ─▶ Pool ─▶ 7×Matryoshka Heads
| Component | Detail |
|---|---|
| ByteEmbed | Lookup-free embedding — each byte value (0–255) maps to a learned 96-dim vector |
| Conv1D | 3 parallel depthwise convolutions (kernel widths 3, 5, 7) with residual + layer norm |
| BiAttention | Bidirectional self-attention with Rotary Position Embeddings (RoPE), 4 heads |
| Pool | Mean + max + std deviation concatenation → fixed-size representation |
| Heads | Matryoshka-style: slice pool dim to 16/64/192/576, project to 7 linear classifiers |
Total parameters: 1.43M (tiny) / 1.45M (small) / 1.48M (base) / 1.56M (pro)
| Tier | Dim | Params | ONNX Size | Accuracy Multiplier |
|---|---|---|---|---|
| tiny | 16 | 1.43M | 9.09 MB | 0.65× |
| small | 64 | 1.45M | 9.13 MB | 0.82× |
| base | 192 | 1.48M | 9.25 MB | 1.0× (reference) |
| pro | 576 | 1.56M | 9.61 MB | 1.05× |
ONNX sizes are single-file FP32 exports (graph-only files are 203–206 KB).
All tiers share the same backbone; only the final linear projection layers differ. Higher-tier models use more dimensions for finer-grained classification.
| Head | Classes | What It Detects |
|---|---|---|
| coarse | 12 | text, code, link, image, file, config, markup, data, error, secret, archive, binary |
| modality | 8 | textual, binary_image, binary_archive, binary_executable, binary_document, etc. |
| subtype | 24 | json, yaml, toml, csv, html, markdown, sql, log, dockerfile, makefile, etc. |
| code_lang | 62 | python, javascript, typescript, java, c, cpp, go, rust, ruby, php, swift, kotlin, and 50 more |
| text_lang | 30 | en, es, fr, de, it, pt, nl, ru, zh, ja, ko, vi, th, id, and 15 more |
| file_mime | 90 | application/json, image/png, video/mp4, font/ttf, application/wasm, and 84 more |
| risk | 6 | api_key, jwt, password, email, phone, ssh_key |
| Platform | Link | Notes |
|---|---|---|
| HuggingFace Space | eulogik/pico-type | Gradio web UI, no GPU needed |
| HuggingFace Model | eulogik/pico-type | ONNX models + export metadata |
| GitHub | eulogik/pico-type | Source code, training, paper |
| PyPI | pip install picotype |
Python package |
| ONNX Runtime | Use with onnxruntime.js | Browser/Node.js deployment |
- Paper — Architecture, training, and evaluation details
- Model Card — Detailed architecture and training configuration
- Walkthrough — Development log and decisions
- Architecture Plan — Original design document
Apache 2.0