Skip to content

Repository files navigation

pico-type 🔍

A tiny byte-level multi-head content classifier — ~1.5M params, ~9MB single-file ONNX (FP32), ~18ms CPU inference.

Classifies any content from raw bytes: coarse type · modality · subtype · code language · text language · file MIME · risk flags

License Python PyPI ONNX CI HuggingFace Space HuggingFace Model


✨ Features

  • No tokenizer — operates directly on raw UTF-8 bytes (supports all languages, no preprocessing)
  • 7 heads, one forward pass — coarse type, modality, subtype, code language, text language, file MIME, risk flags
  • 4 Matryoshka tiers — tiny (16d) → small (64d) → base (192d) → pro (576d) — same trunk, accuracy scales with dim
  • ~9MB single-file ONNX (FP32) — deploy on edge devices, serverless, browser (WebAssembly/ONNX Runtime Web)
  • ~18ms inference on CPU via ONNX Runtime
  • CLI, Python API, Gradio Space, MCP server — ready to use

📊 Evaluation

Overall Accuracy (v2 — trained on real data)

Head Classes Accuracy Dataset
coarse 12 100% Synthetic eval
modality 8 100% Synthetic eval
subtype 24 93.8% Synthetic eval
code_lang 62 60.3% The Heap — 24 real-world langs, 1,200 samples
text_lang 30 98.3% Wikipedia — 30 langs, 1,500 samples
file_mime 90 100% Synthetic eval
risk (mAP) 6 100% Synthetic eval

v0.1 baseline (synthetic-only): code_lang 3%, text_lang 19%. Real-data training in v2 improves code by 57pp and text by 79pp.

Code Language — Per-Language Accuracy

Excellent (90%+) Good (70–89%) Needs Work (<50%)
cpp 96%, dart 98%, erlang 98%, rust 98%, r 94%, swift 92%, python 88%, lua 88% go 86%, ruby 86%, ocaml 84%, php 78%, csharp 76%, java 76%, kotlin 76%, c 62% perl 50%, haskell 24%, scala 4%, javascript 2%, clojure 0%, elixir 0%, julia 0%, sql 0%

Note: Low-accuracy languages have fewer real training samples. More data will improve them.

🚀 Quick Start

Install

pip install picotype

CLI

# Classify from stdin
echo "def hello(name):\n    return f'Hi {name}'" | picotype --pretty

# Classify a file
picotype --file document.txt

# Classify clipboard content
picotype --clip

# All 4 tiers available
echo "..." | picotype --tier pro

Python API

from picotype import load_onnx_model, run_onnx

session = load_onnx_model("base")
result = run_onnx(session, "def hello(): pass")
print(result)
# {
#   "coarse": "code",
#   "code_language": "python",
#   "modality": "textual",
#   "confidence": 0.98,
#   ...
# }

MCP Server (for Claude Desktop, Cursor, etc.)

pip install picotype
PICOTYPE_MODEL_DIR=./checkpoints python -m model.pico_type.mcp_server

Then add to your MCP config:

{
  "mcpServers": {
    "pico-type": {
      "command": "python",
      "args": ["-m", "model.pico_type.mcp_server"],
      "env": { "PICOTYPE_MODEL_DIR": "./checkpoints" }
    }
  }
}

Gradio Web UI

Try it live: huggingface.co/spaces/eulogik/pico-type

🏗 Architecture

Bytes ─▶ ByteEmbed(256→96d) ─▶ 3×Conv1D(k=3,5,7) ─▶ 2×BiAttention(RoPE) ─▶ Pool ─▶ 7×Matryoshka Heads
Component Detail
ByteEmbed Lookup-free embedding — each byte value (0–255) maps to a learned 96-dim vector
Conv1D 3 parallel depthwise convolutions (kernel widths 3, 5, 7) with residual + layer norm
BiAttention Bidirectional self-attention with Rotary Position Embeddings (RoPE), 4 heads
Pool Mean + max + std deviation concatenation → fixed-size representation
Heads Matryoshka-style: slice pool dim to 16/64/192/576, project to 7 linear classifiers

Total parameters: 1.43M (tiny) / 1.45M (small) / 1.48M (base) / 1.56M (pro)

🔧 Model Tiers

Tier Dim Params ONNX Size Accuracy Multiplier
tiny 16 1.43M 9.09 MB 0.65×
small 64 1.45M 9.13 MB 0.82×
base 192 1.48M 9.25 MB 1.0× (reference)
pro 576 1.56M 9.61 MB 1.05×

ONNX sizes are single-file FP32 exports (graph-only files are 203–206 KB).

All tiers share the same backbone; only the final linear projection layers differ. Higher-tier models use more dimensions for finer-grained classification.

🧪 Classification Heads

Head Classes What It Detects
coarse 12 text, code, link, image, file, config, markup, data, error, secret, archive, binary
modality 8 textual, binary_image, binary_archive, binary_executable, binary_document, etc.
subtype 24 json, yaml, toml, csv, html, markdown, sql, log, dockerfile, makefile, etc.
code_lang 62 python, javascript, typescript, java, c, cpp, go, rust, ruby, php, swift, kotlin, and 50 more
text_lang 30 en, es, fr, de, it, pt, nl, ru, zh, ja, ko, vi, th, id, and 15 more
file_mime 90 application/json, image/png, video/mp4, font/ttf, application/wasm, and 84 more
risk 6 api_key, jwt, password, email, phone, ssh_key

🌐 Deployment

Platform Link Notes
HuggingFace Space eulogik/pico-type Gradio web UI, no GPU needed
HuggingFace Model eulogik/pico-type ONNX models + export metadata
GitHub eulogik/pico-type Source code, training, paper
PyPI pip install picotype Python package
ONNX Runtime Use with onnxruntime.js Browser/Node.js deployment

📚 Resources

  • Paper — Architecture, training, and evaluation details
  • Model Card — Detailed architecture and training configuration
  • Walkthrough — Development log and decisions
  • Architecture Plan — Original design document

📄 License

Apache 2.0


Built with PyTorch · ONNX · Gradio · HuggingFace

About

A tiny byte-level multi-head content classifier (~1.5M params, ~200KB ONNX, <6ms). Classifies code, text, markup, config, images, binary, secrets, 62 code languages, 30 text languages, 90 MIME types from raw bytes — no tokenizer needed.

Topics

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages