A powerful toolkit to visualize dataset clusters and automatically detect labeling errors using Ultralytics YOLO models (v8, v11, v12).
This tool extracts deep feature vectors from your images, projects them into 2D space using t-SNE, uses k-Nearest Neighbors (k-NN) to identify "Suspicious" data points (e.g., a "Cat" image sitting deep inside a "Dog" cluster), and generates High-Fidelity Heatmaps to visualize exactly which pixels the model is looking at.
- Universal Support: Works with YOLOv8, YOLOv11, YOLOv12 (Detect, Classify, Pose, Segment, OBB).
- Smart Inspection: Automatically identifies the best layers to hook for feature extraction.
- Robust Caching: Supports massive datasets (100k+ images). If interrupted, it resumes exactly where it left off.
- Ghost Mode Visualization: A specialized plotting mode that makes clean data transparent and highlights potential errors.
- Actionable Reports: Generates a CSV list of mislabeled images to fix.
- X-Ray Vision: Generates smooth, high-resolution heatmaps. Handles rectangular images perfectly (no padding distortion) using smart aspect-ratio alignment.
.
├── util/
│ └── YoloFeatureExtractor.py # Core logic engine
│
├── 1_inspect_model.py # Step 1: Find the right layer
├── 2_generate_tsne.py # Step 2: Extract features & t-SNE
├── 3_view_plot.py # Step 3: Interactive Scatter Plot
├── 4_analyze_errors.py # Step 4: AI Conflict Detection
├── 5_generate_heatmap.py # Step 5: X-Ray Vision (Heatmaps)
│
├── models/
│ └── yolo11n-cls.pt # (Default) Small classification model
│
├── DATASET/
│ └── example/ # Included demo dataset
│ ├── cheeseburger/
│ ├── flamingo/
│ ├── violin/
│ └── ... (12 classes total)
│
└── requirements.txt
- Clone the repository (or download the files).
- Install dependencies:
pip install ultralytics scikit-learn pandas plotly tqdm(Note: GPU is recommended for Step 2, but CPU works fine for smaller datasets.)
This guide uses the included DATASET/example and yolo11n-cls.pt so you can run it immediately.
Different YOLO versions have different architectures. Run this to find the best feature layer.
- Open
1_inspect_model.pyand ensureMODEL_PATH = "models/yolo11n-cls.pt". - Run the script:
python 1_inspect_model.py
- Result: It identifies the classification vector.
(Copy the layer name
Inspecting Model: models\yolo11n-cls.pt Model loaded: models\yolo11n-cls.pt on cpu ==================== RECOMMENDED LAYERS ==================== Layer: model.model.2 | Type: C3k2 Layer: model.model.4 | Type: C3k2 Layer: model.model.6 | Type: C3k2 Layer: model.model.8 | Type: C3k2 Layer: model.model.9 | Type: C2PSA <--- ATTENTION BLOCK Layer: model.model.10 | Type: Classify <--- HEAD Layer: model.model.10.pool | Type: AdaptiveAvgPool2d <--- PRE-VECTOR POOLING Layer: model.model.10.linear | Type: Linear <--- CLASSIFICATION VECTOR ============================================================model.model.10.linearfor the next step).
This extracts features from all images and calculates t-SNE coordinates.
- Open
2_generate_tsne.pyand configure it:MODEL_PATH = "models/yolo11n-cls.pt" DATA_DIR = "./DATASET/example" # Using the demo data TARGET_LAYER = "model.model.10.linear" # Paste layer from Step 1
- Run:
python 2_generate_tsne.py
- Result: It generates JSON files for both 2D and 3D visualizations in
tsne_results/.
💡 Pro Tips:
- 2D vs 3D: Some layers reveal more patterns in 3D!
- Feature Layers (e.g.,
model.model.10): Usually work better in 3D because they contain complex spatial manifolds.- Linear Layers (e.g.,
model.model.10.linear): Usually work better in 2D because they are already flattened and optimized for separation.- Handling Large Datasets: It caches progress in the
cache/folder. You can interrupt the script and resume later. If you retrain your model, change the model filename (e.g.,v1.pt->v2.pt) to force a fresh cache.- Dataset Size: t-SNE requires more images than the Perplexity setting (default 30). We recommend at least 100 images for meaningful clusters.
Check the general health of your dataset.
-
Open
3_view_plot.pyand setJSON_FILEto the result from Step 2. -
Run:
python 3_view_plot.py
-
Result: Opens an interactive HTML plot.
2D 3D 

Find the wrong labels. This uses k-NN to find images surrounded by the wrong class.
- Open
4_analyze_errors.pyand set theJSON_FILE. - Run:
python 4_analyze_errors.py
- Result:
Visualize why the model made a decision. This uses a "Squash-and-Stretch" technique to ensure heatmaps fit rectangular images perfectly without padding distortion.
-
Open
5_generate_heatmap.py. -
Set your target image and model path.
-
Choose your layers (based on Step 1 inspection).
-
Run:
python 5_generate_heatmap.py
model.model.0 model.model.2 model.model.4 model.model.6 model.model.8 model.model.9 





(Raw Edges) (Object Shape) (Final Attention)
💡 Heatmap Interpretation:
- Red: High Activation (The model "loves" this area).
- Dark Blue/Black: Inhibition (The model is actively ignoring or suppressing this area).
- How to choose layers: Use Middle Layers (e.g.,
model.6) to check if the model found the object's shape. Use Deep Layers (e.g.,model.9) to see what specific features the model used for classification.
MIT License. Feel free to modify and use for your own projects!

