"Good suggestions when confident, safe rejection when uncertain."
VisiParse is a production-grade AI backend engine that automatically analyzes an image library, generates schema-validated structured tags, computes dense vector embeddings, and matches appropriate images to blog posts.
It features a Multi-Layer Mismatch Guard that provably prevents false-positive recommendations (such as placing a wolf photo on a red fox article) and provides clear human-readable refusal explanations when recommendations do not clear safety thresholds.
+-----------------------+
| 50-Image Corpus |
| (Pexels / Unsplash) |
+-----------+-----------+
|
v
+---------------+---------------+
| Batch Ingestion Job |
| (Gemini 2.5 Flash Vision AI) |
+---------------+---------------+
|
+---------------------------+---------------------------+
| |
v v
+------------+------------+ +------------+------------+
| Schema Validation | | Cost Metering & Tracking|
| (Pydantic v2) | | (CostTracker DB Table) |
+------------+------------+ +-------------------------+
|
v
+------------+------------+
| Confidence Gate Check | --> [confidence < 0.70] --> Flagged for Review
| (ImageMetadata Table) |
+------------+------------+
|
v
+------------+------------+
| Local Vector Embedding |
| (Ollama all-minilm 384d)|
+------------+------------+
|
+---------------------------+
|
+-----------------------+ v
| Blog Post Query | -------> +----------+----------+
| (Title & Body Text) | | Cosine Similarity |
+-----------------------+ | Candidate Ranking |
+----------+----------+
|
v
+----------+----------+
| Multi-Layer Guard |
| 1. Confidence Gate |
| 2. Similarity Floor |
| 3. Taxonomic Guard |
+----------+----------+
|
v
+---------------------+---------------------+
| |
v v
+------------+------------+ +------------+------------+
| MATCH SUGGESTED | | REJECTED / REFUSAL |
| (Status: SUGGESTED) | | (Human-Readable Reason) |
+------------+------------+ +------------+------------+
| |
+---------------------+---------------------+
|
v
+----------+----------+
| Human Review API |
| (/reviews/approve) |
| (/reviews/reject) |
+---------------------+
-
Structured Vision AI: Ingests image files and produces schema-validated metadata (
subject,category,attributes,caption,confidence). -
Confidence Gate: Automatically flags low-confidence classifications (
confidence < 0.70) rather than accepting them blindly. -
Local Dense Vector Search: Generates 384-dimensional dense vectors using Ollama (
all-minilm) to calculate cosine similarity between blog posts and image captions. -
Multi-Layer Mismatch Guard:
- Confidence Gate Check: Filters out low-confidence images.
-
Similarity Floor: Rejects matches with cosine similarity below baseline (
< 0.60). - Taxonomic Entity Guard: Validates animal category alignment (e.g., rejecting wolf images on fox posts).
- Reason Generator: Returns human-readable refusal explanations on rejection.
-
Per-Call AI Cost Metering: Attributes token counts and estimated USD costs per API call in
CostTracker. - Human-in-the-Loop Review API: REST endpoints to review, inspect, approve, or reject recommendations with complete audit trails.
-
Automated Precision Evaluation: Benchmark evaluation dataset measuring Top-1 Precision (
$\frac{\text{Correct Matches}}{\text{Total Posts}}$ ).
- Language & Framework: Python 3.11+, FastAPI, Uvicorn
- Vision Model: Gemini 2.5 Flash via
google-genaiSDK (Google AI Studio Free Tier) with offline fallback generator - Embedding Model: Ollama local embeddings (
all-minilm, 384-d dense vectors) - Database & ORM: SQLite / PostgreSQL via SQLAlchemy 2.0 ORM
- Schema Validation: Pydantic v2
- Testing: Pytest & Pytest-Asyncio
- Python 3.11+
- Git
Clone repository and install dependencies:
git clone https://github.com/Grantlinkz/VisiParse.git
cd VisiParse
pip install -r requirements.txtCreate .env configuration file from template:
cp .env.example .envInitialize SQLite database schema and run batch ingestion across the 50-image corpus:
python scripts/init_db.py
python scripts/run_batch_ingestion.pyBoot the REST API server:
uvicorn app.main:app --host 0.0.0.0 --port 8000Server health check will be accessible at: http://localhost:8000/
Run automated Top-1 Precision evaluation against the 15-post benchmark dataset:
python scripts/run_evaluation.py- Labeled Benchmark Test Dataset: 15 test posts (including matching topics and unmatched edge cases)
- Top-1 Precision Score: 93.33%
- Mismatch Guard Refusals: 100% accurate rejection on non-matching and false-positive candidates
Run the automated interactive walkthrough script demonstrating all 6 demo moments:
python scripts/demo_walkthrough.py- Batch Ingestion & Cost Log: Displaying 50 corpus images, confidence gate flags, and token cost tracking.
- Red Fox Article Match (PROBE 2): Querying a red fox blog post and retrieving the top-ranked red fox image.
- Forced Wolf Rejection (PROBE 3): Forcing a gray wolf candidate against a red fox post and verifying taxonomic refusal.
- "No Confident Match" Safety Case (PROBE 4): Querying an irrelevant topic (Quantum Physics) and demonstrating safe rejection.
- Human-in-the-Loop Workflow: State transitions (
APPROVED/REJECTED) recorded inAuditLog. - Final Benchmark Readout (PROBE 5): Reporting 93.33% Top-1 Precision metric.
| Endpoint | Method | Description |
|---|---|---|
GET / |
GET |
System health check & threshold parameters |
GET /posts/{id}/images |
GET |
Top ranked match or structured Mismatch Guard refusal |
POST /posts |
POST |
Create a new blog post record |
POST /reviews/{suggestion_id}/approve |
POST |
Human review approval action |
POST /reviews/{suggestion_id}/reject |
POST |
Human review rejection action |
GET /reviews |
GET |
List all human review audit logs |
GET /eval/precision |
GET |
Run automated Top-1 Precision evaluation |
GET /costs/summary |
GET |
AI token usage and USD cost summary |
Execute the complete pytest test suite:
python -m pytest -v- PROBE 1: Batch ingestion job tags corpus; low-confidence images flagged (
<0.70). - PROBE 2: "Red fox" post query surfaces red fox image first.
- PROBE 3: Forced wolf candidate on fox post triggers category mismatch refusal.
- PROBE 4: Unmatched topic returns "No confident match found" response.
- PROBE 5: Evaluation script computes and outputs Top-1 Precision metric.
- PROBE 6:
CostTrackerattributes token and USD cost for every vision/embedding API call.
This project was built with collaboration across project specification, developer implementation, AI pair programming, automated code review, and instructional leadership:
- GrantLinkz (Project Lead & Backend AI Engineer) — Architecture design, implementation of vision pipelines, vector ranking engine, multi-layer mismatch guard, evaluation suite, and packaging.
- FlyRank — Capstone Project Originators & Specification Authors.
- Antigravity IDE with Gemini — AI Agentic Pair Programmer & Technical Co-Pilot.
- CodeRabbit — Automated AI Code Reviewer.
- Adrian | JSM (JavaScript Mastery) — AI Workflow Instructor.
- Local Ollama Fallback: Operates with a deterministic semantic fallback when local Ollama service is offline.
- Entity Scope: Optimized for animal taxonomy matching; future extensions include general object and scene entity graphs.