Skip to content

Repository files navigation

VisiParse — AI Image Understanding & Content Matching Engine

"Good suggestions when confident, safe rejection when uncertain."

VisiParse is a production-grade AI backend engine that automatically analyzes an image library, generates schema-validated structured tags, computes dense vector embeddings, and matches appropriate images to blog posts.

It features a Multi-Layer Mismatch Guard that provably prevents false-positive recommendations (such as placing a wolf photo on a red fox article) and provides clear human-readable refusal explanations when recommendations do not clear safety thresholds.


📐 System Architecture

                                  +-----------------------+
                                  |  50-Image Corpus      |
                                  |  (Pexels / Unsplash)  |
                                  +-----------+-----------+
                                              |
                                              v
                              +---------------+---------------+
                              | Batch Ingestion Job           |
                              | (Gemini 2.5 Flash Vision AI)  |
                              +---------------+---------------+
                                              |
                  +---------------------------+---------------------------+
                  |                                                       |
                  v                                                       v
     +------------+------------+                             +------------+------------+
     | Schema Validation       |                             | Cost Metering & Tracking|
     | (Pydantic v2)           |                             | (CostTracker DB Table)  |
     +------------+------------+                             +-------------------------+
                  |
                  v
     +------------+------------+
     | Confidence Gate Check   |  --> [confidence < 0.70] --> Flagged for Review
     | (ImageMetadata Table)   |
     +------------+------------+
                  |
                  v
     +------------+------------+
     | Local Vector Embedding  |
     | (Ollama all-minilm 384d)|
     +------------+------------+
                  |
                  +---------------------------+
                                              |
+-----------------------+                     v
| Blog Post Query       | -------> +----------+----------+
| (Title & Body Text)   |          | Cosine Similarity   |
+-----------------------+          | Candidate Ranking   |
                                   +----------+----------+
                                              |
                                              v
                                   +----------+----------+
                                   | Multi-Layer Guard   |
                                   | 1. Confidence Gate  |
                                   | 2. Similarity Floor |
                                   | 3. Taxonomic Guard  |
                                   +----------+----------+
                                              |
                                              v
                        +---------------------+---------------------+
                        |                                           |
                        v                                           v
           +------------+------------+                 +------------+------------+
           | MATCH SUGGESTED         |                 | REJECTED / REFUSAL      |
           | (Status: SUGGESTED)     |                 | (Human-Readable Reason) |
           +------------+------------+                 +------------+------------+
                        |                                           |
                        +---------------------+---------------------+
                                              |
                                              v
                                   +----------+----------+
                                   | Human Review API    |
                                   | (/reviews/approve)  |
                                   | (/reviews/reject)   |
                                   +---------------------+

✨ Key Features

  • Structured Vision AI: Ingests image files and produces schema-validated metadata (subject, category, attributes, caption, confidence).
  • Confidence Gate: Automatically flags low-confidence classifications (confidence < 0.70) rather than accepting them blindly.
  • Local Dense Vector Search: Generates 384-dimensional dense vectors using Ollama (all-minilm) to calculate cosine similarity between blog posts and image captions.
  • Multi-Layer Mismatch Guard:
    1. Confidence Gate Check: Filters out low-confidence images.
    2. Similarity Floor: Rejects matches with cosine similarity below baseline (< 0.60).
    3. Taxonomic Entity Guard: Validates animal category alignment (e.g., rejecting wolf images on fox posts).
    4. Reason Generator: Returns human-readable refusal explanations on rejection.
  • Per-Call AI Cost Metering: Attributes token counts and estimated USD costs per API call in CostTracker.
  • Human-in-the-Loop Review API: REST endpoints to review, inspect, approve, or reject recommendations with complete audit trails.
  • Automated Precision Evaluation: Benchmark evaluation dataset measuring Top-1 Precision ($\frac{\text{Correct Matches}}{\text{Total Posts}}$).

🛠️ Tech Stack ($0 Free Tier Commitment)

  • Language & Framework: Python 3.11+, FastAPI, Uvicorn
  • Vision Model: Gemini 2.5 Flash via google-genai SDK (Google AI Studio Free Tier) with offline fallback generator
  • Embedding Model: Ollama local embeddings (all-minilm, 384-d dense vectors)
  • Database & ORM: SQLite / PostgreSQL via SQLAlchemy 2.0 ORM
  • Schema Validation: Pydantic v2
  • Testing: Pytest & Pytest-Asyncio

🚀 Quickstart & Setup Guide

1. Prerequisites

  • Python 3.11+
  • Git

2. Environment Setup

Clone repository and install dependencies:

git clone https://github.com/Grantlinkz/VisiParse.git
cd VisiParse
pip install -r requirements.txt

Create .env configuration file from template:

cp .env.example .env

3. Database Initialization & Ingestion Seed

Initialize SQLite database schema and run batch ingestion across the 50-image corpus:

python scripts/init_db.py
python scripts/run_batch_ingestion.py

4. Single Run Command (FastAPI Server)

Boot the REST API server:

uvicorn app.main:app --host 0.0.0.0 --port 8000

Server health check will be accessible at: http://localhost:8000/


📊 Benchmark Evaluation & Precision

Run automated Top-1 Precision evaluation against the 15-post benchmark dataset:

python scripts/run_evaluation.py

Benchmark Results Overview

  • Labeled Benchmark Test Dataset: 15 test posts (including matching topics and unmatched edge cases)
  • Top-1 Precision Score: 93.33%
  • Mismatch Guard Refusals: 100% accurate rejection on non-matching and false-positive candidates

🎬 6-Minute Live Demo Rehearsal Script

Run the automated interactive walkthrough script demonstrating all 6 demo moments:

python scripts/demo_walkthrough.py

Demo Moments Covered:

  1. Batch Ingestion & Cost Log: Displaying 50 corpus images, confidence gate flags, and token cost tracking.
  2. Red Fox Article Match (PROBE 2): Querying a red fox blog post and retrieving the top-ranked red fox image.
  3. Forced Wolf Rejection (PROBE 3): Forcing a gray wolf candidate against a red fox post and verifying taxonomic refusal.
  4. "No Confident Match" Safety Case (PROBE 4): Querying an irrelevant topic (Quantum Physics) and demonstrating safe rejection.
  5. Human-in-the-Loop Workflow: State transitions (APPROVED / REJECTED) recorded in AuditLog.
  6. Final Benchmark Readout (PROBE 5): Reporting 93.33% Top-1 Precision metric.

🔌 API Reference

Endpoint Method Description
GET / GET System health check & threshold parameters
GET /posts/{id}/images GET Top ranked match or structured Mismatch Guard refusal
POST /posts POST Create a new blog post record
POST /reviews/{suggestion_id}/approve POST Human review approval action
POST /reviews/{suggestion_id}/reject POST Human review rejection action
GET /reviews GET List all human review audit logs
GET /eval/precision GET Run automated Top-1 Precision evaluation
GET /costs/summary GET AI token usage and USD cost summary

🧪 Automated Testing

Execute the complete pytest test suite:

python -m pytest -v

📋 Behavioral Acceptance Probes

  • PROBE 1: Batch ingestion job tags corpus; low-confidence images flagged (<0.70).
  • PROBE 2: "Red fox" post query surfaces red fox image first.
  • PROBE 3: Forced wolf candidate on fox post triggers category mismatch refusal.
  • PROBE 4: Unmatched topic returns "No confident match found" response.
  • PROBE 5: Evaluation script computes and outputs Top-1 Precision metric.
  • PROBE 6: CostTracker attributes token and USD cost for every vision/embedding API call.

👥 Contributors & Acknowledgments

This project was built with collaboration across project specification, developer implementation, AI pair programming, automated code review, and instructional leadership:

  • GrantLinkz (Project Lead & Backend AI Engineer) — Architecture design, implementation of vision pipelines, vector ranking engine, multi-layer mismatch guard, evaluation suite, and packaging.
  • FlyRank — Capstone Project Originators & Specification Authors.
  • Antigravity IDE with Gemini — AI Agentic Pair Programmer & Technical Co-Pilot.
  • CodeRabbit — Automated AI Code Reviewer.
  • Adrian | JSM (JavaScript Mastery) — AI Workflow Instructor.

⚠️ Limitations & Future Work

  • Local Ollama Fallback: Operates with a deterministic semantic fallback when local Ollama service is offline.
  • Entity Scope: Optimized for animal taxonomy matching; future extensions include general object and scene entity graphs.

About

VisiParse is a Python-based AI system that matches images to blog content. Using Gemini Flash and Ollama embeddings, it extracts structured metadata, runs dense vector searches, and applies multi-layer validation to filter low-confidence or mismatched images. It also features cost tracking and REST APIs for human-in-the-loop audit review.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages