Skip to content

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AI-Powered IT Support Assistant

A local, human-in-the-loop AI assistant designed to support Tier 1 IT technicians with incident triage, knowledge retrieval, diagnostic guidance, safe command recommendations, and escalation decisions.

Python Streamlit Ollama RAG Status

Project Overview

This project demonstrates how locally hosted artificial intelligence can assist IT support operations without sending incident data to an external AI service.

The technician submits an incident through a Streamlit interface. The application retrieves relevant internal documentation using semantic search, provides the retrieved context to a local language model, and generates a structured troubleshooting recommendation.

The assistant provides guidance only. The technician remains responsible for reviewing, validating, and approving every action.

Rather than functioning as a generic chatbot, this project applies AI directly to IT support workflows by combining incident triage, internal knowledge retrieval, diagnostic guidance, safety controls, and escalation logic.

Core Capabilities

  • Structured incident intake with category and priority.
  • Local inference using Ollama and qwen3.5:4b.
  • Semantic knowledge retrieval using embeddinggemma.
  • Persistent SQLite-backed vector index.
  • Retrieval-Augmented Generation (RAG).
  • Diagnostic questions and read-only troubleshooting steps.
  • Suggested command validation using an allowlist.
  • Escalation criteria.
  • Knowledge-source attribution with vector distances.
  • Human-in-the-loop safety model.
  • Multi-domain RAG validation.
  • Local-only processing for incident descriptions and knowledge-base content.

Supported Knowledge Domains

The current knowledge base contains runbooks derived from completed technical projects and sanitized simulated incidents:

  • Active Directory account lockouts.
  • Windows domain DNS resolution.
  • Microsoft Intune device policy synchronization.
  • Unresponsive Windows workstations.

The documents were created from completed hands-on implementations, documented procedures, and simulated support scenarios.

The application does not connect directly to production systems or expired cloud trial tenants.

Architecture

flowchart TD
    A["Technician submits incident"] --> B["Streamlit interface"]
    B --> C["EmbeddingGemma creates query vector"]
    C --> D["SQLite vector index"]
    D --> E["Top matching knowledge chunks"]
    E --> F["Qwen3.5 4B generates analysis"]
    F --> G["Safety validation and source display"]
    G --> H["Technician reviews final action"]
Loading

RAG Workflow

  1. The technician enters an incident description.
  2. embeddinggemma converts the incident into a 768-dimensional vector.
  3. The application compares the query vector against the local knowledge index.
  4. The three most relevant knowledge chunks are retrieved.
  5. The retrieved context is included in the Ollama prompt.
  6. qwen3.5:4b produces a structured recommendation.
  7. Suggested commands pass through a local safety validator.
  8. The interface displays the recommendation and its knowledge sources.
  9. The technician reviews the recommendation and makes the final decision.

Structured Recommendation

Each analysis contains:

  1. Incident summary.
  2. Three diagnostic questions.
  3. Three safe diagnostic steps.
  4. Suggested commands and their purposes.
  5. Escalation criteria.
  6. Retrieved knowledge sources.

Project Structure

ai-powered-it-support-assistant/
├── app.py
├── requirements.txt
├── README.md
├── knowledge_base/
│   ├── active_directory/
│   │   └── account_lockout.md
│   ├── intune/
│   │   └── device_policy_sync.md
│   ├── networking/
│   │   └── dns_name_resolution.md
│   └── windows/
│       └── workstation_unresponsive.md
├── src/
│   ├── __init__.py
│   ├── llm_client.py
│   ├── rag.py
│   ├── safety.py
│   └── schemas.py
├── tests/
│   ├── __init__.py
│   └── validate_rag.py
└── docs/
    └── screenshots/

The generated knowledge_index.db file is excluded from Git because it can be recreated from the Markdown knowledge base.

Technology Stack

  • Python 3.12
  • Streamlit
  • Ollama
  • Qwen3.5 4B
  • EmbeddingGemma
  • SQLite
  • Pydantic
  • PowerShell
  • Git and GitHub

Local Installation

Prerequisites

Install:

  • Python 3.12
  • Git
  • Ollama
  • Visual Studio Code

Create the virtual environment

py -3.12 -m venv .venv
Set-ExecutionPolicy -Scope Process -ExecutionPolicy RemoteSigned
.\.venv\Scripts\Activate.ps1

Install Python dependencies

python -m pip install --upgrade pip
python -m pip install -r requirements.txt

Download the local models

ollama pull qwen3.5:4b
ollama pull embeddinggemma

Build the knowledge index

python -m src.rag

Validate multi-domain retrieval

python -m tests.validate_rag

Expected result:

Result: 4/4 domains passed

Start the application

python -m streamlit run app.py

Open:

http://localhost:8501

The port may change if another Streamlit process is already running.

Validation

The retrieval validation tests four different incident categories:

Domain Example scenario Expected knowledge source
Networking DNS requests time out while IP connectivity works DNS resolution runbook
Active Directory Account repeatedly locks after a password change Account lockout runbook
Microsoft Intune Enrolled device does not synchronize policies Intune policy-sync runbook
Windows Desktop becomes unresponsive after sign-in Windows workstation runbook

A successful result confirms that the same local assistant can select different internal documentation based on the incident context.

Additional validation confirmed:

  • All four knowledge domains returned the expected primary source.
  • The application generated structured recommendations consistently.
  • Suggested commands were processed through the local safety validator.
  • Retrieved knowledge sources were displayed for technician review.
  • The assistant did not automatically execute system commands.
  • The final action remained under technician control.

Evidence

Multi-Domain RAG Validation

The retrieval validation confirms that the assistant can select different knowledge sources based on the technical domain of the submitted incident.

Multi-domain RAG validation

AI-Assisted Incident Analysis

The final incident-analysis interface displays diagnostic questions, recommended troubleshooting steps, suggested commands, and escalation criteria.

Final incident analysis

Retrieved Knowledge Sources

The interface displays the knowledge sources used during retrieval so the technician can review the documentation that influenced the recommendation.

Retrieved knowledge sources

Additional implementation screenshots are available in the docs/screenshots directory.

Safety Model

The application implements several safety controls:

  • Recommendations are advisory.
  • The technician makes the final decision.
  • Commands are validated against an allowlist.
  • State-changing commands are blocked.
  • Retrieved internal documentation is displayed for review.
  • Diagnostic steps prioritize evidence collection.
  • Escalation criteria are included in every response.
  • The application does not automatically execute generated commands.

This human-in-the-loop model is intended to assist technicians rather than replace technical judgment.

Design Decisions

Local AI

Ollama keeps incident descriptions and knowledge-base content on the local workstation rather than sending them to an external AI service.

SQLite Vector Index

ChromaDB was initially evaluated as the vector-store implementation.

The final design uses SQLite and Python cosine similarity because Windows Application Control blocked an unsigned native gRPC dependency required by the initial implementation.

The SQLite-based approach:

  • Avoided weakening Windows security controls.
  • Removed the blocked native dependency.
  • Kept the knowledge index fully local.
  • Provided sufficient retrieval performance for the current knowledge-base size.

CPU-First Operation

The project was developed on a Windows laptop with 16 GB of RAM and integrated graphics.

Local response times generally range from approximately 45 to 85 seconds. This limitation is documented as a hardware constraint rather than hidden from the evaluation.

Human-in-the-Loop Design

The assistant does not replace the technician.

Its purpose is to accelerate knowledge retrieval, initial incident analysis, documentation review, and escalation decisions while preserving human approval.

Troubleshooting

ChromaDB Dependency Blocked by Windows Application Control

The initial RAG implementation evaluated ChromaDB as the vector store.

During setup, Windows Application Control blocked an unsigned native gRPC dependency required by the package.

Rather than weakening the host security configuration, the project was redesigned to use a lightweight SQLite-backed vector index with Python cosine similarity.

This change:

  • Preserved the existing Windows security posture.
  • Removed the blocked native dependency.
  • Kept the knowledge index fully local.
  • Provided sufficient retrieval performance for the current knowledge-base size.

Local Inference Performance

The assistant was developed on a Windows laptop with 16 GB of RAM and integrated graphics.

CPU-only local inference produced response times of approximately 45–85 seconds.

The limitation was documented and accepted for the prototype rather than bypassed with external cloud inference.

Knowledge Retrieval Validation

Because the knowledge base contains multiple technical domains, retrieval behavior was validated independently rather than assuming that the highest-scoring result was always correct.

The validation script tested representative incidents for:

  • Networking.
  • Active Directory.
  • Microsoft Intune.
  • Windows workstation support.

All four domains returned the expected knowledge source during final validation.

Limitations

  • The current prototype uses a small project-derived knowledge base.
  • It does not connect directly to Active Directory, Intune, GLPI, osTicket, or production devices.
  • It does not automatically execute PowerShell or system commands.
  • Local CPU inference is slower than cloud-hosted models.
  • Recommendations depend on the quality of the incident description and available documentation.
  • Generated guidance must be reviewed before use.
  • The current retrieval system is optimized for a small local knowledge base rather than enterprise-scale document collections.

Skills Demonstrated

  • Python development
  • Retrieval-Augmented Generation (RAG)
  • Local AI inference with Ollama
  • Semantic search and vector embeddings
  • SQLite data management
  • Prompt and context engineering
  • Human-in-the-loop AI design
  • AI safety controls and command validation
  • IT incident triage
  • Knowledge-base design
  • Windows troubleshooting
  • Active Directory troubleshooting concepts
  • Networking and DNS troubleshooting
  • Microsoft Intune troubleshooting concepts
  • PowerShell diagnostics
  • Multi-domain retrieval validation
  • Technical documentation
  • Git and GitHub

Future Improvements

  • Add additional runbooks from completed IT projects.
  • Generate ticket notes for GLPI or osTicket.
  • Add authenticated and controlled PowerShell data collection.
  • Add role-based access controls.
  • Add automated retrieval-quality evaluations.
  • Add technician feedback and approval tracking.
  • Evaluate faster local models or GPU acceleration.
  • Expand the knowledge base with additional networking, Windows, Active Directory, and endpoint-management scenarios.

Portfolio Context

This project extends previous work in:

  • Windows administration.
  • Active Directory.
  • Networking and DNS.
  • Microsoft Intune and Entra ID.
  • PowerShell automation.
  • Help Desk operations.
  • Local artificial intelligence.

It demonstrates the practical use of AI within IT Operations rather than a generic chatbot implementation.

The project connects infrastructure, troubleshooting, documentation, automation, and AI-assisted decision support into a single technician-focused workflow.

Author

Carlos Cabrera
CompTIA A+ Certified | IT Support | Windows Administration | Networking | PowerShell | AI-Assisted IT Operations

About

Local AI-powered IT support assistant using Python, Streamlit, Ollama, RAG, SQLite vector search, and human-in-the-loop safety controls.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages