Skip to content

Repository files navigation

LLM Adaptation Lab

End-to-end LLM adaptation comparison covering Base Inference, Prompting, RAG, LoRA, QLoRA, GGUF export, Ollama deployment, and Streamlit evaluation dashboard.

Python Ollama HuggingFace QLoRA Streamlit

Pareto Frontier Quality vs latency across adaptation methods. LoRA gives the strongest quality-latency tradeoff, while QLoRA remains a memory-efficient alternative.


Project Overview

This project compares multiple LLM adaptation strategies in one reproducible workflow.

Base → Prompting → RAG → LoRA → QLoRA → Evaluation → GGUF → Ollama → Dashboard

Unlike many LLM projects that stop at prompting or fine-tuning, this repo completes the loop from experimentation to local deployment and evaluation.


Key Differentiators

Area What this repo demonstrates
Multi-method comparison Base, Prompting, RAG, LoRA, and QLoRA in one pipeline
Fair evaluation Same questions and metrics across all methods
LoRA vs QLoRA analysis Compared using quality, latency, loss, and VRAM notes
Deployment loop Fine-tuned model converted to GGUF and served through Ollama
Live local testing Streamlit dashboard can compare base vs fine-tuned Ollama models live
Reliability checks Forgetting tracker and dataset contamination probe
Interpretability Adapter SVD analysis for LoRA/QLoRA weights

Methods Compared

Method Purpose
Base Inference Raw model behavior
Prompting No-training adaptation
RAG External knowledge grounding
LoRA Parameter-efficient fine-tuning
QLoRA Memory-efficient 4-bit fine-tuning
GGUF + Ollama Local fine-tuned model deployment

Evaluation Summary

A small fixed evaluation set is used to demonstrate the complete pipeline.

Method Outputs Avg Keyword Score Avg Composite Score Avg Latency Key Observation
Base 5 0.44 0.40 18.98s Raw model often confused LoRA with other terms
Prompting 25 0.28-0.36 0.21-0.40 9.92s-20.85s Prompt wording changed output quality and latency
RAG 5 0.64 0.67 9.14s Strong grounding, but retrieval adds latency
LoRA 5 0.80 0.86 5.31s Best quality-latency tradeoff
QLoRA 5 0.76 0.84 7.01s Close to LoRA with better memory efficiency
Total 45 - - - Combined evaluation outputs

Keyword score measures expected concept coverage, while composite score combines quality, latency, and answer length.

Core Finding

LoRA achieved the best quality-latency tradeoff with the highest keyword score, highest composite score, and lowest latency among the top methods. QLoRA delivered close performance with better memory efficiency, while RAG improved grounding but added retrieval latency.


Dashboard

Dashboard Preview

Run locally:

python -m streamlit run dashboard/app.py

Dashboard features:

  • method summary
  • question-level comparison
  • response viewer
  • prompt sensitivity analysis
  • LoRA vs QLoRA loss plot
  • LoRA vs QLoRA VRAM plot
  • forgetting tracker
  • adapter SVD plot
  • contamination plot
  • CSV download buttons
  • live Ollama comparison

Note: Live Ollama comparison works locally when Ollama is running on your machine. Streamlit Cloud cannot access your laptop's local Ollama server, so this feature is mainly for local demo.


Live Ollama Demo

After GGUF conversion, the fine-tuned model was loaded into Ollama as:

llm-adaptation-tuned

Manual comparison:

ollama run llama3.2:3b "Explain LoRA in simple terms."
ollama run llm-adaptation-tuned "Explain LoRA in simple terms."

Batch comparison:

python 07_gguf_export/ollama_compare.py

The generated comparison file is saved at:

07_gguf_export/ollama_gguf_comparison.csv

Tech Stack

Area Tools
Inference Ollama
Fine-tuning Hugging Face Transformers, PEFT
QLoRA bitsandbytes
RAG ChromaDB, Sentence Transformers
Evaluation Custom keyword,composite scoring
Visualization matplotlib
Dashboard Streamlit
Deployment GGUF, Ollama

Project Structure

00_baseline/        base model inference
01_prompting/       prompt experiments
02_rag/             RAG pipeline
03_lora/            LoRA training and results
04_qlora/           QLoRA training and results
05_evaluation/      evaluation CSVs
06_analysis/        analysis scripts and plots
07_gguf_export/     GGUF export and Ollama comparison
assets/             generated plots
dashboard/          Streamlit dashboard
data/               prompts and training data
docs/               VRAM notes
paper_notes/        short method notes

Key Outputs

File Purpose
05_evaluation/evaluation_matrix.csv All evaluated responses
05_evaluation/method_summary.csv Method-level scores
05_evaluation/vram_summary.csv LoRA vs QLoRA memory notes
05_evaluation/adapter_svd_summary.csv Adapter SVD results
05_evaluation/contamination_probe.csv Train/eval overlap check
07_gguf_export/ollama_gguf_comparison.csv Base vs fine-tuned Ollama comparison
docs/VRAM_USAGE.md VRAM documentation

Run Pipeline

The deployed dashboard uses lightweight requirements and displays saved CSVs, plots, and analysis results.

pip install -r requirements.txt
python -m streamlit run dashboard/app.py

For full local pipeline:

ollama pull llama3.2:3b

python 00_baseline/run_baseline.py
python 01_prompting/run_prompting.py
python 02_rag/run_rag.py
python 05_evaluation/build_evaluation_matrix.py
python 06_analysis/pareto_frontier.py
python -m streamlit run dashboard/app.py

Advanced Analysis

Analysis Key Finding
Prompt sensitivity Prompt wording changed output quality and latency across variants.
Pareto frontier LoRA achieved the strongest quality-latency tradeoff in this small evaluation.
LoRA vs QLoRA QLoRA gave close quality to LoRA with better memory efficiency through 4-bit loading.
Forgetting tracker Fine-tuned models were checked for basic general-knowledge retention.
Adapter SVD LoRA/QLoRA adapter matrices were inspected through singular value energy concentration.
Contamination probe Evaluation prompts were checked against training data for overlap risk.
Rank ablation notes Rank 8 was completed; rank 4 and rank 16 are documented as future comparison points.
Paper notes Short notes connect the implementation to LoRA, QLoRA, RAG, and instruction tuning.

Deployment Notes

  • GGUF model files are excluded from GitHub because they are large.
  • Adapter/model artifacts are ignored using .gitignore.
  • Streamlit Cloud version shows saved CSVs, plots, and dashboard analysis.
  • Local version additionally supports live Ollama comparison.

Limitations

  • Evaluation set is intentionally small.
  • Training data is small and project-focused.
  • Rank 8 is completed; rank 4 and rank 16 are documented as planned.
  • This is a prototype-level comparison, not a large benchmark.

Future Work

  • Expand evaluation set
  • Complete rank 4 and rank 16 training
  • Add human evaluation
  • Train on larger domain data
  • Integrate full PDF RAG pipeline

About

This is an end-to-end LLM adaptation pipeline with RAG, LoRA, QLoRA, GGUF deployment, Ollama inference, evaluation, and Streamlit dashboard.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages