End-to-end LLM adaptation comparison covering Base Inference, Prompting, RAG, LoRA, QLoRA, GGUF export, Ollama deployment, and Streamlit evaluation dashboard.
Quality vs latency across adaptation methods. LoRA gives the strongest quality-latency tradeoff, while QLoRA remains a memory-efficient alternative.
This project compares multiple LLM adaptation strategies in one reproducible workflow.
Base → Prompting → RAG → LoRA → QLoRA → Evaluation → GGUF → Ollama → Dashboard
Unlike many LLM projects that stop at prompting or fine-tuning, this repo completes the loop from experimentation to local deployment and evaluation.
| Area | What this repo demonstrates |
|---|---|
| Multi-method comparison | Base, Prompting, RAG, LoRA, and QLoRA in one pipeline |
| Fair evaluation | Same questions and metrics across all methods |
| LoRA vs QLoRA analysis | Compared using quality, latency, loss, and VRAM notes |
| Deployment loop | Fine-tuned model converted to GGUF and served through Ollama |
| Live local testing | Streamlit dashboard can compare base vs fine-tuned Ollama models live |
| Reliability checks | Forgetting tracker and dataset contamination probe |
| Interpretability | Adapter SVD analysis for LoRA/QLoRA weights |
| Method | Purpose |
|---|---|
| Base Inference | Raw model behavior |
| Prompting | No-training adaptation |
| RAG | External knowledge grounding |
| LoRA | Parameter-efficient fine-tuning |
| QLoRA | Memory-efficient 4-bit fine-tuning |
| GGUF + Ollama | Local fine-tuned model deployment |
A small fixed evaluation set is used to demonstrate the complete pipeline.
| Method | Outputs | Avg Keyword Score | Avg Composite Score | Avg Latency | Key Observation |
|---|---|---|---|---|---|
| Base | 5 | 0.44 | 0.40 | 18.98s | Raw model often confused LoRA with other terms |
| Prompting | 25 | 0.28-0.36 | 0.21-0.40 | 9.92s-20.85s | Prompt wording changed output quality and latency |
| RAG | 5 | 0.64 | 0.67 | 9.14s | Strong grounding, but retrieval adds latency |
| LoRA | 5 | 0.80 | 0.86 | 5.31s | Best quality-latency tradeoff |
| QLoRA | 5 | 0.76 | 0.84 | 7.01s | Close to LoRA with better memory efficiency |
| Total | 45 | - | - | - | Combined evaluation outputs |
Keyword score measures expected concept coverage, while composite score combines quality, latency, and answer length.
LoRA achieved the best quality-latency tradeoff with the highest keyword score, highest composite score, and lowest latency among the top methods. QLoRA delivered close performance with better memory efficiency, while RAG improved grounding but added retrieval latency.
Run locally:
python -m streamlit run dashboard/app.pyDashboard features:
- method summary
- question-level comparison
- response viewer
- prompt sensitivity analysis
- LoRA vs QLoRA loss plot
- LoRA vs QLoRA VRAM plot
- forgetting tracker
- adapter SVD plot
- contamination plot
- CSV download buttons
- live Ollama comparison
Note: Live Ollama comparison works locally when Ollama is running on your machine. Streamlit Cloud cannot access your laptop's local Ollama server, so this feature is mainly for local demo.
After GGUF conversion, the fine-tuned model was loaded into Ollama as:
llm-adaptation-tuned
Manual comparison:
ollama run llama3.2:3b "Explain LoRA in simple terms."
ollama run llm-adaptation-tuned "Explain LoRA in simple terms."Batch comparison:
python 07_gguf_export/ollama_compare.pyThe generated comparison file is saved at:
07_gguf_export/ollama_gguf_comparison.csv
| Area | Tools |
|---|---|
| Inference | Ollama |
| Fine-tuning | Hugging Face Transformers, PEFT |
| QLoRA | bitsandbytes |
| RAG | ChromaDB, Sentence Transformers |
| Evaluation | Custom keyword,composite scoring |
| Visualization | matplotlib |
| Dashboard | Streamlit |
| Deployment | GGUF, Ollama |
00_baseline/ base model inference
01_prompting/ prompt experiments
02_rag/ RAG pipeline
03_lora/ LoRA training and results
04_qlora/ QLoRA training and results
05_evaluation/ evaluation CSVs
06_analysis/ analysis scripts and plots
07_gguf_export/ GGUF export and Ollama comparison
assets/ generated plots
dashboard/ Streamlit dashboard
data/ prompts and training data
docs/ VRAM notes
paper_notes/ short method notes
| File | Purpose |
|---|---|
05_evaluation/evaluation_matrix.csv |
All evaluated responses |
05_evaluation/method_summary.csv |
Method-level scores |
05_evaluation/vram_summary.csv |
LoRA vs QLoRA memory notes |
05_evaluation/adapter_svd_summary.csv |
Adapter SVD results |
05_evaluation/contamination_probe.csv |
Train/eval overlap check |
07_gguf_export/ollama_gguf_comparison.csv |
Base vs fine-tuned Ollama comparison |
docs/VRAM_USAGE.md |
VRAM documentation |
The deployed dashboard uses lightweight requirements and displays saved CSVs, plots, and analysis results.
pip install -r requirements.txt
python -m streamlit run dashboard/app.pyFor full local pipeline:
ollama pull llama3.2:3b
python 00_baseline/run_baseline.py
python 01_prompting/run_prompting.py
python 02_rag/run_rag.py
python 05_evaluation/build_evaluation_matrix.py
python 06_analysis/pareto_frontier.py
python -m streamlit run dashboard/app.py| Analysis | Key Finding |
|---|---|
| Prompt sensitivity | Prompt wording changed output quality and latency across variants. |
| Pareto frontier | LoRA achieved the strongest quality-latency tradeoff in this small evaluation. |
| LoRA vs QLoRA | QLoRA gave close quality to LoRA with better memory efficiency through 4-bit loading. |
| Forgetting tracker | Fine-tuned models were checked for basic general-knowledge retention. |
| Adapter SVD | LoRA/QLoRA adapter matrices were inspected through singular value energy concentration. |
| Contamination probe | Evaluation prompts were checked against training data for overlap risk. |
| Rank ablation notes | Rank 8 was completed; rank 4 and rank 16 are documented as future comparison points. |
| Paper notes | Short notes connect the implementation to LoRA, QLoRA, RAG, and instruction tuning. |
- GGUF model files are excluded from GitHub because they are large.
- Adapter/model artifacts are ignored using
.gitignore. - Streamlit Cloud version shows saved CSVs, plots, and dashboard analysis.
- Local version additionally supports live Ollama comparison.
- Evaluation set is intentionally small.
- Training data is small and project-focused.
- Rank 8 is completed; rank 4 and rank 16 are documented as planned.
- This is a prototype-level comparison, not a large benchmark.
- Expand evaluation set
- Complete rank 4 and rank 16 training
- Add human evaluation
- Train on larger domain data
- Integrate full PDF RAG pipeline
