python -m venv .venv
.venv/Scripts/activate
pip install -r requirements.txtpython app.py- Upload files (
txt,pdf,docx,doc) - Extract metadata + keywords (full text is NOT saved in extraction JSON)
- Search semantically over extracted keywords
- Ask the LM questions over document metadata/keywords, optionally grounded by current semantic-search matches
Use the search box on the home page to find files by meaning (paraphrases work, not only exact words).
The embedding model (all-MiniLM-L6-v2) runs locally after the first download. Hugging Face Hub is only contacted once to cache weights; restart the Flask server after code changes (Ctrl+C, then python app.py again).
Start Ollama (ollama serve) and pull a model (ollama pull llama3.2).
Then use "Ask language model" in the UI after searching, e.g.:
"Which matched document talks about mounting from a USB?"
uploads/ # raw uploaded files, grouped by UUID
extractions/ # extracted metadata and keywords per file (no text_content)
keyword_extraction.py # keyword and chunk extraction
semantic_search.py # FAISS + sentence-transformers search
document_repository.py# document metadata/keyword loading + cleanup
lm_assistant.py # local Ollama integration for question answering
template/ # HTML templates