Skip to content

About

Project COS783

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

COS783 Project

Setup

python -m venv .venv
.venv/Scripts/activate
pip install -r requirements.txt

Run

python app.py

Flow

  1. Upload files (txt, pdf, docx, doc)
  2. Extract metadata + keywords (full text is NOT saved in extraction JSON)
  3. Search semantically over extracted keywords
  4. Ask the LM questions over document metadata/keywords, optionally grounded by current semantic-search matches

Semantic search

Use the search box on the home page to find files by meaning (paraphrases work, not only exact words).

The embedding model (all-MiniLM-L6-v2) runs locally after the first download. Hugging Face Hub is only contacted once to cache weights; restart the Flask server after code changes (Ctrl+C, then python app.py again).

LM interface

Start Ollama (ollama serve) and pull a model (ollama pull llama3.2). Then use "Ask language model" in the UI after searching, e.g.: "Which matched document talks about mounting from a USB?"

Structure

uploads/              # raw uploaded files, grouped by UUID
extractions/          # extracted metadata and keywords per file (no text_content)
keyword_extraction.py # keyword and chunk extraction
semantic_search.py    # FAISS + sentence-transformers search
document_repository.py# document metadata/keyword loading + cleanup
lm_assistant.py       # local Ollama integration for question answering
template/             # HTML templates

About

Project COS783

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages