NoteScanner is a FastAPI + React platform for uploading notes, extracting text, indexing with embeddings, and generating grounded learning outputs.
Core capabilities:
- Upload PDFs/images/text and extract text with Vision OCR.
- Store persistent note chunks + embeddings in ChromaDB for retrieval.
- Support temporary chat-only uploads from the chatbot
+button (ephemeral cache, no DB persistence). - Answer questions and generate flashcards/MCQs/mind maps from retrieved context.
NoteScanner uploads course materials, indexes them in ChromaDB, and uses Sarvam AI: Document Intelligence (Sarvam Vision) for PDF/image OCR and Chat Completions (sarvam-30b / sarvam-105b) for note-grounded Q&A, flashcards, MCQs, and mind maps. Tesseract remains a fallback extractor. It features a FastAPI backend and a React/Vite frontend, orchestrated with Docker Compose.
You upload your handwritten notes → the system scans, organizes, and lets you chat with your notes using Retrieval-Augmented Generation (RAG).
NoteScanner Demo
- Docker & Docker Compose
- Python 3.10+
- Node.js 20+
frontend/my-app: React + Vite UIbackend/api.py: FastAPI API layerbackend/chroma_store.py: Chroma client + collection accessorsbackend/ingest_api.py: chunking + embedding ingestionbackend/llm_pipeline.py: OCR + LLM callsbackend/vfs_tree.py: virtual folder tree helpersbackend/file_meta.py: file metadata policy layer
Endpoint: POST /upload_note
Flow:
- Client uploads file to backend.
- Backend reads file bytes in memory.
- Text extraction:
- PDF: Sarvam Document Intelligence first, parser fallback.
- Image: Sarvam Vision OCR first, Tesseract fallback.
- Plain text formats: UTF-8 decode.
- Backend persists:
- full extracted text in
user_documents, - chunked text + embeddings in
user_{user_id}_notes.
- VFS is updated for explorer visibility.
Endpoint: POST /chat/upload_ephemeral
Flow:
- Frontend sends file with
chat_session_id. - Backend extracts text in memory.
- Backend stores extracted text in process memory cache keyed by
(user_id, chat_session_id).
Important:
- No
user_documentswrite. - No
user_{user_id}_noteswrite. - No Chroma persistence for this path.
Cache clear:
POST /chat/session/clear- Triggered by chat refresh/new chat/unmount.
Endpoint: POST /query_folder
Retrieval order:
- Ephemeral chat uploads (
+cache) - Opened file in viewer
- Other files in same course folder
Selection logic:
- Embedding-based relevance scoring is computed per stage.
- The highest-priority relevant stage is selected.
- If relevance is weak at a higher stage, system falls to next stage.
- This reduces unrelated-domain leakage (for example RL queries answered from generic ML notes).
Endpoints:
POST /study/generate(task:flashcardsormcq)POST /study/mindmap
Context policy:
- Opened file is primary context.
- Same-course files can be included as supporting context.
POST /registerPOST /loginGET /meGET /guest_id
GET /list_treePOST /create_folderPOST /create_filePOST /upload_notePOST /move_pathPOST /delete_pathDELETE /notesGET /file_metaPOST /file_meta
POST /query_folderPOST /chat/upload_ephemeralPOST /chat/session/clear
POST /study/generatePOST /study/mindmap
GET /integrations/onenote/statusGET /integrations/onenote/auth_urlGET /integrations/onenote/callbackPOST /integrations/onenote/sync
Chroma is used as central storage for both metadata and vector retrieval.
Collections:
users: email, hashed password, namesessions: session_id -> user_iduser_documents: full extracted text per logical pathuser_{user_id}_notes: chunk docs + embeddings + chunk metadatauser_{user_id}_vfs: virtual tree + file meta JSONonenote_tokens: OAuth tokensonenote_oauth_states: OAuth state for callback
Notes:
- Guest mode uses effective IDs like
guest_<uuid>. - Persistent note uploads are stored/indexed.
- Chatbot
+uploads are intentionally ephemeral.
From backend/llm_pipeline.py:
- OCR:
- Sarvam Document Intelligence / Vision for PDF-image text extraction.
- Local fallback extractors where needed.
- Q&A:
sarvam_rag_answer(question, context)
- Study generation:
cheap_study_json(context, task, n)for flashcards/MCQmind_map_json(context)for mind maps
Model env controls:
SARVAM_MODEL_RAG(default:sarvam-105b)SARVAM_MODEL_STUDY(default:sarvam-30b)
Required:
SARVAM_API_KEY
Optional Sarvam:
SARVAM_API_BASE(defaulthttps://api.sarvam.ai)SARVAM_MODEL_RAGSARVAM_MODEL_STUDYSARVAM_DOC_INTEL_LANGUAGE(defaulten-IN)SARVAM_DOC_INTEL_TIMEOUT(default180)
Chroma options:
- Cloud:
CHROMA_API_KEYCHROMA_TENANTCHROMA_DATABASE
- Self-hosted:
CHROMA_HOSTCHROMA_PORTCHROMA_SSL(optional)
Other:
USER_DOCUMENT_MAX_BYTESMICROSOFT_CLIENT_ID,MICROSOFT_CLIENT_SECRET,MICROSOFT_REDIRECT_URI(OneNote)
docker compose up --buildServices:
- Backend:
http://localhost:8000 - Frontend:
http://localhost:5173
Rebuild backend after backend changes:
docker compose build backend && docker compose up -d backendBackend:
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt
uvicorn backend.api:app --host 0.0.0.0 --port 8000Frontend:
cd frontend/my-app
npm install
npm run dev -- --host- Chatbot
+uploads are temporary per chat session. - Temporary uploads are not persisted in Chroma collections.
- Query routing order is: chat cache -> open file -> same-course files.
- Chat refresh/new chat clears temporary cache.
backend/API, ingestion, retrieval, auth, VFS, LLM integrationfrontend/my-app/UI and interaction layerdocker-compose.ymllocal orchestrationrequirements.txtbackend dependencies