DocLens AI turns research papers into an evidence-grounded research workspace. Researchers can upload papers, organize them into workspaces and collections, search across documents, ask questions, compare papers, generate literature reviews, and trace answers back to page-level source evidence.
The platform follows one core rule:
No evidence = no claim
The implementation is intentionally separated into a lightweight application control plane and a Python AI service. The NestJS backend owns authentication, authorization, uploads, persistence, and API contracts. The FastAPI service owns PDF extraction, embeddings, retrieval, generation, and verification.
- Workspace, collection, and paper library management
- Authenticated PDF upload with processing status
- PyMuPDF text extraction with page metadata
- Section-aware text chunking
- Native PDF table evidence extraction
- Figure/caption evidence representation
- BGE-M3 embeddings stored in PostgreSQL with pgvector
- Hybrid semantic and keyword retrieval
- Reciprocal Rank Fusion (RRF)
- BGE cross-encoder reranking
- Conversational multi-paper Q&A
- Follow-up question rewriting and bounded multi-query retrieval
- Structured Gemini answers
- Page- and chunk-level citations
- Citation overlap validation and claim verification
- Paper comparison
- Literature-review generation
- Research notes and reading progress
- Retrieval, citation, grounding, and latency evaluation
The current ingestion path does not use Docling, Semantic Scholar, a separate vector database, or a mandatory knowledge-graph processing stage.
Browser
β
βΌ
ββββββββββββββββββββββββββββββββ
β React + Vite + Tailwind β
β Research workspace UI β
ββββββββββββββββ¬ββββββββββββββββ
β REST / WebSocket
βΌ
ββββββββββββββββββββββββββββββββ
β NestJS Backend β
β Auth, ACL, APIs, persistence β
β Uploads, jobs, AI proxy β
βββββββββ¬βββββββββββββββ¬ββββββββββ
β β
βΌ βΌ
ββββββββββββββββββ ββββββββββββββββββββββββ
β Redis β β FastAPI AI Service β
β Cache/session β β Ingestion and RAG β
β job coordinationβ β Models and evaluationβ
ββββββββββββββββββ ββββββββββββ¬ββββββββββββ
β
βΌ
ββββββββββββββββββββββββββββββββββββ
β PostgreSQL + pgvector β
β Application data, chunks, vectors β
ββββββββββββββββββββββββββββββββββββ
β²
β
ββββββββββββββββββββββββ΄βββββββββββββ
β Persistent PDF upload volume β
β Backend read/write, AI read-only β
βββββββββββββββββββββββββββββββββββββ
| Service | Responsibility |
|---|---|
| Frontend | Research UI, uploads, chat, citations, comparison, reviews |
| NestJS backend | Authentication, authorization, API contracts, uploads, persistence, orchestration |
| FastAPI AI service | PDF extraction, chunking, embeddings, retrieval, reranking, generation, verification |
| PostgreSQL | Users, workspaces, documents, chunks, chats, citations, reviews, notes |
| pgvector | 1024-dimensional BGE-M3 chunk embeddings |
| Redis | Cache, session support, and processing coordination |
| Upload volume | Persistent PDF files shared between backend and AI service |
DocLens-AI/
βββ frontend/
β βββ src/
β β βββ components/ # Reusable UI components
β β βββ contexts/ # Auth and application state
β β βββ pages/ # Workspace, library, chat, comparison, review views
β β βββ services/ # Backend API clients
β β βββ hooks/ # Frontend hooks
β β βββ types/ # Frontend types
β β βββ App.* # Application shell and routes
β β βββ main.* # Browser entrypoint
β βββ Dockerfile
β βββ nginx.conf
β βββ package.json
β βββ vite.config.*
β
βββ backend/
β βββ src/
β β βββ auth/ # Registration, login, JWT, guards
β β βββ users/ # User APIs
β β βββ workspaces/ # Workspace APIs
β β βββ collections/ # Collection APIs
β β βββ documents/ # Uploads, document metadata, access checks
β β βββ processing/ # Ingestion and embedding job orchestration
β β βββ query/ # Search, chat, comparisons, literature reviews
β β βββ ai-proxy/ # NestJS-to-FastAPI client
β β βββ gateway/ # WebSocket updates
β β βββ prisma/ # Prisma service and database access
β β βββ common/ # Shared guards, DTOs, and utilities
β β βββ config/ # Environment configuration
β β βββ app.module.ts
β β βββ main.ts
β βββ prisma/
β β βββ schema.prisma
β β βββ migrations/
β βββ Dockerfile
β βββ package.json
β βββ tsconfig.json
β
βββ ai-service/
β βββ main.py # FastAPI endpoints
β βββ ingest.py # PyMuPDF extraction, chunking, embeddings
β βββ query.py # AI use cases: ask, summarize, compare, review
β βββ requirements.txt
β βββ Dockerfile
β βββ rag/
β β βββ chain.py # Conversational RAG and structured generation
β β βββ retriever.py # Hybrid candidates, RRF, reranking adapter
β β βββ schemas.py # Structured answers, claims, citations, reviews
β β βββ verification.py # Claim and citation validation
β β βββ evaluation.py # Retrieval and grounding metrics
β β βββ test_phase_pipeline.py
β βββ vector_store/
β βββ pg_store.py # PostgreSQL/pgvector persistence and search
β
βββ docker-compose.yml # PostgreSQL, Redis, migrations, backend, AI, frontend
βββ package.json # Workspace commands
βββ .env.example # Environment variable template
βββ README.md
Ingestion is intentionally separate from answering questions.
PDF upload
β
βΌ
NestJS validation and persistent file storage
β
βΌ
Processing job
β
βΌ
FastAPI POST /ingest
β
βΌ
PyMuPDF extraction
β
βββ Page text and document metadata
βββ Section detection
βββ Section-aware text chunks
βββ Native table evidence
βββ Figure/caption evidence
β
βΌ
DocumentChunk records in PostgreSQL
β
βΌ
FastAPI POST /embed
β
βΌ
BAAI/bge-m3, 1024 dimensions
β
βΌ
pgvector + EmbeddingMetadata
β
βΌ
Document status = READY
Every persisted chunk contains content plus metadata such as:
{
"pageNumber": 7,
"chunkIndex": 12,
"contentType": "table",
"section": "Experiments",
"contentHash": "..."
}Text chunks preserve page and section context. Native tables are represented as searchable Markdown-like content with table indexes. Figure evidence preserves the page, caption, image count, and section. This allows answers to identify visual evidence and connect it to surrounding paper text.
The base ingestion path records figure evidence and captions; full pixel-level interpretation of arbitrary charts, diagrams, and scanned pages is not assumed for every document.
The canonical schema is in backend/prisma/schema.prisma.
User
βββ Workspace
βββ Collection
βββ Document
βββ DocumentChunk
βββ EmbeddingMetadata
βββ ProcessingJob
βββ DocumentEntity
βββ Relationship
User
βββ ChatSession ββ Query ββ Citation ββ DocumentChunk
βββ PaperComparison
βββ LiteratureReview
βββ ResearchNote
βββ ReadingProgress
Important models:
User: identity, role, authentication, and ownership relationshipsWorkspace: top-level research environmentCollection: thematic grouping of documentsDocument: uploaded PDF metadata and processing stateDocumentChunk: persistent text/table/figure evidence unitEmbeddingMetadata: model, dimensions, vector-store, and chunk linkageChatSession,Query, andCitation: conversational history and traceabilityPaperComparison: multi-paper comparison resultLiteratureReview: generated review sections and Markdown outputResearchNote: user-authored research notesReadingProgress: document reading state
Document processing states include:
PENDING β UPLOADED β EXTRACTING β CHUNKING β EMBEDDING β INDEXING β READY
ββββββββββββββββββββββββββββββββ FAILED
Question and conversation history
β
βΌ
Conditional standalone-question rewrite
β
βΌ
Bounded multi-query expansion
β
βΌ
Query embedding with BGE-M3
β
ββββββββ΄βββββββ
βΌ βΌ
Semantic search Keyword search
β β
ββββββββ¬βββββββ
βΌ
Reciprocal Rank Fusion, k = 60
β
βΌ
BGE cross-encoder reranking
β
βΌ
Evidence context construction
β
βΌ
Gemini structured generation
β
βΌ
Claim verification
β
βΌ
Citation validation
β
βΌ
Grounded answer with citations
The question is embedded with BAAI/bge-m3 and compared with chunk vectors in
pgvector. This handles conceptually similar wording.
The chunk text is searched for exact terms. This is important for model names, dataset names, acronyms, equations, identifiers, and numeric table values.
Independent semantic and keyword candidate lists are combined using:
RRF contribution = 1 / (60 + rank)
Chunks present in both lists receive contributions from both retrieval signals. Source and rank provenance are preserved in the result metadata.
The fused candidates are scored with:
BAAI/bge-reranker-v2-m3
The reranker evaluates the complete (question, chunk) pair before final
top-K evidence selection.
Retrieved chunks are converted into explicit evidence blocks:
[Evidence 1]
chunk_id: ...
document_id: ...
document_title: ...
page: 8
evidence_type: table
section: Experiments
---
Original chunk content
Gemini receives the evidence, question, and conversation context through the structured LangChain pipeline. The generation schema requires:
- Claims to be supported by retrieved evidence
- Exact chunk and document identifiers
- Near-verbatim source excerpts
- Explicit insufficient-evidence responses
- Preservation of table labels and values
- Separation of visible figure observations from author-reported conclusions
The AI response is parsed into typed Pydantic models rather than treated as unstructured text.
The verification layer runs after generation.
Generated answer
β
βΌ
Extract claims and citations
β
βΌ
Check chunk/document identity
β
βΌ
Check non-empty source text
β
βΌ
Check source-text overlap
β
βΌ
Reject or remove unsupported citations/claims
β
βΌ
Persist final answer and citations
Validation rejects unknown chunk IDs, empty source excerpts, and citations with insufficient overlap against the retrieved source content. Evaluation tracks unsupported-claim rate and citation coverage.
Questions can be scoped to a collection or selected document IDs. Access control is enforced by NestJS before the AI service is called.
The comparison workflow synthesizes selected papers across:
- Methods
- Datasets
- Models
- Metrics
- Findings
- Similarities
- Differences
- Limitations
The literature-review workflow retrieves evidence from selected papers and generates structured sections plus Markdown output.
The schema supports research entities and relationships such as authors, concepts, datasets, methods, models, and metrics. Entity and relationship queries are available through the AI service. Advanced graph exploration features are not part of the core RAG path.
The FastAPI entrypoint is ai-service/main.py.
GET /health
POST /ingest
POST /embed
GET /status/{document_id}
POST /search
POST /search/semantic
POST /search/chunk
POST /search/hybrid
POST /ask
POST /summarise
POST /review
POST /compare
POST /literature-review
POST /kg/entities
POST /kg/relationships
POST /evaluate
The NestJS backend exposes the public /api/v1 routes and uses the AI service
as an internal dependency.
The evaluation implementation is in ai-service/rag/evaluation.py.
- Recall@K
- Mean Reciprocal Rank (MRR)
- Reranking improvement
- Faithfulness
- Answer relevance
- Unsupported-claim rate
- Citation correctness
- Citation precision
- Citation coverage
- Retrieval and generation latency
- Processing latency
Negative-evidence tests verify that questions with no supporting source material produce transparent insufficient-evidence responses instead of fabricated claims.
Docker Compose runs:
postgres
redis
backend-migrate
backend
ai-service
frontend
Startup order:
PostgreSQL and Redis
β
Prisma migrations
β
FastAPI AI service
β
NestJS backend
β
React/Nginx frontend
The AI container caches Hugging Face models in a persistent volume so BGE-M3 and the reranker are not downloaded on every restart.
Required production configuration includes:
DATABASE_URL
POSTGRES_PASSWORD
JWT_SECRET
INTERNAL_API_SECRET
GEMINI_API_KEY
CORS_ORIGINS
See .env.example for the available configuration.
- The frontend never accesses PostgreSQL directly.
- The backend owns authentication and collection/document authorization.
- The AI service is an internal service behind the NestJS control plane.
- Uploaded PDFs are stored in a persistent backend volume.
- The AI service mounts uploaded files read-only.
- DTO validation and access checks occur before AI requests.
- Processing failures are represented in document/job status rather than silently marking a document ready.
- Retrieval provenance, page metadata, chunk IDs, and source text are preserved for traceability.
Install JavaScript dependencies:
npm installGenerate the Prisma client:
npm run prisma:generateStart infrastructure and services:
docker compose up --buildUseful validation commands:
npm run build:all
npm --prefix backend run test
python -m compileall -q ai-service
python -m pytest ai-service\rag\test_phase_pipeline.py -q
docker compose configThe implemented pipeline is:
Ingestion
β Embeddings
β Hybrid retrieval
β RRF
β Cross-encoder reranking
β Multi-query retrieval
β Evidence construction
β Gemini structured generation
β Citation validation
β Claim verification
β Evaluation
The architecture avoids adding separate orchestration frameworks, a second vector database, or heavyweight document-processing dependencies. The main quality boundary is the evidence metadata preserved from ingestion through the final answer.