Skip to content

About

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

DocLens AI

Citation-First Research Intelligence Platform

DocLens AI turns research papers into an evidence-grounded research workspace. Researchers can upload papers, organize them into workspaces and collections, search across documents, ask questions, compare papers, generate literature reviews, and trace answers back to page-level source evidence.

The platform follows one core rule:

No evidence = no claim

The implementation is intentionally separated into a lightweight application control plane and a Python AI service. The NestJS backend owns authentication, authorization, uploads, persistence, and API contracts. The FastAPI service owns PDF extraction, embeddings, retrieval, generation, and verification.


Core capabilities

  • Workspace, collection, and paper library management
  • Authenticated PDF upload with processing status
  • PyMuPDF text extraction with page metadata
  • Section-aware text chunking
  • Native PDF table evidence extraction
  • Figure/caption evidence representation
  • BGE-M3 embeddings stored in PostgreSQL with pgvector
  • Hybrid semantic and keyword retrieval
  • Reciprocal Rank Fusion (RRF)
  • BGE cross-encoder reranking
  • Conversational multi-paper Q&A
  • Follow-up question rewriting and bounded multi-query retrieval
  • Structured Gemini answers
  • Page- and chunk-level citations
  • Citation overlap validation and claim verification
  • Paper comparison
  • Literature-review generation
  • Research notes and reading progress
  • Retrieval, citation, grounding, and latency evaluation

The current ingestion path does not use Docling, Semantic Scholar, a separate vector database, or a mandatory knowledge-graph processing stage.


System architecture

                                Browser
                                   β”‚
                                   β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚ React + Vite + Tailwind      β”‚
                    β”‚ Research workspace UI        β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                   β”‚ REST / WebSocket
                                   β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚ NestJS Backend                β”‚
                    β”‚ Auth, ACL, APIs, persistence  β”‚
                    β”‚ Uploads, jobs, AI proxy       β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                            β”‚              β”‚
                            β–Ό              β–Ό
                 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                 β”‚ Redis          β”‚  β”‚ FastAPI AI Service   β”‚
                 β”‚ Cache/session  β”‚  β”‚ Ingestion and RAG    β”‚
                 β”‚ job coordinationβ”‚ β”‚ Models and evaluationβ”‚
                 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                                β”‚
                                                β–Ό
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚ PostgreSQL + pgvector             β”‚
                         β”‚ Application data, chunks, vectors β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                                β–²
                                                β”‚
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚ Persistent PDF upload volume       β”‚
                         β”‚ Backend read/write, AI read-only   β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Service responsibilities

Service Responsibility
Frontend Research UI, uploads, chat, citations, comparison, reviews
NestJS backend Authentication, authorization, API contracts, uploads, persistence, orchestration
FastAPI AI service PDF extraction, chunking, embeddings, retrieval, reranking, generation, verification
PostgreSQL Users, workspaces, documents, chunks, chats, citations, reviews, notes
pgvector 1024-dimensional BGE-M3 chunk embeddings
Redis Cache, session support, and processing coordination
Upload volume Persistent PDF files shared between backend and AI service

Repository structure

DocLens-AI/
β”œβ”€β”€ frontend/
β”‚   β”œβ”€β”€ src/
β”‚   β”‚   β”œβ”€β”€ components/       # Reusable UI components
β”‚   β”‚   β”œβ”€β”€ contexts/         # Auth and application state
β”‚   β”‚   β”œβ”€β”€ pages/            # Workspace, library, chat, comparison, review views
β”‚   β”‚   β”œβ”€β”€ services/         # Backend API clients
β”‚   β”‚   β”œβ”€β”€ hooks/            # Frontend hooks
β”‚   β”‚   β”œβ”€β”€ types/            # Frontend types
β”‚   β”‚   β”œβ”€β”€ App.*             # Application shell and routes
β”‚   β”‚   └── main.*            # Browser entrypoint
β”‚   β”œβ”€β”€ Dockerfile
β”‚   β”œβ”€β”€ nginx.conf
β”‚   β”œβ”€β”€ package.json
β”‚   └── vite.config.*
β”‚
β”œβ”€β”€ backend/
β”‚   β”œβ”€β”€ src/
β”‚   β”‚   β”œβ”€β”€ auth/             # Registration, login, JWT, guards
β”‚   β”‚   β”œβ”€β”€ users/            # User APIs
β”‚   β”‚   β”œβ”€β”€ workspaces/       # Workspace APIs
β”‚   β”‚   β”œβ”€β”€ collections/      # Collection APIs
β”‚   β”‚   β”œβ”€β”€ documents/        # Uploads, document metadata, access checks
β”‚   β”‚   β”œβ”€β”€ processing/       # Ingestion and embedding job orchestration
β”‚   β”‚   β”œβ”€β”€ query/            # Search, chat, comparisons, literature reviews
β”‚   β”‚   β”œβ”€β”€ ai-proxy/         # NestJS-to-FastAPI client
β”‚   β”‚   β”œβ”€β”€ gateway/          # WebSocket updates
β”‚   β”‚   β”œβ”€β”€ prisma/           # Prisma service and database access
β”‚   β”‚   β”œβ”€β”€ common/           # Shared guards, DTOs, and utilities
β”‚   β”‚   β”œβ”€β”€ config/           # Environment configuration
β”‚   β”‚   β”œβ”€β”€ app.module.ts
β”‚   β”‚   └── main.ts
β”‚   β”œβ”€β”€ prisma/
β”‚   β”‚   β”œβ”€β”€ schema.prisma
β”‚   β”‚   └── migrations/
β”‚   β”œβ”€β”€ Dockerfile
β”‚   β”œβ”€β”€ package.json
β”‚   └── tsconfig.json
β”‚
β”œβ”€β”€ ai-service/
β”‚   β”œβ”€β”€ main.py               # FastAPI endpoints
β”‚   β”œβ”€β”€ ingest.py             # PyMuPDF extraction, chunking, embeddings
β”‚   β”œβ”€β”€ query.py              # AI use cases: ask, summarize, compare, review
β”‚   β”œβ”€β”€ requirements.txt
β”‚   β”œβ”€β”€ Dockerfile
β”‚   β”œβ”€β”€ rag/
β”‚   β”‚   β”œβ”€β”€ chain.py          # Conversational RAG and structured generation
β”‚   β”‚   β”œβ”€β”€ retriever.py      # Hybrid candidates, RRF, reranking adapter
β”‚   β”‚   β”œβ”€β”€ schemas.py        # Structured answers, claims, citations, reviews
β”‚   β”‚   β”œβ”€β”€ verification.py   # Claim and citation validation
β”‚   β”‚   β”œβ”€β”€ evaluation.py     # Retrieval and grounding metrics
β”‚   β”‚   └── test_phase_pipeline.py
β”‚   └── vector_store/
β”‚       └── pg_store.py       # PostgreSQL/pgvector persistence and search
β”‚
β”œβ”€β”€ docker-compose.yml        # PostgreSQL, Redis, migrations, backend, AI, frontend
β”œβ”€β”€ package.json              # Workspace commands
β”œβ”€β”€ .env.example              # Environment variable template
└── README.md

Document ingestion pipeline

Ingestion is intentionally separate from answering questions.

PDF upload
    β”‚
    β–Ό
NestJS validation and persistent file storage
    β”‚
    β–Ό
Processing job
    β”‚
    β–Ό
FastAPI POST /ingest
    β”‚
    β–Ό
PyMuPDF extraction
    β”‚
    β”œβ”€β”€ Page text and document metadata
    β”œβ”€β”€ Section detection
    β”œβ”€β”€ Section-aware text chunks
    β”œβ”€β”€ Native table evidence
    └── Figure/caption evidence
    β”‚
    β–Ό
DocumentChunk records in PostgreSQL
    β”‚
    β–Ό
FastAPI POST /embed
    β”‚
    β–Ό
BAAI/bge-m3, 1024 dimensions
    β”‚
    β–Ό
pgvector + EmbeddingMetadata
    β”‚
    β–Ό
Document status = READY

Evidence representation

Every persisted chunk contains content plus metadata such as:

{
  "pageNumber": 7,
  "chunkIndex": 12,
  "contentType": "table",
  "section": "Experiments",
  "contentHash": "..."
}

Text chunks preserve page and section context. Native tables are represented as searchable Markdown-like content with table indexes. Figure evidence preserves the page, caption, image count, and section. This allows answers to identify visual evidence and connect it to surrounding paper text.

The base ingestion path records figure evidence and captions; full pixel-level interpretation of arbitrary charts, diagrams, and scanned pages is not assumed for every document.


Data model

The canonical schema is in backend/prisma/schema.prisma.

User
 └── Workspace
      └── Collection
           └── Document
                β”œβ”€β”€ DocumentChunk
                β”œβ”€β”€ EmbeddingMetadata
                β”œβ”€β”€ ProcessingJob
                β”œβ”€β”€ DocumentEntity
                └── Relationship

User
 β”œβ”€β”€ ChatSession ── Query ── Citation ── DocumentChunk
 β”œβ”€β”€ PaperComparison
 β”œβ”€β”€ LiteratureReview
 β”œβ”€β”€ ResearchNote
 └── ReadingProgress

Important models:

  • User: identity, role, authentication, and ownership relationships
  • Workspace: top-level research environment
  • Collection: thematic grouping of documents
  • Document: uploaded PDF metadata and processing state
  • DocumentChunk: persistent text/table/figure evidence unit
  • EmbeddingMetadata: model, dimensions, vector-store, and chunk linkage
  • ChatSession, Query, and Citation: conversational history and traceability
  • PaperComparison: multi-paper comparison result
  • LiteratureReview: generated review sections and Markdown output
  • ResearchNote: user-authored research notes
  • ReadingProgress: document reading state

Document processing states include:

PENDING β†’ UPLOADED β†’ EXTRACTING β†’ CHUNKING β†’ EMBEDDING β†’ INDEXING β†’ READY
                                      └──────────────────────────────→ FAILED

RAG architecture

Question and conversation history
              β”‚
              β–Ό
Conditional standalone-question rewrite
              β”‚
              β–Ό
Bounded multi-query expansion
              β”‚
              β–Ό
Query embedding with BGE-M3
              β”‚
       β”Œβ”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”
       β–Ό             β–Ό
Semantic search   Keyword search
       β”‚             β”‚
       β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
              β–Ό
Reciprocal Rank Fusion, k = 60
              β”‚
              β–Ό
BGE cross-encoder reranking
              β”‚
              β–Ό
Evidence context construction
              β”‚
              β–Ό
Gemini structured generation
              β”‚
              β–Ό
Claim verification
              β”‚
              β–Ό
Citation validation
              β”‚
              β–Ό
Grounded answer with citations

Retrieval stages

Semantic retrieval

The question is embedded with BAAI/bge-m3 and compared with chunk vectors in pgvector. This handles conceptually similar wording.

Keyword retrieval

The chunk text is searched for exact terms. This is important for model names, dataset names, acronyms, equations, identifiers, and numeric table values.

Reciprocal Rank Fusion

Independent semantic and keyword candidate lists are combined using:

RRF contribution = 1 / (60 + rank)

Chunks present in both lists receive contributions from both retrieval signals. Source and rank provenance are preserved in the result metadata.

Cross-encoder reranking

The fused candidates are scored with:

BAAI/bge-reranker-v2-m3

The reranker evaluates the complete (question, chunk) pair before final top-K evidence selection.


Evidence-grounded generation

Retrieved chunks are converted into explicit evidence blocks:

[Evidence 1]
chunk_id: ...
document_id: ...
document_title: ...
page: 8
evidence_type: table
section: Experiments
---
Original chunk content

Gemini receives the evidence, question, and conversation context through the structured LangChain pipeline. The generation schema requires:

  • Claims to be supported by retrieved evidence
  • Exact chunk and document identifiers
  • Near-verbatim source excerpts
  • Explicit insufficient-evidence responses
  • Preservation of table labels and values
  • Separation of visible figure observations from author-reported conclusions

The AI response is parsed into typed Pydantic models rather than treated as unstructured text.


Citation and claim verification

The verification layer runs after generation.

Generated answer
      β”‚
      β–Ό
Extract claims and citations
      β”‚
      β–Ό
Check chunk/document identity
      β”‚
      β–Ό
Check non-empty source text
      β”‚
      β–Ό
Check source-text overlap
      β”‚
      β–Ό
Reject or remove unsupported citations/claims
      β”‚
      β–Ό
Persist final answer and citations

Validation rejects unknown chunk IDs, empty source excerpts, and citations with insufficient overlap against the retrieved source content. Evaluation tracks unsupported-claim rate and citation coverage.


Research workflows

Multi-paper Q&A

Questions can be scoped to a collection or selected document IDs. Access control is enforced by NestJS before the AI service is called.

Paper comparison

The comparison workflow synthesizes selected papers across:

  • Methods
  • Datasets
  • Models
  • Metrics
  • Findings
  • Similarities
  • Differences
  • Limitations

Literature reviews

The literature-review workflow retrieves evidence from selected papers and generates structured sections plus Markdown output.

Knowledge graph data

The schema supports research entities and relationships such as authors, concepts, datasets, methods, models, and metrics. Entity and relationship queries are available through the AI service. Advanced graph exploration features are not part of the core RAG path.


AI service API

The FastAPI entrypoint is ai-service/main.py.

GET  /health

POST /ingest
POST /embed
GET  /status/{document_id}

POST /search
POST /search/semantic
POST /search/chunk
POST /search/hybrid

POST /ask
POST /summarise
POST /review
POST /compare
POST /literature-review

POST /kg/entities
POST /kg/relationships

POST /evaluate

The NestJS backend exposes the public /api/v1 routes and uses the AI service as an internal dependency.


Evaluation

The evaluation implementation is in ai-service/rag/evaluation.py.

Retrieval metrics

  • Recall@K
  • Mean Reciprocal Rank (MRR)
  • Reranking improvement

Generation and grounding metrics

  • Faithfulness
  • Answer relevance
  • Unsupported-claim rate

Citation metrics

  • Citation correctness
  • Citation precision
  • Citation coverage

Operational metrics

  • Retrieval and generation latency
  • Processing latency

Negative-evidence tests verify that questions with no supporting source material produce transparent insufficient-evidence responses instead of fabricated claims.


Deployment

Docker Compose runs:

postgres
redis
backend-migrate
backend
ai-service
frontend

Startup order:

PostgreSQL and Redis
        ↓
Prisma migrations
        ↓
FastAPI AI service
        ↓
NestJS backend
        ↓
React/Nginx frontend

The AI container caches Hugging Face models in a persistent volume so BGE-M3 and the reranker are not downloaded on every restart.

Required production configuration includes:

DATABASE_URL
POSTGRES_PASSWORD
JWT_SECRET
INTERNAL_API_SECRET
GEMINI_API_KEY
CORS_ORIGINS

See .env.example for the available configuration.


Security and reliability boundaries

  • The frontend never accesses PostgreSQL directly.
  • The backend owns authentication and collection/document authorization.
  • The AI service is an internal service behind the NestJS control plane.
  • Uploaded PDFs are stored in a persistent backend volume.
  • The AI service mounts uploaded files read-only.
  • DTO validation and access checks occur before AI requests.
  • Processing failures are represented in document/job status rather than silently marking a document ready.
  • Retrieval provenance, page metadata, chunk IDs, and source text are preserved for traceability.

Local development

Install JavaScript dependencies:

npm install

Generate the Prisma client:

npm run prisma:generate

Start infrastructure and services:

docker compose up --build

Useful validation commands:

npm run build:all
npm --prefix backend run test
python -m compileall -q ai-service
python -m pytest ai-service\rag\test_phase_pipeline.py -q
docker compose config

Current implementation boundary

The implemented pipeline is:

Ingestion
β†’ Embeddings
β†’ Hybrid retrieval
β†’ RRF
β†’ Cross-encoder reranking
β†’ Multi-query retrieval
β†’ Evidence construction
β†’ Gemini structured generation
β†’ Citation validation
β†’ Claim verification
β†’ Evaluation

The architecture avoids adding separate orchestration frameworks, a second vector database, or heavyweight document-processing dependencies. The main quality boundary is the evidence metadata preserved from ingestion through the final answer.

About

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages