A local, offline-first English tutor powered by LLMs. Runs entirely on your machine — no internet required, no data leaves your computer.
LexiMind is designed for kids aged 9–12 at an A1–A2 English level. It uses GGUF quantized models via llama-cpp-python to deliver grammar explanations, vocabulary practice, short stories, and conversational exercises through a clean chat interface.
- 100% offline — once downloaded, everything runs locally. No API keys, no cloud, no tracking.
- Streaming responses — tokens appear in real time as the model generates them.
- Persistent chat history — conversations are saved in SQLite and restored when you come back.
- Multiple sessions — start new chats, switch between them, delete old ones.
- Topic shortcuts — one-click prompts for colors, animals, numbers, and short stories.
- Custom username — editable profile name, stored in your browser.
- PDF & OCR support (via
pdfjs-distandtesseract.json the frontend). - Portable build — packaged as a single
.exewith PyInstaller. No Python installation needed.
| Layer | Technology |
|---|---|
| Backend | Python 3, FastAPI, uvicorn |
| LLM runtime | llama-cpp-python |
| Storage | SQLite |
| Frontend | Svelte 5, Vite |
| Styling | Tailwind CSS, Lucide icons |
| Distribution | PyInstaller |
Download the model and place it inside a model/ folder in the app directory. The app looks for model/*.gguf at startup.
| Model | File size |
|---|---|
| microsoft_Phi-4-mini-instruct-Q6_K.gguf | 3.16 GB |
How to download: open the link above, click the "Download" button on the HuggingFace page, and save the file as
microsoft_Phi-4-mini-instruct-Q6_K.ggufinside themodel/folder.
Older models like Phi-3 use MHA (Multi-Head Attention) with 32 key/value heads, which creates a large KV cache (~804 MB) and high RAM usage. Phi-4-mini solves this with two architectural improvements:
- GQA (Grouped Query Attention): 8 key/value heads instead of 32 → KV cache is only ~268 MB
- Tied embeddings: input and output embedding matrices are shared, saving ~600–800 MB
Combined with mmap (memory mapping — the .gguf file stays mostly on disk; only accessed pages load into physical RAM), a 3.16 GB model uses only ~500 MB of real RAM during normal use (short conversations). It runs smoothly on both 8 GB and 16 GB PCs.
LexiMind limits the model to 2048 tokens of context (the model supports up to 131,072). You will see this message on startup:
llama_context: n_ctx_seq (2048) < n_ctx_train (131072) -- the full capacity of the model will not be utilized
This is normal and intentional — the KV cache scales linearly with context size. At 2048 it uses ~268 MB; at 131072 it would use ~17 GB. For short children's tutoring sessions, 2048 is more than enough and keeps RAM usage low on any PC.
Only keep one model in the folder at a time — the first .gguf found is loaded automatically. Use the LEXIMIND_MODEL environment variable to select a specific model by filename substring:
set LEXIMIND_MODEL=phi # loads the first .gguf containing "phi" in its name-
Python 3.10 or higher
-
Node.js 18+ and npm
-
Windows: Visual Studio 2022 Build Tools (needed to compile
llama-cpp-python)winget install Microsoft.VisualStudio.2022.BuildTools --override "--add Microsoft.VisualStudio.Workload.VCTools --includeRecommended"
# Clone the repo
git clone https://github.com/jezmn/leximind.git
cd leximind
# Backend
python -m venv venv
venv\Scripts\activate # Windows
pip install --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cpu -r requirements.txt
# If llama-cpp-python fails to install, see "Building llama-cpp-python" below.
# Frontend
cd chat
npm install
npm run build
cd ..
# Download the model (see Recommended Model above) and place it in model/
# Run
python main.pyThe app will build the frontend, start the server on http://127.0.0.1:5050, and open your browser.
- Download the latest
LexiMind-vX.X.X-win-x64.exefrom the Releases page. - Create a folder
leximind/anywhere and inside it create amodel/subfolder, then move the.exeintoleximind/:
leximind/ ← create this folder
├── model/ ← create this folder inside
│ └── microsoft_Phi-4-mini-instruct-Q6_K.gguf ← place the model here
└── LexiMind.exe ← move the downloaded .exe here
- Download the model file (see Recommended model above) and place it inside
model/. - Run
LexiMind.exe. - Run
LexiMind.exe.
That's it. No Python, no Node, no dependencies to install.
The prebuilt llama-cpp-python wheels may be compiled with AVX512 instructions, causing an OSError: [WinError -1073741795] on CPUs without AVX512 support (Intel 12th/13th gen hybrid cores, AMD Ryzen, older Intel).
python -c "import llama_cpp; print(llama_cpp.llama_print_system_info().decode())"If you see AVX512 = 1, your build will crash on non-AVX512 CPUs.
-
Download the source
pip download llama-cpp-python --no-deps --no-binary llama-cpp-python -d C:\temp\llama_src cd C:\temp\llama_src tar -xf llama_cpp_python-<VERSION>.tar.gz
-
Edit
vendor\llama.cpp\ggml\CMakeLists.txt- Change
set(GGML_NATIVE_DEFAULT ON)toset(GGML_NATIVE_DEFAULT OFF) - Change
if (GGML_NATIVE OR NOT GGML_NATIVE_DEFAULT)toif (GGML_NATIVE)
- Change
-
Delete precompiled DLLs
Remove-Item "llama_cpp\lib\*" -ErrorAction SilentlyContinue
-
Install from the modified source
pip uninstall llama-cpp-python -y python -m pip install "C:\temp\llama_src\llama_cpp_python-<VERSION>" --no-cache-dir --no-build-isolation
-
Verify
python -c "import llama_cpp; print(llama_cpp.llama_print_system_info().decode())"
Expected:
AVX2 = 1,AVX512absent.
Important: If you update
llama-cpp-pythonto a newer version, repeat these steps from step 1.
LexiMind uses SQLite for persistence — no database server to install, no config, no setup. The file leximind_chat.db is created automatically in the app directory the first time you run it.
chat_sessions — one row per conversation.
| Column | Type | Notes |
|---|---|---|
session_id |
TEXT (UUID) | Primary key |
user_id |
TEXT | Defaults to "anonymous" |
title |
TEXT | Auto-generated from the first message |
created_at |
TIMESTAMP | Set on creation |
last_activity |
TIMESTAMP | Updated on every message |
chat_messages — individual messages within a session.
| Column | Type | Notes |
|---|---|---|
id |
INTEGER | Auto-increment primary key |
session_id |
TEXT | Foreign key → chat_sessions |
user_message |
TEXT | What the user typed |
ai_response |
TEXT | What the model replied |
timestamp |
TIMESTAMP | When it was sent |
- When you start a new chat, a
session_id(UUID) is created and stored inchat_sessions. - Each exchange (user message + AI reply) is inserted into
chat_messages. - The sidebar loads session titles and message counts via a
LEFT JOINquery. - Clicking a past session loads its last 6 message pairs and restores the conversation.
- Deleting a session removes both its row in
chat_sessionsand all related rows inchat_messages.
The database is local to your machine and never touched by anything other than LexiMind. No telemetry, no sync, no cloud.
If you want to generate the portable .exe from source:
# Make sure the frontend is built
cd chat
npm run build
cd ..
# Install PyInstaller
pip install pyinstaller
# Build
pyinstaller LexiMind.specThe output goes to dist/LexiMind.exe. Zip that folder (including model/ and config/) and you have a portable release.
leximind/
├── api/ # FastAPI endpoints
├── chat/ # Svelte frontend
│ ├── src/ # source files
│ └── dist/ # built output (generated)
├── config/ # prompt templates and settings
├── core/ # LLM engine, prompt building
├── database/ # SQLite chat repository
├── model/ # place your .gguf here
├── schemas/ # Pydantic request/response models
├── utils/ # helpers (input handling, progress)
├── main.py # app entry point
├── config.py # paths and defaults
└── LexiMind.spec # PyInstaller build spec
MIT