Skip to content

About

A voice-controlled, Qwen3-4B-powered desktop assistant with draggable widgets

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

Repository files navigation

Jarvis Vision

A voice-controlled desktop assistant with draggable widgets — weather, maps, notes, calendar — all controlled through a terminal-style chat interface. Powered by the Qwen3-4B model. Jarvis speaks, listens through your microphone, and tries to do what you ask.


Status: Unfinished and Honestly Kind of Janky

This is a prototype that I've been tinkering with. It mostly works, but it has its moments:

  • Audio transcription is unreliable — works best with Groq Whisper (free tier at console.groq.com), but there's a fallback to local faster-whisper which is slower and less accurate.
  • Widget windows land in weird spots — you can drag them around fine, but the first time they open the positioning is a bit random.
  • Jarvis sometimes gets it wrong — the Qwen3-4B model handles command parsing pretty well, but complex multi-step requests can confuse it.
  • Refresh the page and the conversation is gone — history lives in Flask session, not a database. Good for a single sitting, gone on reload.
  • No tests. Like, at all. Zero. Use at your own risk.
  • Everything saves to flat JSON files. Notes and events live in data/ as plain JSON. Fine for one person tinkering, not for production.

If something breaks or you've got an idea, open an issue. If you fix something, send a PR. I'd genuinely like to see this thing improve.


Features

Feature What it does
Voice input Microphone -> WebM -> Groq Whisper (or local faster-whisper) -> text
LLM backend Any OpenAI-compatible API — defaults to Qwen3-4B via Ollama/LM Studio
Text-to-speech Edge TTS — Jarvis reads responses out loud
Weather widget Draggable window showing live Open-Meteo forecasts
Maps widget OSM-based, geocodes locations through Nominatim
Calendar widget Full month grid with event dots and clickable days
Notes widget Create and browse sticky notes
Login Simple password gate via environment variables
Theme Dark navy and gold, Playfair Display font — kind of an Edwardian thing

Quick Start

Prerequisites

  • Python 3.10+
  • An LLM server running an OpenAI-compatible API (Ollama, LM Studio, vLLM, etc.)
  • Optional: a Groq API key for better speech transcription (free tier: 2,000 requests/day)

Setup

git clone https://github.com/TheAhmadZeb/jarvis-open-source.git
cd jarvis-open-source

python3 -m venv venv
source venv/bin/activate

pip install -r requirements.txt

cp .env.example .env
# now edit .env with your LLM endpoint, credentials, and model

Environment variables

Variable Purpose Example
JARVIS_USERNAME Login username admin
JARVIS_PASSWORD Login password keep-it-secret
OLLAMA_URL LLM API endpoint http://192.168.1.42:11434/v1/chat/completions
OLLAMA_MODEL Model name qwen/qwen3-4b-thinking-2507
GROQ_API_KEY Transcription (optional) gsk_...
JARVIS_PORT Server port (optional) 8094

Run

python3 launch.py

Then open http://localhost:8094 and log in.


The Model

This project was built around the Qwen3-4B model (specifically qwen/qwen3-4b-thinking-2507). A few reasons why:

  • Small enough to run on a consumer GPU — an RTX 3060 12GB handles it without breaking a sweat.
  • Fast enough for near-real-time chat — response latency stays low.
  • Instruction-following is genuinely good — important when you're asking it to parse "schedule dentist next Tuesday at 3pm" into a calendar event.
  • Open weights — no API costs, no rate limits.

The system prompt keeps Jarvis in character as a Wodehouse-style valet: polite, slightly witty, always ready with a suggestion. The model pulls it off nicely.

You can swap in any other OpenAI-compatible model by changing OLLAMA_URL and OLLAMA_MODEL. Smaller models tend to struggle with the instruction-following, and the really big ones (70B+) introduce noticeable latency — your mileage will vary.


Architecture

Browser (terminal page)
  ├── Chat interface
  ├── Draggable widgets
  └── Microphone -> WebM recording
         │
         ▼ POST /transcribe
    Groq Whisper (or local faster-whisper)
         │
         ▼
    Flask App (app.py)
      ├── Weather widget -> Open-Meteo API
      ├── Maps widget     -> Nominatim / OpenStreetMap
      ├── Notes widget    -> data/notes.json
      └── Calendar widget -> data/events.json
         │
         ▼ POST /v1/chat/completions
    LLM Server (Ollama / LM Studio)
      └── Qwen3-4B running on your GPU

Contributing

Open issues. Send PRs. Fork it, break it, fix it. All welcome.

Things I'd love help with:

  • Making transcription more reliable without Groq
  • Persisting conversation history across sessions
  • More widgets — todo lists, crypto prices, RSS feeds, whatever
  • Writing tests (there are genuinely none)
  • Untangling the frontend JS — it works but it's not pretty
  • A docker-compose setup that just works

MIT © TheAhmadZeb

About

A voice-controlled, Qwen3-4B-powered desktop assistant with draggable widgets

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages