Skip to content

Repository files navigation

Verso

CI MIT licensed Free, no subscription No account required Works offline Windows 10 and 11

Push-to-talk dictation for Windows. Hold a hotkey, speak, release, and the cleaned-up text is typed into whatever app has focus.

The software is free, and it stays free. No subscription, no account, no sign-in, no trial that expires, no per-word limit, no paid tier holding the good features. Privacy and productivity should not be things you rent.

Hold the hotkey and speak: the Verbar shows the level, a live caption and the elapsed time, then cleans the text up and gets out of the way

Why it exists

Dictation is not a luxury feature. For a lot of people it is not even a convenience, it is how they use a computer at all: because typing hurts, because their hands do not cooperate, because reading their own writing is easier than producing it, or simply because they think faster than they type. Putting a monthly fee in front of that means putting a monthly fee in front of someone's ability to do their work.

So I made this one free. Not free for now, and not free until there are enough users to start charging. It is MIT licensed, which means "free" is a fact about the licence rather than a promise from me that I could quietly change my mind about in a year. Anyone can use it, read it, change it or fork it, and nobody can take that back, including me.

Free is also the only honest way to make the privacy claim. A program that listens to your microphone all day is asking for an unusual amount of trust, and "trust us" is not good enough when the business model needs your audio on somebody's server. The version you can actually verify is the one that runs on your machine and whose source you can read. So the on-device path is the default rather than a downgrade, and PRIVACY.md accounts for every byte that leaves your computer, including the two that are easy to forget: the local engine downloads its model once, and there is a once-a-day check for a new release that sends nothing but the version number and is one switch to turn off. With no API key, the models already on disk and that switch off, Verso makes no network connection at all.

None of that would matter if the thing were bad, and free software earns a certain amount of suspicion on that front. Windows already ships dictation and most people stop using it within a week, because it transcribes what you said instead of what you meant: no punctuation where you paused, every "um" left in, and no idea what to do when you correct yourself halfway through a sentence.

That last part is the whole job, and it is where the work went. Raw speech-to-text gives you um so like send that to the team i mean the whole team by tomorrow. Verso gives you Send that to the whole team by tomorrow.

What it costs

Nothing, and there is no version of it that costs something. It is MIT licensed, so it is free to use, free to read, free to modify and free to fork, and that cannot be revoked.

The one thing worth being precise about: the optional cloud tier talks to Groq using your own API key, not a Verso account. Whatever Groq charges or gives away on their free tier is between you and Groq, and Verso takes no cut and never sees the key. Turn the cloud off, or never add a key at all, and the app is complete and free end to end, because the on-device engines are the default rather than a crippled fallback.

No telemetry funds this. Nothing is monetised in the background. There is no data being resold, because nothing is collected in the first place, which PRIVACY.md accounts for line by line.

The cleanup is the part that matters. Raw speech-to-text gives you um so like send that to the team i mean the whole team by tomorrow. Verso gives you Send that to the whole team by tomorrow.

What it does

Two transcription tiers, chosen automatically. With a Groq API key it uses whisper-large-v3-turbo in the cloud. Without one it runs entirely on device: NVIDIA Parakeet through ONNX Runtime on CPU for English, faster-whisper otherwise. The keyless path is the default, not a degraded fallback.

A consensus engine, not a single guess. In Grade-A mode a second on-device engine transcribes in parallel and an LLM referee resolves disagreements, but only when it is worth it: a confident cloud result ships immediately, and the dual-engine path is reserved for audio that is genuinely poor. The thresholds were tuned from real latency logs rather than picked.

Cleanup that degrades honestly. Filler removal, spoken punctuation and self-corrections ("scratch that, send it to Sarah instead") go through an LLM, which can be Groq, an on-device Qwen3-1.7B, or Ollama. If none is available a deterministic rules pass still handles punctuation and clear fillers, so the feature never silently stops working.

It learns your vocabulary. Fix a word in History and Verso records the correction. Starred terms are fed to the recogniser as a bias prompt, so proper nouns it has never heard start landing correctly.

A live overlay. The Verbar shows level, state and an optional live caption while you speak, then gets out of the way.

What it looks like

The dashboard reports what the engine is actually doing: which tiers are live, how often the two transcriptions agree, and where the latency is going.

The Verso dashboard

Every dictation is kept with what was heard alongside what was changed, so the cleanup is auditable rather than something that happens to your words offscreen. The whole history is searchable.

History, showing the cleaned text above the raw transcript it came from

Starred terms are fed to the recogniser as a bias prompt, and replacement rules rewrite a word it keeps mishearing.

The dictionary

These are rendered from the real widgets by tools/make_gifs.py against a seeded throwaway profile rather than recorded off a desktop, so they contain no real dictations. The mouse pointer is drawn on, because the capture drives a virtual cursor; what the interface does in response to it is not staged.

Install

Download the latest release, or take a build directly from the table below. No admin rights required; it installs per-user.

Two builds are published:

Build Download Installed Use it when
VersoSetup-CPU.exe 98 MB 367 MB Almost always. Runs entirely on CPU
VersoSetup.exe 1.5 GB several GB You have an NVIDIA GPU and want faster-whisper on CUDA

Those two links always resolve to the newest release, so they do not go stale when a version is cut.

The CPU build is the default recommendation. The vendored CUDA and cuDNN libraries add 1.4 GB to the download on their own, against 41 MB of ONNX Runtime and onnx-asr for the entire CPU inference stack, and Parakeet on CPU is fast enough that most people never notice. Both figures are measured from the built installers rather than estimated.

Neither build downloads a speech model at install time. The first run of a local engine fetches its weights from Hugging Face, so the first dictation after a fresh install is slower than every one after it.

The installer is not code-signed yet, so SmartScreen will warn about an unknown publisher on first run.

Known issues

Found in a pre-release audit and not yet fixed. Listed because you should know what you are installing, and because a project that claims no known problems has usually just not looked.

  • Starting a new dictation while the previous one is still processing can lose the older one from History. The text still gets pasted correctly, but the entry may not be recorded and you may see a "transcription failed" toast after a delivery that actually worked. Leave a beat between takes for now.
  • Changing the theme needs a restart. Switching between Ink and Paper re-themes some of the interface immediately and leaves the rest until the next launch, which can leave text hard to read in between.
  • The overlay is not free while idle. It repaints continuously and photographs the area behind itself for its glass effect, which costs a few percent of one CPU core even when you are not dictating.
  • Deleting a history entry does not delete its recovery audio. Nothing on a row identifies which recording belongs to it. That audio is capped at 7 days regardless, and most dictations never leave one behind. PRIVACY.md explains this in full.
  • First launch downloads a speech model with no progress indicator. It is several hundred MB and runs in the background, but the only sign is a single toast when it starts. If you dictate before it finishes, that dictation waits for the download rather than failing, which on a slow connection can mean a long pause with nothing on screen explaining it.
  • The setup wizard cannot be closed or skipped until you reach the last page. There is no minimize and Escape does nothing.
  • Clicking a shortcut field and then clicking away leaves it listening. The recorder keeps a global key hook until you press Esc or record a combination, so the next key you type anywhere is captured. Press Esc to cancel out of it.

Running without an API key

This is the default. On first launch with no key, Verso uses the on-device engine and rules-based cleanup, and no audio and no text ever leaves your machine.

Two connections still happen and it would be dishonest to round them to zero: the local engine downloads its model from Hugging Face the first time it runs, and there is a once-a-day check against the GitHub releases page that sends nothing but the version number. Both are in the table in PRIVACY.md, the update check is a single switch in Settings, Version, and once the model is on disk and that switch is off, Verso makes no network connection at all.

To add the cloud tier, paste a Groq key into Settings, Cloud access. It goes straight into Windows Credential Manager, where DPAPI encrypts it under your account. The field is masked, it is cleared the moment you save, and Settings afterwards reports only that a key exists: not the value, not a masked prefix, not the length.

If you would rather do it from a terminal, the target name is GROQ_API_KEY:

cmdkey /generic:GROQ_API_KEY /user:verso /pass

Keys live in Credential Manager only. They are never written to a config file, never logged, and never included in an error message.

To force local processing even when a key is present, set the backend to local in Settings or set FORCE_LOCAL_STT=1.

Privacy

Everything is stored under %APPDATA%\Verso, unencrypted. Transcripts are kept for 90 days by default and the retention window is configurable, down to "never store anything". Raw recovery audio is capped at 7 days regardless.

There is no telemetry, no analytics and no crash reporting service. The only outbound connections are to Groq when the cloud tier is active, to Hugging Face when downloading a model, and a once-a-day check for a newer release that sends nothing but the version you are on and can be switched off.

PRIVACY.md is the full accounting, including what the history file contains that is not obvious (it records the title of the window you dictated into).

Build from source

Requires Python 3.11 on Windows 10 or 11, x64.

py -3.11 -m venv .venv
.venv\Scripts\python.exe -m pip install -r requirements-dev.txt
.venv\Scripts\python.exe -m pip install -e . --no-deps
.venv\Scripts\python.exe -m verso

The editable install is what puts src/ on the path; without it python -m verso cannot find the package.

To produce the installers, with Inno Setup 6 installed:

pwsh -File tools\deploy.ps1

Set VERSO_BUILD_CPU_ONLY=1 first for the CPU flavour. Packaging changes must be verified by exercising the feature in the installed executable; a frozen build fails in ways the dev virtualenv cannot reproduce, which is why several of the comments in verso.spec exist.

Tests

.venv\Scripts\python.exe -m pytest tests -q

The automated suite covers the parts where a regression is expensive: history retention and the delete paths, backend routing, the local server's auth, credential handling, and the deterministic text formatting. Tests run against a throwaway VERSO_DATA_DIR and never touch real data.

tests/m*/ holds manual milestone harnesses that drive real audio, real hotkeys and a real window. They are excluded from collection on purpose.

Architecture

hotkey  ->  capture  ->  STT router  ->  cleanup  ->  inject
            (sounddevice)  |             |            (clipboard + paste,
                           |             |             with restore)
                    Groq / Parakeet      rules -> LLM
                    + consensus          (Groq / Qwen3 / Ollama)
                      referee
  • src/verso/pipeline.py is the spine, from key release to pasted text.
  • src/verso/stt/ holds the engines and the router that picks between them.
  • src/verso/cleanup/ is the text layer: deterministic rules first, LLM second.
  • src/verso/ui/ is the Qt application. Tiles are custom-painted, so they ignore stylesheets entirely; styling goes through the painter.
  • src/verso/keys.py is the only module that reads or writes the credential store. Two callers ask it for a key (stt/groq_client.py to build an auth header, llm.py to test whether one exists) and neither persists or logs the value.

Around 16,000 lines across 70 modules. Many carry a dated comment explaining the field failure that produced them, which is usually more useful than the code itself.

Design

DESIGN.md is the committed design contract: the material laws, the page grammar every surface obeys, the instrument-borrowed motion rules, and an explicit list of banned patterns. It exists so that taste is not re-derived from scratch each time, and so a change can be argued against a written standard instead of a preference.

License

MIT. See LICENSE.

Verso links against Qt for Python under the LGPL v3 and python-soxr under the LGPL v2.1. THIRD-PARTY-NOTICES.md lists every component, its license, and how the LGPL obligations are met.

About

Free push-to-talk dictation for Windows. Runs on device with no account and no subscription. Your voice never has to leave your machine.

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages