Use free AI models directly in Copilot Chat without API keys, billing, or token counting.
⚠️ Read before use. This is not an official API: the extension drives the providers' regular web chats. Your account can be banned or rate-limited — temporarily or permanently — and any provider may work unstably or stop working at any time. See Important notice.
AI Free VSCode connects free web-based models through a Playwright browser session and registers them as a Copilot Chat provider.
- ✅ Integrates free web-based AI models into Copilot Chat
- ✅ Requires no API keys and does not use paid OpenAI endpoints
- ✅ Authenticates via a real browser session
- ✅ Supports streaming, automatic thinking mode, and tool calling
- ✅ Exposes models through a unified provider for VS Code
- ✅ Generates commit messages from your staged diff
- ✅ Inline code suggestions (ghost text) on demand
- ✅ "Fix with AI" Quick Fix on errors/warnings, with a diff preview
- ✅ Add your own OpenAI-compatible endpoints (Ollama, LM Studio, any gateway)
- ✅ One settings page for everything:
AI Free VSCode — Open Settings
Only models from providers you are signed into appear in the picker — no dead entries that fail on use.
| Model | ID | Context |
|---|---|---|
| Qwen3.8 Max | qwen3.8-max |
1M |
| Qwen3.7 Plus | qwen3.7-plus |
1M |
| Qwen3.7 Max | qwen3.7-max |
1M |
| Qwen3.6 Plus | qwen3.6-plus |
1M |
| Qwen3.5 Plus | qwen3.5-plus |
1M |
| Model | ID | Context |
|---|---|---|
| DeepSeek | deepseek-default |
128K |
| DeepSeek Expert | deepseek-expert |
128K |
| Model | ID | Context |
|---|---|---|
| Kimi K2.5 | kimi-k2.5 |
128K |
⚠️ Kimi is kept out of Copilot Chat (chat: false). Short tool sessions work, but once the context grows K2.5 escapes into its own server-side code interpreter: it recreates the project in Kimi's sandbox, "fixes" it there and reports success while nothing changed in your workspace (seen in the agent test: 33 hiddenipythoncalls in one turn). It stays available for commit messages, inline suggestions and Quick Fix, where no tools are involved.
| Model | ID | Context |
|---|---|---|
| GLM-5.3 | glm-5.3 |
128K |
| GLM-5.2 | glm-5.2 |
128K |
| GLM-5.3 Flash | glm-5.3-flash |
128K |
| GLM-4.7 | glm-4.7 |
128K |
Anything that serves POST {baseUrl}/chat/completions can be added as a
provider: Ollama, LM Studio, llama.cpp, vLLM, OpenRouter, a corporate gateway.
Its models show up in the same picker, next to the free web ones.
Run AI Free VSCode — Open Settings, then + Add endpoint:
| Field | Example | Notes |
|---|---|---|
| Name | LM Studio |
Shown in the model picker |
| Base URL | http://localhost:1234/v1 |
A bare origin gets /v1 appended automatically |
| API key | sk-... |
Stored in SecretStorage, never in settings.json |
| Models | (empty) | Empty = fetched from GET {baseUrl}/models |
Test & load models checks the connection and fills the list; leave the list empty to keep it in sync with the endpoint, or pin an exact set by hand (useful on OpenRouter, where the catalogue is hundreds of models long).
Unlike the web backends, these speak the OpenAI tools API, so tool calls and tool results travel structurally — agent mode works as it does with a paid provider, as far as the model behind the endpoint supports it.
Everything except the key is kept in freeAI.custom.providers, so it syncs with
your settings and can be edited as JSON if you prefer:
- The extension registers a single Copilot Chat provider:
free-ai-vscode. - Signing in launches Playwright and stores the authenticated session in
SecretStorage. - Requests are routed through the providers' private API streams (see Supported models).
- Responses are delivered to VS Code as streamed chunks.
- The extension handles tool calling and thinking-mode when available.
chat.qwen.ai sits behind the Aliyun WAF and Alibaba's anti-bot layer, which
reject plain server-side requests. Qwen keeps working through a two-tier flow:
1. Direct request (fast path). The Qwen client calls the API directly from
Node while mirroring the request headers the web app sends (source, version,
x-request-id, timezone, bx-v). That is enough to pass the WAF for most
endpoints (e.g. creating a chat), and responses stream natively.
2. Browser bridge (fallback). When a response is a challenge instead of the
expected stream — an HTML challenge page, or an x5sec / RGV587
(FAIL_SYS_USER_VALIDATE) JSON body — the request is transparently replayed
inside the real browser session, which carries the genuine cookies and browser
network fingerprint the WAF trusts. The SSE response is still streamed back to
VS Code chunk by chunk.
The bridge reuses the persistent profile created at sign-in
(~/.ai-free-vscode/browser-profile), so no extra login is needed, and it
degrades gracefully:
- Headless + stealth first — no visible window (fingerprint is masked so the anti-bot treats the session as a normal browser).
- If the anti-bot detects it, the bridge escalates to a real (headed) Chrome automatically, with its window minimized so it stays out of your way.
- If a one-time captcha is required, that window is brought to the front and a notification appears — solve it once, then resend your request. The verification cookie is stored in the profile, so you are not asked again for a while.
You normally don't need to touch this, but you can pin the browser mode in
Settings → AI Free VSCode → freeAI.qwen.browserMode (auto / headed /
headless; changing it needs a window reload).
This is a best-effort resilience layer. Provider-side anti-bot rules change frequently, so a blocked Qwen request may still occasionally surface — retrying usually clears it.
chat.z.ai signs each completion request inside its own web app and runs an
invisible captcha before it. So Z.ai has no direct client at all: every message
is typed into the real chat page, in a Chrome window kept minimized off-screen,
and the page sends it itself. The answer is copied from the page's own stream
back to VS Code as it arrives; the model and thinking mode are set on the
outgoing request.
- The window uses the profile created at sign-in
(
~/.ai-free-vscode/zai-browser-profile), so no extra login is needed. - Each request is a throwaway chat, deleted afterwards — your chat.z.ai history stays clean.
- If the captcha ever turns interactive, the window is brought to the front with a notification. Complete it there — the request continues by itself.
- Requests are sent one at a time, and each opens a fresh chat page, so a reply starts a couple of seconds later than with the other providers.
- In VS Code, open the Extensions view (
Cmd+Shift+X/Ctrl+Shift+X), search for AI Free VSCode, and click Install. - Or from the command line:
code --install-extension AppsGanin.free-ai-vscode
- Or open it directly on the Marketplace page.
- Download the latest
free-ai-vscode-*.vsixfrom the Releases page. - In VS Code, open Extensions →
···→ Install from VSIX.... - Reload VS Code.
- Open Command Palette (
Cmd+Shift+P/Ctrl+Shift+P). - Run the sign-in command:
AI Free VSCode — Sign In
- Choose a provider from the list (see Supported models).
- Sign in with your browser and wait for success.
- Open Copilot Chat and select a model.
Prefer a UI? Run AI Free VSCode — Open Settings instead: it has sign-in for
every provider, your own OpenAI-compatible endpoints, and all settings on one
page.
To sign out:
AI Free VSCode — Sign Out
To check auth status:
AI Free VSCode — Status
Besides Copilot Chat, the extension adds a few editor integrations powered by the same free models. All of them respect your sign-in state and use thinking off for speed.
A ✨ button in the Source Control view title bar generates a commit message from your staged diff (falls back to the working-tree diff). The message streams straight into the commit input box.
- Model:
freeAI.commit.model(auto= first available) - Prompt:
freeAI.commit.prompt(Conventional Commits, English by default) - Pick a model quickly: command
AI Free VSCode — Change Commit Model
Every model declares which features it is fit for — chat, commit,
suggestions, fix — and each picker lists only the fit ones. The flags are
independent: commit, suggestions and Quick Fix talk to the providers directly,
so a model hidden from chat still works for them.
Only models of providers you are signed into are ever picked. If the chosen one fails before it has produced any text, the request is retried against your other signed-in providers, one model each — so a provider that is rate-limited or has lost its session does not break the feature while another is available.
On-demand code completion at the cursor. It is manual only — it never fires while typing.
- Trigger:
Ctrl+Alt+\(MacCmd+Alt+\), or commandAI Free VSCode — Suggest Code at Cursor, or the built-in Trigger Inline Suggestion - Accept with
Tab, dismiss withEsc - Enable first: set
freeAI.suggestions.enabledtotrue(the hotkey will offer to enable it) - Model:
freeAI.suggestions.model(auto= first available) - Pick a model quickly: command
AI Free VSCode — Change Completions Model
While a suggestion is being generated the status bar shows
AI Free: suggesting…; if nothing usable comes back it briefly says so instead
of failing silently. Generation stops on its own after 30 seconds, 12 lines or
600 characters — whichever comes first — and whatever arrived by then is offered
as the suggestion.
These backends run through a web session (DeepSeek also solves a PoW challenge), so expect noticeably higher latency than native Copilot. If Copilot itself is active, its suggestion usually wins the race and you see that one instead.
When there is a red error or yellow warning, open the lightbulb (Cmd+. /
Ctrl+.) and choose ✨ Fix with AI Free. The model rewrites the affected
lines (indentation is preserved) and shows a diff preview with Apply / Cancel
before changing the file.
- Toggle the action:
freeAI.fix.enabled - Model:
freeAI.fix.model(auto= first available) - Pick a model quickly: command
AI Free VSCode — Change Fix Model
Everything below is also editable on one page: AI Free VSCode — Open Settings
(sign-in status, your own endpoints, and every setting).
| Setting | Default | Description |
|---|---|---|
freeAI.custom.providers |
[] |
Your own OpenAI-compatible endpoints (keys live in SecretStorage) |
freeAI.playwright.timeout |
120000 |
Browser sign-in timeout in milliseconds |
freeAI.qwen.browserMode |
auto |
Qwen anti-bot fallback browser: auto / headed / headless |
freeAI.zai.browserMode |
auto |
Z.ai bridge browser: auto / headed / headless |
freeAI.proxy |
— | Proxy for the extension's browser windows (see below) |
freeAI.debug |
false |
Verbose DEBUG logs in the output channel (see below) |
freeAI.commit.enabled |
true |
Show the ✨ commit message generation button in Source Control |
freeAI.commit.model |
auto |
Model for commit messages (auto = first available) |
freeAI.commit.prompt |
— | Instruction prepended to the diff for commit generation |
freeAI.suggestions.enabled |
false |
Enable manual inline ghost-text suggestions |
freeAI.suggestions.model |
auto |
Model for inline suggestions |
freeAI.suggestions.maxPrefixChars |
2000 |
Chars of code before the cursor sent to the model |
freeAI.suggestions.maxSuffixChars |
800 |
Chars of code after the cursor sent to the model |
freeAI.fix.enabled |
true |
Show the "Fix with AI Free" Quick Fix on diagnostics |
freeAI.fix.model |
auto |
Model for fixing problems |
Thinking (reasoning) is automatic: enabled in plain chat, disabled when tools are active (reasoning is unreliable with tool calling on these backends).
Requests the extension makes itself already follow VS Code's http.proxy (and
http.noProxy). The browser windows — sign-in and the Qwen / Z.ai bridges —
run as a separate Chrome process and take the same proxy too, so setting
http.proxy alone covers everything. freeAI.proxy overrides it for the
browsers only, e.g. for a SOCKS proxy: http://host:port,
http://user:pass@host:port or socks5://host:port. A change applies to
newly opened browser windows; reload VS Code to apply it to a running bridge.
- Turn on
freeAI.debug(Verbose debug logs on the settings page). - Reproduce the problem.
- Run
AI Free VSCode — Copy Debug Report(or the button next to the checkbox) and paste the result into a GitHub issue, together with what you did and what you expected.
The report holds your VS Code / OS / extension versions, the relevant settings,
which providers are signed in, and the recent log: every chat request with its
own [#id], timings, HTTP statuses, the browser that was launched, errors with
their stack, and an excerpt of the model's raw output. Tokens, cookies, keys,
proxy passwords and e-mail addresses are masked automatically, but the log can
contain fragments of your prompts and the model's answers — skim it before
posting.
git clone https://github.com/AppsGanin/ai-free-vscode
cd ai-free-vscode
npm install
# If you need local Chromium for development:
npx playwright install chromiumRun the extension in VS Code with F5.
npm test # offline: parsers, adapter, every provider client (~1 s)
npm run test:live # real requests to Qwen, DeepSeek, Kimi and Z.aiOffline tests (test/) need no network or accounts. Each provider client
is fed canned upstream streams through a stubbed fetch (Z.ai through a fake
browser bridge), split into small pieces the way real streams arrive. They
cover the tool-call dialects, text/reasoning separation, error mapping, the
order of parts the chat UI receives, and argument re-typing against tool
schemas.
Live tests (test/live/) send real requests through the same path a chat
takes (adapter → provider → upstream) and catch what canned data cannot: a
changed stream format, a new anti-bot step, a model that stopped following the
tool protocol. They read each provider's session from its browser profile in
~/.ai-free-vscode/, so:
- a provider that is not signed in opens a browser window on its sign-in
page and the run waits for you (5 min;
LIVE_LOGIN_TIMEOUT=600for longer, never on CI) — or sign in through the extension beforehand; - close the VS Code dev host — Chrome lets one process hold a profile, and a locked profile is skipped with the reason.
test/live/agent.test.ts runs a full agent session per model: a small
project with planted bugs and a six-part task (run tests, fix the code, add a
function and its test, update the README), worked through with Copilot's own
tools for dozens of steps. It passes only if the project really ends up
correct — tests green, the new function right, the original tests untouched —
and every tool call along the way matched its schema. A step-by-step
transcript of each run lands in test/live/.reports/.
npm run test:live -- test/live/agent.test.ts # every provider (~10–30 min)
npm run test:live -- test/live/agent.test.ts -t qwen # one providerA provider that is overloaded on its side ("at capacity", rate limits) marks the affected test skipped with that reason rather than failed, so a red live run always points at something in this repository. The Z.ai bridge tests run in a throwaway guest profile and answer the completion endpoint themselves, so they need no account and cost no model calls.
Each provider can be excluded from a build via a PROVIDER_<NAME>=false
environment variable (accepted falsy values: false, 0, off, no).
By default all providers are included. The build prints the active set, e.g.
[build] enabled providers: deepseek, kimi.
# Build without Qwen
PROVIDER_QWEN=false npm run bundle
# Build with only DeepSeek (drop Qwen and Kimi)
PROVIDER_QWEN=false PROVIDER_KIMI=false npm run bundle
# Package a VSIX without Qwen
PROVIDER_QWEN=false npm run bundle && npm run packageProvider keys: qwen, deepseek, kimi, zai. Adding a new provider means
registering its key in src/providers/providerConfig.ts,
esbuild.js, and the factory map in src/extension.ts.
- VS Code
^1.105.0 - System Chrome or network access for Playwright to download Chromium
This is an unofficial extension. Authentication is handled through a browser session, and none of the providers offers this as an API — the extension automates their regular web chats.
- Account bans. Providers can detect automated use and restrict your account or IP: a temporary block (a captcha, a pause of several minutes or hours, "too many requests") or a permanent ban. Use an account you can afford to lose, and avoid bursts of requests.
- Instability. Because this is not a real API, every provider may work unstably: slow or empty answers, captchas, expired sessions, or sudden breakage when a site changes its web app, anti-bot rules or terms. A fix may take a while, and some models may disappear.
- No guarantees. Availability, speed and quality are decided by the providers, not by this extension.
Use at your own risk. Web session automation may violate provider policies.
AI Free VSCode is built in spare time and distributed for free. If the extension has been useful, you can support its development.
The easiest way is via DonationAlerts or Boosty (one-time or recurring). You can also donate in USDT — pick a network and carefully copy the address (the sender's and recipient's networks must match):
| Network | Token | Address |
|---|---|---|
| TRC20 (Tron) | USDT | TJwyrPVEZVZ1YrcmDiZTyFjLo3Q2DmEGzs |
| ERC20 (Ethereum) | USDT | 0xf9d663146ce902da91911b214c71cc73a5269d1d |
| Solana | USDT | 2qAZRTbaUMTfYuZbD1dCYHjkYgxkw4dUYE9XY3JhC2Cs |
| TON | USDT | UQDoat731MLYuIw8ayL3Vhhw7zTBbLvRaQFmDvab--CNNI7e |
If you also use Bybit, the simplest option is a transfer by UID (instant and fee-free):
| Exchange | UID |
|---|---|
| Bybit | 136462734 |
The cheapest fees are on the TRON (TRC20) and TON networks. Send only USDT and only on the specified network — a transfer on the wrong network is unrecoverable.
Thanks for your support! 🙏






