This document tracks all local (WebGPU/WebLLM) and cloud models supported in PortfoliOS, including their quantization settings, repository sources, hardware requirements, and verification history.
These models run entirely in the browser using WebGPU and the WebLLM runtime. They are compiled to MLC format (q4f16_1 quantization).
| Model ID & Name | Size / VRAM | Repo Source (HF / GitHub) | Status / Verification Notes |
|---|---|---|---|
SmolLM2 360M SmolLM2-360M-Instruct-q4f16_1-MLC |
376MB | mlc-ai/SmolLM2-360M-Instruct-q4f16_1-MLC | Active & Verified (Default) - Official prebuilt model from MLC-AI catalog. - Loads WASM directly from GitHub raw CDN. - Highly stable, works out of the box. |
Qwen 2.5 0.5B Qwen2.5-0.5B-Instruct-q4f16_1-MLC |
420MB | mlc-ai/Qwen2.5-0.5B-Instruct-q4f16_1-MLC | Active & Verified - Official prebuilt model from MLC-AI catalog. - Stable performance for low-mid spec devices. |
Llama 3.2 1B Llama-3.2-1B-Instruct-q4f16_1-MLC |
980MB | mlc-ai/Llama-3.2-1B-Instruct-q4f16_1-MLC | Active & Verified - Official prebuilt model from MLC-AI catalog. - Recommended for devices with 4GB+ system VRAM. |
Gemma 3 1B gemma-3-1b-it-q4f16_1-MLC |
600MB | mlc-ai/gemma-3-1b-it-q4f16_1-MLC | Active - Official prebuilt model. - Requires shader-f16 GPU feature support.- May prompt for HuggingFace authentication depending on gating terms. |
Gemma 2 2B gemma-2-2b-it-q4f16_1-MLC |
1.6GB | mlc-ai/gemma-2-2b-it-q4f16_1-MLC | Active - High-quality local reasoning. - Needs 3GB+ VRAM allocated for browser execution. |
Only models that provide useful instruction/chat responses and are available through the maintained WebLLM catalog belong in the application selector.
- No community gating bypasses: Do not route users through an unofficial model fork solely to avoid provider authentication or license agreements. Hide or remove a model that cannot be loaded reliably from its maintained source.
- Minimum capability: Ultra-small models that return control tokens, repetitive fragments, or cannot sustain basic chat should not be exposed as assistants.
- Avoid broken mirrors:
hf-mirror.com(China CDN) must remain disabled because it fails to resolve outside of China or fails to verify LFS assets correctly.- Run tests directly to
huggingface.coor GitHub raw CDNs.
- Root Cause: The downloaded
.wasmfile is a 25-byte text placeholder containing Git LFS pointer metadata instead of the actual compiled binary. - Fix: Check other community repos or pull the compiled WASM binary (which should be ~1-6 MB in size) and ensure the URL resolves to the direct file, not the LFS text pointer.
- Root Cause: Incognito mode limits storage quotas or denies Cache API calls if the browser context restricts site storage.
- Fix: Ensure standard Cache API support is active, or fall back to memory-only compilation (which will download the model weights on each load).