Executant supports running local LLMs via llama.cpp with Apple Silicon Metal GPU acceleration. The architecture keeps LLM inference fast and native while the coding agent (opencode/claude) runs sandboxed in Docker.
┌─────────────────────────────────────────────────┐
│ macOS host (Apple Silicon Metal GPU) │
│ │
│ llama-server :8080 Qwen2.5-Coder 7B │
│ llama-server :8081 Qwen2.5-Coder 14B │
│ llama-server :8082 Llama 3.1 8B │
│ ↑ native binaries, Metal-accelerated ~80 t/s │
└──────────────────────┬──────────────────────────┘
│ HTTP via host-gateway
┌──────────────────────▼──────────────────────────┐
│ Docker container (coding agent) │
│ │
│ opencode / claude-code │
│ can only see /workspace mount │
│ no SSH keys, no ~/.config, no secrets │
└─────────────────────────────────────────────────┘
Security model: The agent that executes code and touches your files is sandboxed in Docker — it can only see what you mount into /workspace. The LLM inference server is just matrix multiplication over an HTTP API; it has no file system access and no security concern running natively.
Performance: Docker on macOS has no Metal GPU passthrough (Linux VM layer). Running llama-server natively bypasses this, giving full Apple Silicon Metal throughput (~80 tokens/sec on M-series chips vs ~11 tokens/sec CPU-only in Docker).
brew install llama.cppThis installs llama-server to /opt/homebrew/bin/llama-server. No daemon, no background service, no hidden data directories — just a binary.
npm run models:downloadDownloads Q4_K_M quantized GGUF files to ~/.executant/models/:
| Model | Size | Port |
|---|---|---|
| Qwen2.5-Coder 7B | ~4.7 GB | 8080 |
| Qwen2.5-Coder 14B | ~9 GB | 8081 |
| Llama 3.1 8B | ~4.7 GB | 8082 |
Downloads are idempotent — already-present files are skipped.
npm run models:startStarts all three llama-server processes in the background. Each loads its model into Metal GPU memory and begins accepting requests on its port. Give them ~30 seconds to warm up.
npm run models:status # check which are running
npm run models:stop # stop all serverscurl http://localhost:8080/health # should return {"status":"ok"}
npm run setup # full dependency check# Single step
executant --provider opencode --model llama-qwen7b/qwen2.5-coder-7b workflow.yaml
# Or set env vars for the session
export EXECUTANT_PROVIDER=opencode
export EXECUTANT_MODEL=llama-qwen7b/qwen2.5-coder-7b
executant workflow.yamlopencode.json registers the three llama.cpp providers with URLs like http://localhost:8080/v1. These resolve correctly in both contexts:
- macOS host:
localhostis the loopback → hits native llama-server directly - Docker dev container:
extra_hosts: localhost:host-gatewaymapslocalhostto the Docker host bridge IP → routes to the native llama-server on the macOS host
No configuration changes needed when switching between host and container contexts.
To start model servers automatically on login:
# Create a launchd agent (adjust paths as needed)
cat > ~/Library/LaunchAgents/com.executant.models.plist << 'EOF'
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>com.executant.models</string>
<key>ProgramArguments</key>
<array>
<string>/opt/homebrew/bin/node</string>
<string>/path/to/executant/src/model-server.ts</string>
<string>start</string>
</array>
<key>RunAtLoad</key>
<true/>
</dict>
</plist>
EOF
launchctl load ~/Library/LaunchAgents/com.executant.models.plistOr just run npm run models:start manually before each session.
To free disk space:
npm run models:stop
rm -rf ~/.executant/models # removes ~18 GB of GGUF files
rmdir ~/.executant/pids 2>/dev/null || true
brew uninstall llama.cpp # optional — removes the binaryThe ~/.executant/models directory is the only thing on your host Mac besides the Homebrew binary.
With all three servers running, compare local models against Claude:
npm run eval:compareResults are written to results/*.csv. Use npm run eval:compare:merge to combine into a single CSV.