Private meeting transcription for Mac. No bots. No cloud. No subscriptions.
EchoNotes is a native macOS desktop app that records audio from any call — Zoom, Google Meet, FaceTime, Discord, phone calls, whatever — and transcribes it locally using WhisperKit. Optionally generate AI-powered meeting summaries with your existing ChatGPT subscription or any API key.
🎤 AUDIO-ONLY APP — We capture system audio output (other people on calls) and microphone input (your voice). No video, screen recording, or visual data whatsoever — and no Screen Recording permission either: system audio comes from a Core Audio process tap.
- Universal recording — works with any app that plays audio through your system
- Live transcription — see text appear in ~5-second intervals while you record
- Post-recording transcription — transcribe the full audio after you stop for higher accuracy
- Speaker diarization — distinguish between different speakers in the transcript
- Recording library — browse, search, and manage all your recordings with full-text search
- AI summaries — generate structured meeting summaries (key points, action items, decisions)
- Multi-provider AI — OpenAI, Anthropic, Google Gemini, or local Ollama
- ChatGPT OAuth — sign in with your existing ChatGPT Plus/Pro subscription (no API key needed)
- Global hotkey — start/stop recording with ⌘⇧R from anywhere
- 100% local transcription — all speech-to-text happens on-device via Apple's Neural Engine
- Free — no accounts or usage limits for recording and transcription
Core Audio process tap (system audio) ──→ system.caf ──┐
├──→ session folder + meta.json
AVAudioEngine (microphone) ──→ mic.caf ──┘
│
WhisperKit (CoreML, per track)
│
merged Transcript ("You"/"Them")
(transcript.txt + transcript.json)
│
AI Provider (optional)
│
Meeting Summary
(key points, actions, decisions)
- macOS 15.0+ (Sequoia or later — system audio capture uses Core Audio process taps)
- Apple Silicon (M1 or later recommended for fast transcription)
git clone <repo>
cd echonotes./scripts/build-app.shThis creates EchoNotes.app in the project root. Double-click to launch, or drag to /Applications to install.
open Package.swiftPress ⌘R to build and run.
swift build
.build/debug/EchoNotesOn first launch, the app downloads a Whisper model (~150MB for Base English). This only happens once.
macOS will ask for:
- Microphone — to capture your voice
- System Audio Recording Only — to capture system audio output (prompted on first recording)
No Screen Recording permission is needed: system audio is captured through a Core Audio process tap, which has its own audio-only permission under System Settings → Privacy & Security → Screen & System Audio Recording.
- Launch EchoNotes — a window opens with a sidebar and detail area
- Choose transcription mode in the toolbar: Live or After Recording
- Start a call in any app
- Click Record in the toolbar (or press ⌘⇧R)
- Click Stop when done
- Your recording (mic + system tracks) + transcript are saved to a session folder in
~/Documents/EchoNotes/ - Select a recording in the sidebar to view the transcript
- Click Ask EchoNotes to generate an AI summary
- Open Settings (⌘,) to configure AI provider, Whisper model, and transcription defaults
EchoNotes can generate structured meeting summaries from your transcripts. Configure in Settings (⌘,) → AI tab:
Click Sign in with ChatGPT to use your existing Plus/Pro subscription. Uses OAuth 2.1 + PKCE to authenticate, then calls the ChatGPT backend Responses API directly.
Paste an API key for any supported provider:
- OpenAI —
api.openai.com - Anthropic — Claude models
- Google Gemini — Gemini models
- Ollama — local models (no key needed)
EchoNotes/
├── AI/
│ ├── AIProvider.swift # Multi-provider config (OpenAI, Anthropic, Google, Ollama)
│ └── AIService.swift # Summarization logic + ChatGPT backend streaming
├── App/
│ ├── EchoNotesApp.swift # Entry point (WindowGroup + Settings scene)
│ ├── AppDelegate.swift # Owns shared state (RecordingEngine, RecordingLibrary)
│ └── RecordingEngine.swift # Coordinates capture + writing + transcription
├── Audio/
│ ├── SystemAudioCapture.swift # ScreenCaptureKit (audio-only)
│ ├── MicrophoneCapture.swift # AVAudioEngine mic input
│ └── AudioFileWriter.swift # Stereo M4A output
├── Auth/
│ ├── OAuthManager.swift # OAuth 2.1 + PKCE flow, JWT parsing, token storage
│ └── OAuthCallbackServer.swift # Local HTTP server for OAuth redirect
├── Models/
│ ├── Recording.swift # Recording model + library management
│ ├── Speaker.swift # Speaker enum for diarization
│ └── Transcript.swift # Segment model + export (.txt, .json)
├── Transcription/
│ ├── WhisperEngine.swift # WhisperKit wrapper
│ ├── StreamingTranscriber.swift # Live transcription (5s chunks)
│ ├── TranscriptionManager.swift # Orchestrates transcription + AI config + OAuth state
│ └── ModelManager.swift # Model loading + caching
├── Views/
│ ├── MainWindowView.swift # Root NavigationSplitView (sidebar + detail)
│ ├── SidebarView.swift # Sidebar meeting list with search
│ ├── ActiveRecordingView.swift # Recording session UI (timer, levels, live transcript)
│ ├── RecordingDetailView.swift # Transcript + AI summary with DisclosureGroups
│ ├── DesktopSettingsView.swift # Tabbed Settings window (General, AI, About)
│ ├── SettingsView.swift # AI provider config + OAuth sign-in
│ ├── LibraryView.swift # Recording library (date formatter shared)
│ ├── RecordingControlsView.swift # Start/stop button
│ ├── SummaryView.swift # Meeting summary display
│ ├── TranscriptDisplayView.swift # Completed transcript display
│ └── LevelMeterView.swift # Audio level visualization
├── Utils/
│ ├── Constants.swift # Audio configuration
│ └── Permissions.swift # Mic + screen recording checks
├── scripts/
│ └── build-app.sh # Build release app bundle
├── research/
│ └── competitive-analysis.md # Market research
├── CLAUDE.md # AI coding assistant rules
├── CONTRIBUTING.md # Build + test requirements for PRs
└── EchoNotesTests/ # ~90 tests across 10 files
- Swift / SwiftUI — native macOS desktop app with NavigationSplitView
- SwiftPM — package management (no Xcode project file needed)
- ScreenCaptureKit — system audio capture (audio-only mode)
- AVAudioEngine — microphone capture
- WhisperKit — on-device speech-to-text via CoreML + Apple Neural Engine
- Network.framework — OAuth callback server (NWListener)
- AAC/M4A — compressed audio output
See CONTRIBUTING.md. Every PR must:
- Pass
./scripts/build-app.sh - Include tests for new functionality
- Use feature branches (never push directly to main)
Proprietary — All rights reserved.