This MVP lays the foundation for Genori's goal: a bracelet-style device with cameras, microphones, sensors, lightweight compute, and connectivity. The assistant would:
- Support continuous, hands-free voice dialogue with natural listening and responses
- Analyze real-world context via video, detecting scenarios like cooking pasta and proactively suggesting timers, recipes, or reminders
- Perform tasks such as reservations, scheduling, and blended digital-physical reminders
The macOS demo establishes multimodal input, proactive cues, and modular core components adaptable to wearable hardware for responsive assistance.
Genori is a multimodal AI assistant that pairs a SwiftUI macOS shell with a reusable Swift package (AIAssistantCore). The current prototype is video-first: it watches for useful context, asks the model to reason about one concrete next step, then presents that step for explicit user approval.
- Swift package (AIAssistantCore) containing models, services, utilities, and an observable
AssistantViewModelthat manages chat, speech, camera capture, scene reasoning, and pending proactive actions - macOS SwiftUI prototype (genoriApp) centered on camera feed streaming, text chat, and explicit approval controls for proactive suggestions
- OpenAI-backed services (via MacPaw/OpenAI) for chat responses, speech transcription, and image understanding, accessed through a concurrency-safe
OpenAIServiceactor - Structured scene suggestions: the assistant receives a short visual summary, generates a Planner/Helper reasoning pass, proposes one action, and waits for the user to approve or dismiss it
- Local timer execution: approved timer suggestions start an in-app countdown with progress, completion messaging, and cancel/clear controls
Sources/AIAssistantCore/Models: Domain models such asMessageandSceneSuggestionSources/AIAssistantCore/Services: Protocol-based integrations, includingOpenAIService,OpenAIClient, andSceneSuggestionParserSources/AIAssistantCore/Utilities: Utilities for image encoding, permission handling, and capture sessionsSources/AIAssistantCore/ViewModels:AssistantViewModelfor state management, chat, voice, video analysis, and pending action coordinationgenoriApp: SwiftUI views (MainView,ChatView,VoiceView,VideoView) rendering the prototype and hosting the shared view modelTests/AIAssistantCoreTests: XCTest suite covering media encoding and structured scene suggestion parsing
- macOS 15 (Sequoia) or newer, with Xcode 16 and Swift 6 toolchain
- OpenAI API key stored in the
OPENAI_API_KEYenvironment variable (export in shell or configure in Xcode scheme) - Approved microphone and camera permissions for voice and video features
# Compile the Swift package
swift build
# Run the unit tests
swift test
# Open the macOS prototype in Xcode for previews or launch
open genoriApp.xcodeprojBefore launching, ensure the environment variable is set:
export OPENAI_API_KEY=sk-...In Xcode, add OPENAI_API_KEY to the scheme's Run environment, then build and run the genori target. Grant microphone and camera access when prompted.
- Make the video-first loop feel useful: capture frames, summarize the scene, and surface one actionable suggestion at a time.
- Keep the assistant honest: suggestions pause analysis until the user approves, dismisses, or replies in chat.
- Execute low-risk approved actions locally, starting with in-app timers and expanding next to notes and reminders.
- Expand from the macOS shell toward wearable constraints by keeping core logic reusable and UI-independent.
VideoCaptureUtilityprovides preview frames toAssistantViewModel.OpenAIService.analyzeImagecompresses the latest frame into a short scene summary.OpenAIService.reasonAboutSceneasks for structured JSON with Planner/Helper reasoning, a user-facing question, and an action payload.SceneSuggestionParserconverts that JSON intoSceneSuggestion.VideoViewpresents the suggestion and pauses further analysis until the user approves, dismisses, or replies.- Approved timer actions create an
ActiveTimer, show countdown progress, and post completion/cancel messages into chat.
The scene reasoning response should include an action object:
{
"kind": "timer",
"title": "Pasta timer",
"duration_seconds": 600,
"detail": "Track pasta cooking time."
}Supported kind values are timer, note, reminder, and general. Timer actions require duration_seconds; other action kinds are currently queued in chat and are ready for future local integrations.
- Video-first view: Start the camera to stream preview frames and perform periodic image analysis.
- Suggested actions: When the assistant sees a useful opportunity, it shows a structured action panel with the Planner/Helper reasoning, the exact question, and approve/dismiss controls.
- Timers: Approving a timer action starts a visible countdown and posts a completion message when it finishes.
- Chat: Type follow-up messages; responses can use the latest scene summary as context.
- Voice prototype:
VoiceViewremains available in the app sources for push-to-talk transcription experiments, though the current shell prioritizes the camera loop.
swift buildchecks the reusable package.swift testruns parser, timer, and media utility coverage.- App-level SwiftUI type checking can be run with:
swift build && swiftc -typecheck -parse-as-library -target arm64-apple-macosx15.0 -I .build/arm64-apple-macosx/debug/Modules genoriApp/AIAssistantPrototypeApp.swift genoriApp/Views/ChatView.swift genoriApp/Views/MainView.swift genoriApp/Views/VideoView.swift genoriApp/Views/VoiceView.swift