Skip to content
adamlin1009Public

About

Multimodal SwiftUI AI assistant prototype for text, voice, and camera-aware contextual help on macOS.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Genori

Vision: Genori as a Wearable Assistant

This MVP lays the foundation for Genori's goal: a bracelet-style device with cameras, microphones, sensors, lightweight compute, and connectivity. The assistant would:

  • Support continuous, hands-free voice dialogue with natural listening and responses
  • Analyze real-world context via video, detecting scenarios like cooking pasta and proactively suggesting timers, recipes, or reminders
  • Perform tasks such as reservations, scheduling, and blended digital-physical reminders

The macOS demo establishes multimodal input, proactive cues, and modular core components adaptable to wearable hardware for responsive assistance.

Genori is a multimodal AI assistant that pairs a SwiftUI macOS shell with a reusable Swift package (AIAssistantCore). The current prototype is video-first: it watches for useful context, asks the model to reason about one concrete next step, then presents that step for explicit user approval.

What's in This Demo

  • Swift package (AIAssistantCore) containing models, services, utilities, and an observable AssistantViewModel that manages chat, speech, camera capture, scene reasoning, and pending proactive actions
  • macOS SwiftUI prototype (genoriApp) centered on camera feed streaming, text chat, and explicit approval controls for proactive suggestions
  • OpenAI-backed services (via MacPaw/OpenAI) for chat responses, speech transcription, and image understanding, accessed through a concurrency-safe OpenAIService actor
  • Structured scene suggestions: the assistant receives a short visual summary, generates a Planner/Helper reasoning pass, proposes one action, and waits for the user to approve or dismiss it
  • Local timer execution: approved timer suggestions start an in-app countdown with progress, completion messaging, and cancel/clear controls

Architecture at a Glance

  • Sources/AIAssistantCore/Models: Domain models such as Message and SceneSuggestion
  • Sources/AIAssistantCore/Services: Protocol-based integrations, including OpenAIService, OpenAIClient, and SceneSuggestionParser
  • Sources/AIAssistantCore/Utilities: Utilities for image encoding, permission handling, and capture sessions
  • Sources/AIAssistantCore/ViewModels: AssistantViewModel for state management, chat, voice, video analysis, and pending action coordination
  • genoriApp: SwiftUI views (MainView, ChatView, VoiceView, VideoView) rendering the prototype and hosting the shared view model
  • Tests/AIAssistantCoreTests: XCTest suite covering media encoding and structured scene suggestion parsing

Prerequisites

  • macOS 15 (Sequoia) or newer, with Xcode 16 and Swift 6 toolchain
  • OpenAI API key stored in the OPENAI_API_KEY environment variable (export in shell or configure in Xcode scheme)
  • Approved microphone and camera permissions for voice and video features

Setup and Build

# Compile the Swift package
swift build

# Run the unit tests
swift test

# Open the macOS prototype in Xcode for previews or launch
open genoriApp.xcodeproj

Before launching, ensure the environment variable is set:

export OPENAI_API_KEY=sk-...

In Xcode, add OPENAI_API_KEY to the scheme's Run environment, then build and run the genori target. Grant microphone and camera access when prompted.

Current Project Plan

  1. Make the video-first loop feel useful: capture frames, summarize the scene, and surface one actionable suggestion at a time.
  2. Keep the assistant honest: suggestions pause analysis until the user approves, dismisses, or replies in chat.
  3. Execute low-risk approved actions locally, starting with in-app timers and expanding next to notes and reminders.
  4. Expand from the macOS shell toward wearable constraints by keeping core logic reusable and UI-independent.

Scene Action Flow

  1. VideoCaptureUtility provides preview frames to AssistantViewModel.
  2. OpenAIService.analyzeImage compresses the latest frame into a short scene summary.
  3. OpenAIService.reasonAboutScene asks for structured JSON with Planner/Helper reasoning, a user-facing question, and an action payload.
  4. SceneSuggestionParser converts that JSON into SceneSuggestion.
  5. VideoView presents the suggestion and pauses further analysis until the user approves, dismisses, or replies.
  6. Approved timer actions create an ActiveTimer, show countdown progress, and post completion/cancel messages into chat.

Action Payload

The scene reasoning response should include an action object:

{
  "kind": "timer",
  "title": "Pasta timer",
  "duration_seconds": 600,
  "detail": "Track pasta cooking time."
}

Supported kind values are timer, note, reminder, and general. Timer actions require duration_seconds; other action kinds are currently queued in chat and are ready for future local integrations.

Using the Prototype

  • Video-first view: Start the camera to stream preview frames and perform periodic image analysis.
  • Suggested actions: When the assistant sees a useful opportunity, it shows a structured action panel with the Planner/Helper reasoning, the exact question, and approve/dismiss controls.
  • Timers: Approving a timer action starts a visible countdown and posts a completion message when it finishes.
  • Chat: Type follow-up messages; responses can use the latest scene summary as context.
  • Voice prototype: VoiceView remains available in the app sources for push-to-talk transcription experiments, though the current shell prioritizes the camera loop.

Validation

  • swift build checks the reusable package.
  • swift test runs parser, timer, and media utility coverage.
  • App-level SwiftUI type checking can be run with:
    swift build && swiftc -typecheck -parse-as-library -target arm64-apple-macosx15.0 -I .build/arm64-apple-macosx/debug/Modules genoriApp/AIAssistantPrototypeApp.swift genoriApp/Views/ChatView.swift genoriApp/Views/MainView.swift genoriApp/Views/VideoView.swift genoriApp/Views/VoiceView.swift

About

Multimodal SwiftUI AI assistant prototype for text, voice, and camera-aware contextual help on macOS.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages