Skip to content

Feature: Add SenseVoice for local voice input #61

Description

@LauraGPT

Important

Correction: The earlier comparative claims below are withdrawn; they were not established by a matched benchmark. SenseVoiceSmall supports Mandarin Chinese, Cantonese, English, Japanese, and Korean. It can emit language, emotion, and audio-event tags, but speaker diarization requires a separate model or pipeline (for example CAM++) and is not a built-in SenseVoice result. Runtime, timestamps, punctuation, and performance depend on the selected model, interface, hardware, and audio. FunASR and SenseVoice repository source code is MIT; model weights follow each model card. Please evaluate the exact integration on this project's workload.

Hi! VisionClaw is impressive — real-time AI for smart glasses with voice + vision.

For the voice input/ASR component, SenseVoice could improve responsiveness:

Why SenseVoice for smart glasses?

  • 5x faster than Whisper — critical for real-time wearable interaction
  • 234M params — lightweight enough for edge/phone processing
  • Non-autoregressive — constant-time decoding, instant results
  • Emotion detection — understand user's mood from voice
  • Audio event detection — context awareness (music, traffic, crowds)
  • 50+ languages — auto-detection

Edge deployment options

Links

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions