A real-time virtual keyboard that lets you type using finger gestures detected by your webcam. No physical keyboard needed — just point and pinch!
- Air typing — hover your index finger over a key, then pinch (index + thumb) to press it
- Two-hand support — use both hands simultaneously; each hand tracked independently
- Virtual CLOSE button — pinch the on-screen CLOSE button to exit
- Cooldown ring — visual arc around your fingertip shows when the next key press is ready
- Real-time hand skeleton — see your hand landmarks rendered live on screen
- Auto model download — hand landmark model (~25 MB) is downloaded automatically on first run
- Clean dark UI — semi-transparent keyboard overlay with hover and press feedback
Point index finger at a key → amber highlight
Pinch index + thumb → key flashes green, character typed
- Python 3.9+
- Webcam
opencv-python >= 4.8.0
mediapipe >= 0.10.0
numpy >= 1.24.0
Install with:
pip install -r requirements_air_keyboard.txtpython air_keyboard.pyIf you have multiple cameras:
python air_keyboard.py --camera 1First run: the app automatically downloads
hand_landmarker.task(~25 MB) from Google's MediaPipe model registry.
| Action | Gesture |
|---|---|
| Hover over a key | Point index finger tip at the key |
| Type a character | Pinch index finger + thumb together |
| Space | Hover SPACE key + pinch |
| Backspace | Hover BACK key + pinch |
| Clear all text | Hover CLEAR key + pinch |
| Exit | Pinch the CLOSE button (top-right) or press ESC |
Webcam frame
└── MediaPipe HandLandmarker (Tasks API)
└── 21 landmarks per hand (up to 2 hands)
├── Index fingertip (landmark #8) → hit-test against key rects
├── Thumb tip (landmark #4) → pinch distance check
└── Normalised pinch distance < 0.055 → key press event
- Hand detection: MediaPipe
HandLandmarker(Tasks API,RunningMode.IMAGE) - Pinch detection: normalised Euclidean distance between index tip and thumb tip — works at any distance from camera
- Per-key cooldown: 0.45 s per key, so both hands can press different keys simultaneously
- Rendering: OpenCV with semi-transparent overlays and anti-aliased text
air_keyboard.py # main application
requirements_air_keyboard.txt # pip dependencies
hand_landmarker.task # auto-downloaded on first run (not tracked in git)
| Problem | Fix |
|---|---|
| Fingers not detected | Ensure good lighting; keep hands in the middle of the frame |
| Wrong camera opens | Use --camera 1 (or 2, 3...) to select a different device |
| Model download fails | Download manually from the MediaPipe model page and place as hand_landmarker.task in the project folder |
| App exits immediately | Make sure no other app is using the camera |
This project was built entirely through prompt engineering with Claude (Anthropic's AI assistant) — no prior OpenCV or MediaPipe experience was needed.
The development process was fully conversational:
- Described the idea in plain English → Claude generated the initial implementation
- Pasted errors back into the chat → Claude diagnosed and fixed them
- Requested UI changes iteratively — raising the keyboard, adding two-hand support, adding the CLOSE button — each as a natural follow-up prompt
- Asked for the GitHub setup, README, and even this section — all through conversation
What this demonstrates:
Knowing what to build, how to describe it clearly, and how to iterate on AI output is a skill in itself — one that lets a single person ship projects that would otherwise require deep specialist knowledge.
This is AI-native development: using Claude as a thought partner and coding engine, while the human drives the vision, tests the result, and makes the product decisions.
MIT