Summary
Hi @JayWebtech and contributors! 👋
I've been working on an extended fork of AutoShorts to optimize it for long gaming/IRL live streams (1 to 4+ hours) and local hardware efficiency (Ollama, low VRAM, low RAM overhead).
I would love to share the architectural enhancements and features implemented in our fork, discuss the chunking strategies, and see which parts could be upstreamed or inspire the core roadmap.
1. 🧠 Long-Stream Chunking Strategy & Boundary Overlap
- The Problem: Passing a 2 to 4-hour transcript directly to an LLM blows past context windows or produces degraded hallucinations and JSON parse errors.
- The Solution:
- Adaptive 10-minute audio/transcript chunking with a 30-second rolling boundary overlap (
chunk_sec = 600.0, overlap_sec = 30.0).
- The overlap prevents losing epic moments that start at 09:45 and finish at 10:15.
- An intelligent deduplication pass (
merge_overlapping_candidates) merges overlapping candidate timestamps into coherent, single clips with unified hooks and descriptions.
2. 🎯 Acoustic Action Peak Detection for Silent Gaming Clutches
-
The Problem: In tactical shooters or competitive games, streamers often go completely silent during clutch moments, intense 1v3 gunfights, or boss battles. Traditional Whisper-only detection misses these because there is little to no speech.
-
The Solution:
- Lightweight acoustic energy analyzer in pure Python (
wave, struct, math) calculating root-mean-square (RMS) energy windows in milliseconds (0 external dependencies, < 1.6s for a 1-hour audio file).
- Cross-references audio peaks (gunshots, explosions, crowd gasps) with transcript density.
- When acoustic intensity is in the top 15% but speech segment count is $\le 1$, the system automatically synthesizes a Gaming Action / Clutch Highlight candidate and feeds this acoustic cue directly into Ollama/LLM prompt context.
3. ⚡ Zero-VRAM Footprint & Multi-Model Engine (Ollama + Cloud)
- Immediate VRAM Unload: Immediately after Ollama returns JSON candidate planning, the model is unloaded via
keep_alive: 0 (POST /api/generate with empty prompt) so the GPU is freed for video playback, gaming, or other workflows.
- Robust JSON Extraction: A multi-stage parser that handles raw markdown blocks, backtick wrapping (
json ... ), and trailing comma corrections across diverse local models (llama3.2, qwen2.5, mistral, gemma2).
- Cloud Fallback & Rate Limiting: Built-in support for OpenRouter, DeepSeek, Anthropic Claude, OpenAI, Gemini, and Groq with automatic retry handling.
4. 🎬 Auto-Editing Stream Summary Engine (Budget-Based Compilation)
- Target Duration Compilation: Ability to generate a single 3, 5, 8, 10, or 15-minute summary video from detected clips without re-transcribing or extra VRAM usage.
- Editorial Vibes: Selectable narrative focuses:
- Balanced: Complete stream storyline (Hook -> Context -> Climax -> Outro).
- Tryhard / Epic Plays: Prioritizes kills, clutches, and high-tension gameplay.
- Funny / Fails: Prioritizes chat laughter, fails, trolling, and banter.
- Natural Clip Preserving: A cumulative budget algorithm that ensures each selected moment remains 100% whole (never abruptly truncated mid-sentence or mid-play).
- Sequential Low-RAM Concatenation: FFmpeg stream concat with micro-fades maintaining < 120 MB RAM usage.
5. 📺 YouTube Automation: Chapters & 1080p Thumbnails
- Auto YouTube Chapters: Calculates cumulative offsets and exports ready-to-paste timestamps:
00:00 Intro: Teaser Clutch 1v3
01:15 Ranked Match Climax
04:30 Funny Chat Banter
Summary
Hi @JayWebtech and contributors! 👋
I've been working on an extended fork of AutoShorts to optimize it for long gaming/IRL live streams (1 to 4+ hours) and local hardware efficiency (Ollama, low VRAM, low RAM overhead).
I would love to share the architectural enhancements and features implemented in our fork, discuss the chunking strategies, and see which parts could be upstreamed or inspire the core roadmap.
1. 🧠 Long-Stream Chunking Strategy & Boundary Overlap
chunk_sec = 600.0,overlap_sec = 30.0).merge_overlapping_candidates) merges overlapping candidate timestamps into coherent, single clips with unified hooks and descriptions.2. 🎯 Acoustic Action Peak Detection for Silent Gaming Clutches
wave,struct,math) calculating root-mean-square (RMS) energy windows in milliseconds (0 external dependencies, < 1.6s for a 1-hour audio file).3. ⚡ Zero-VRAM Footprint & Multi-Model Engine (Ollama + Cloud)
keep_alive: 0(POST /api/generatewith empty prompt) so the GPU is freed for video playback, gaming, or other workflows.json ...), and trailing comma corrections across diverse local models (llama3.2,qwen2.5,mistral,gemma2).4. 🎬 Auto-Editing Stream Summary Engine (Budget-Based Compilation)
5. 📺 YouTube Automation: Chapters & 1080p Thumbnails