From bf77c945339f8e3fa9fb063bcd9dcbd78996d8a0 Mon Sep 17 00:00:00 2001 From: matt Date: Wed, 12 Aug 2026 22:43:38 +0100 Subject: [PATCH 1/2] Add Grok STT as the default transcription provider MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Grok STT already ships word-level timestamps, diarization, and filler-word retention. Normalize its words into the existing Scribe-shaped schema so pack/render/timeline_view stay unchanged. ElevenLabs remains available via --provider elevenlabs. Vision composites stay local (timeline_view PNGs) — no API to swap. --- .env.example | 4 +- README.md | 21 ++--- SKILL.md | 14 ++-- helpers/transcribe.py | 161 +++++++++++++++++++++++++++++++----- helpers/transcribe_batch.py | 22 +++-- install.md | 47 ++++++----- pyproject.toml | 2 +- 7 files changed, 206 insertions(+), 65 deletions(-) diff --git a/.env.example b/.env.example index 4c49a949..b85fcd2c 100644 --- a/.env.example +++ b/.env.example @@ -1 +1,3 @@ -ELEVENLABS_API_KEY= +XAI_API_KEY= +# Optional fallback (helpers/transcribe.py --provider elevenlabs) +# ELEVENLABS_API_KEY= diff --git a/README.md b/README.md index 0f7b2649..9ed4a10d 100644 --- a/README.md +++ b/README.md @@ -4,9 +4,9 @@ # video-use -Introducing **video-use** — edit videos with Claude Code. 100% open source. +Introducing **video-use** — edit videos with Grok, Claude Code, or any coding agent. 100% open source. -Drop raw footage in a folder, chat with Claude Code, get `final.mp4` back. Works for any content — talking heads, montages, tutorials, travel, interviews — without presets or menus. +Drop raw footage in a folder, chat with your agent, get `final.mp4` back. Works for any content — talking heads, montages, tutorials, travel, interviews — without presets or menus. Try video-use in [Browser Use Cloud](https://cloud.browser-use.com/v4?utm_campaign=video-use-use-in-cloud&utm_source=github). @@ -22,21 +22,21 @@ Try video-use in [Browser Use Cloud](https://cloud.browser-use.com/v4?utm_campai ## Setup prompt -Paste into Claude Code, Codex, Hermes, Openclaw, or any agent with shell access: +Paste into Grok, Claude Code, Codex, Hermes, Openclaw, or any agent with shell access: ```text Set up https://github.com/browser-use/video-use for me. -Read install.md first to install this repo, wire up ffmpeg, register the skill with whichever agent you're running under, and set up the ElevenLabs API key — ask me to paste it when you need it. Then read SKILL.md for daily usage, and always read helpers/ because that's where the editing scripts live. After install, don't transcribe anything on your own — just tell me it's ready and wait for me to drop footage into a folder. +Read install.md first to install this repo, wire up ffmpeg, register the skill with whichever agent you're running under, and set up the xAI API key — ask me to paste it when you need it. Then read SKILL.md for daily usage, and always read helpers/ because that's where the editing scripts live. After install, don't transcribe anything on your own — just tell me it's ready and wait for me to drop footage into a folder. ``` -The agent handles the clone, dependencies, skill registration, and prompts you once for your ElevenLabs API key (grab one at [elevenlabs.io/app/settings/api-keys](https://elevenlabs.io/app/settings/api-keys)). +The agent handles the clone, dependencies, skill registration, and prompts you once for your xAI API key (grab one at [console.x.ai](https://console.x.ai/team/default/api-keys)). Then point your agent at a folder of raw takes: ```bash cd /path/to/your/videos -claude # or codex, hermes, etc. +grok # or claude, codex, hermes, etc. ``` For always-on editing from your own VPS or Telegram, run the agent through [Browser Use Box](https://browser-use.com/bux). [Watch the 15-second demo](https://www.tiktok.com/@browser_use/video/7639824093721758989). @@ -54,7 +54,8 @@ If you'd rather do it by hand: ```bash # 1. Clone and symlink into your agent's skills directory git clone https://github.com/browser-use/video-use ~/Developer/video-use -ln -sfn ~/Developer/video-use ~/.claude/skills/video-use # Claude Code +ln -sfn ~/Developer/video-use ~/.grok/skills/video-use # Grok +# ln -sfn ~/Developer/video-use ~/.claude/skills/video-use # Claude Code # ln -sfn ~/Developer/video-use ~/.codex/skills/video-use # Codex # 2. Install deps @@ -63,9 +64,9 @@ uv sync # or: pip install -e . brew install ffmpeg # required brew install yt-dlp # optional, for downloading online sources -# 3. Add your ElevenLabs API key +# 3. Add your xAI API key cp .env.example .env -$EDITOR .env # ELEVENLABS_API_KEY=... +$EDITOR .env # XAI_API_KEY=... ``` ## How it works @@ -76,7 +77,7 @@ The LLM never watches the video. It **reads** it — through two layers that tog timeline_view composite — filmstrip + speaker track + waveform + word labels + silence-gap cut candidates

-**Layer 1 — Audio transcript (always loaded).** One ElevenLabs Scribe call per source gives word-level timestamps, speaker diarization, and audio events (`(laughter)`, `(applause)`, `(sigh)`). All takes pack into a single ~12KB `takes_packed.md` — the LLM's primary reading view. +**Layer 1 — Audio transcript (always loaded).** One Grok STT call per source gives word-level timestamps, speaker diarization, and filler-word retention (ElevenLabs Scribe is still available via `--provider elevenlabs`). All takes pack into a single ~12KB `takes_packed.md` — the LLM's primary reading view. ``` ## C0103 (duration: 43.0s, 8 phrases) diff --git a/SKILL.md b/SKILL.md index fa4b776d..451d3146 100644 --- a/SKILL.md +++ b/SKILL.md @@ -24,8 +24,8 @@ These are the things where deviation produces silent failures or broken output. 3. **30ms audio fades at every segment boundary** (`afade=t=in:st=0:d=0.03,afade=t=out:st={dur-0.03}:d=0.03`). Otherwise audible pops at every cut. 4. **Overlays use `setpts=PTS-STARTPTS+T/TB`** to shift the overlay's frame 0 to its window start. Otherwise you see the middle of the animation during the overlay window. 5. **Master SRT uses output-timeline offsets**: `output_time = word.start - segment_start + segment_offset`. Otherwise captions misalign after segment concat. -6. **Never cut inside a word.** Snap every cut edge to a word boundary from the Scribe transcript. -7. **Pad every cut edge.** Working window: 30–200ms. Scribe timestamps drift 50–100ms — padding absorbs the drift. Tighter for fast-paced, looser for cinematic. +6. **Never cut inside a word.** Snap every cut edge to a word boundary from the word-level transcript. +7. **Pad every cut edge.** Working window: 30–200ms. ASR timestamps drift tens of milliseconds — padding absorbs the drift. Tighter for fast-paced, looser for cinematic. 8. **Word-level verbatim ASR only.** Never SRT/phrase mode (loses sub-second gap data). Never normalized fillers (loses editorial signal). 9. **Cache transcripts per source.** Never re-transcribe unless the source file itself changed. 10. **Parallel sub-agents for multiple animations.** Never sequential. Spawn N at once via the `Agent` tool; total wall time ≈ slowest one. @@ -45,7 +45,7 @@ The skill lives in `video-use/`. User footage lives wherever they put it. All se ├── project.md ← memory; appended every session ├── takes_packed.md ← phrase-level transcripts, the LLM's primary reading view ├── edl.json ← cut decisions - ├── transcripts/.json ← cached raw Scribe JSON + ├── transcripts/.json ← cached raw STT JSON ├── animations/slot_/ ← per-animation source + render + reasoning ├── clips_graded/ ← per-segment extracts with grade + fades ├── master.srt ← output-timeline subtitles @@ -59,7 +59,7 @@ The skill lives in `video-use/`. User footage lives wherever they put it. All se First-time install lives in `install.md` (clone, deps, ffmpeg, skill registration, API key). Don't re-run it every session; on cold start just verify: -- `ELEVENLABS_API_KEY` resolves — either in the environment or in `.env` at the video-use repo root. If missing, ask the user to paste one and write it to `.env` (never to the user's ``). +- `XAI_API_KEY` resolves — either in the environment or in `.env` at the video-use repo root. If missing, ask the user to paste one and write it to `.env` (never to the user's ``). `ELEVENLABS_API_KEY` is an optional fallback (`--provider elevenlabs`). - `ffmpeg` + `ffprobe` on PATH. - Python deps installed (`uv sync` or `pip install -e .` inside the repo). - Node.js + npm available if the session needs HyperFrames or Remotion slots. HyperFrames currently requires Node.js 22+. @@ -67,11 +67,11 @@ First-time install lives in `install.md` (clone, deps, ffmpeg, skill registratio - First-use animation setup happens inside the slot directory, never at the video-use repo root. HyperFrames can be invoked with `npx --yes hyperframes ...`; Remotion can be scaffolded with `npx create-video@latest` or installed as a project-local dependency before using its `remotion render` command. - This skill vendors `skills/manim-video/`. Read its SKILL.md when building a Manim slot. -Helpers (`helpers/transcribe.py`, `helpers/render.py`, etc.) live alongside this SKILL.md. Resolve their paths relative to the directory containing this file — the skill is typically symlinked at `~/.claude/skills/video-use/` or `~/.codex/skills/video-use/`. +Helpers (`helpers/transcribe.py`, `helpers/render.py`, etc.) live alongside this SKILL.md. Resolve their paths relative to the directory containing this file — the skill is typically symlinked at `~/.grok/skills/video-use/`, `~/.claude/skills/video-use/`, or `~/.codex/skills/video-use/`. ## Helpers -- **`transcribe.py