From bf77c945339f8e3fa9fb063bcd9dcbd78996d8a0 Mon Sep 17 00:00:00 2001
From: matt
Date: Wed, 12 Aug 2026 22:43:38 +0100
Subject: [PATCH 1/2] Add Grok STT as the default transcription provider
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Grok STT already ships word-level timestamps, diarization, and filler-word
retention. Normalize its words into the existing Scribe-shaped schema so
pack/render/timeline_view stay unchanged. ElevenLabs remains available
via --provider elevenlabs.
Vision composites stay local (timeline_view PNGs) — no API to swap.
---
.env.example | 4 +-
README.md | 21 ++---
SKILL.md | 14 ++--
helpers/transcribe.py | 161 +++++++++++++++++++++++++++++++-----
helpers/transcribe_batch.py | 22 +++--
install.md | 47 ++++++-----
pyproject.toml | 2 +-
7 files changed, 206 insertions(+), 65 deletions(-)
diff --git a/.env.example b/.env.example
index 4c49a949..b85fcd2c 100644
--- a/.env.example
+++ b/.env.example
@@ -1 +1,3 @@
-ELEVENLABS_API_KEY=
+XAI_API_KEY=
+# Optional fallback (helpers/transcribe.py --provider elevenlabs)
+# ELEVENLABS_API_KEY=
diff --git a/README.md b/README.md
index 0f7b2649..9ed4a10d 100644
--- a/README.md
+++ b/README.md
@@ -4,9 +4,9 @@
# video-use
-Introducing **video-use** — edit videos with Claude Code. 100% open source.
+Introducing **video-use** — edit videos with Grok, Claude Code, or any coding agent. 100% open source.
-Drop raw footage in a folder, chat with Claude Code, get `final.mp4` back. Works for any content — talking heads, montages, tutorials, travel, interviews — without presets or menus.
+Drop raw footage in a folder, chat with your agent, get `final.mp4` back. Works for any content — talking heads, montages, tutorials, travel, interviews — without presets or menus.
Try video-use in [Browser Use Cloud](https://cloud.browser-use.com/v4?utm_campaign=video-use-use-in-cloud&utm_source=github).
@@ -22,21 +22,21 @@ Try video-use in [Browser Use Cloud](https://cloud.browser-use.com/v4?utm_campai
## Setup prompt
-Paste into Claude Code, Codex, Hermes, Openclaw, or any agent with shell access:
+Paste into Grok, Claude Code, Codex, Hermes, Openclaw, or any agent with shell access:
```text
Set up https://github.com/browser-use/video-use for me.
-Read install.md first to install this repo, wire up ffmpeg, register the skill with whichever agent you're running under, and set up the ElevenLabs API key — ask me to paste it when you need it. Then read SKILL.md for daily usage, and always read helpers/ because that's where the editing scripts live. After install, don't transcribe anything on your own — just tell me it's ready and wait for me to drop footage into a folder.
+Read install.md first to install this repo, wire up ffmpeg, register the skill with whichever agent you're running under, and set up the xAI API key — ask me to paste it when you need it. Then read SKILL.md for daily usage, and always read helpers/ because that's where the editing scripts live. After install, don't transcribe anything on your own — just tell me it's ready and wait for me to drop footage into a folder.
```
-The agent handles the clone, dependencies, skill registration, and prompts you once for your ElevenLabs API key (grab one at [elevenlabs.io/app/settings/api-keys](https://elevenlabs.io/app/settings/api-keys)).
+The agent handles the clone, dependencies, skill registration, and prompts you once for your xAI API key (grab one at [console.x.ai](https://console.x.ai/team/default/api-keys)).
Then point your agent at a folder of raw takes:
```bash
cd /path/to/your/videos
-claude # or codex, hermes, etc.
+grok # or claude, codex, hermes, etc.
```
For always-on editing from your own VPS or Telegram, run the agent through [Browser Use Box](https://browser-use.com/bux). [Watch the 15-second demo](https://www.tiktok.com/@browser_use/video/7639824093721758989).
@@ -54,7 +54,8 @@ If you'd rather do it by hand:
```bash
# 1. Clone and symlink into your agent's skills directory
git clone https://github.com/browser-use/video-use ~/Developer/video-use
-ln -sfn ~/Developer/video-use ~/.claude/skills/video-use # Claude Code
+ln -sfn ~/Developer/video-use ~/.grok/skills/video-use # Grok
+# ln -sfn ~/Developer/video-use ~/.claude/skills/video-use # Claude Code
# ln -sfn ~/Developer/video-use ~/.codex/skills/video-use # Codex
# 2. Install deps
@@ -63,9 +64,9 @@ uv sync # or: pip install -e .
brew install ffmpeg # required
brew install yt-dlp # optional, for downloading online sources
-# 3. Add your ElevenLabs API key
+# 3. Add your xAI API key
cp .env.example .env
-$EDITOR .env # ELEVENLABS_API_KEY=...
+$EDITOR .env # XAI_API_KEY=...
```
## How it works
@@ -76,7 +77,7 @@ The LLM never watches the video. It **reads** it — through two layers that tog
-**Layer 1 — Audio transcript (always loaded).** One ElevenLabs Scribe call per source gives word-level timestamps, speaker diarization, and audio events (`(laughter)`, `(applause)`, `(sigh)`). All takes pack into a single ~12KB `takes_packed.md` — the LLM's primary reading view.
+**Layer 1 — Audio transcript (always loaded).** One Grok STT call per source gives word-level timestamps, speaker diarization, and filler-word retention (ElevenLabs Scribe is still available via `--provider elevenlabs`). All takes pack into a single ~12KB `takes_packed.md` — the LLM's primary reading view.
```
## C0103 (duration: 43.0s, 8 phrases)
diff --git a/SKILL.md b/SKILL.md
index fa4b776d..451d3146 100644
--- a/SKILL.md
+++ b/SKILL.md
@@ -24,8 +24,8 @@ These are the things where deviation produces silent failures or broken output.
3. **30ms audio fades at every segment boundary** (`afade=t=in:st=0:d=0.03,afade=t=out:st={dur-0.03}:d=0.03`). Otherwise audible pops at every cut.
4. **Overlays use `setpts=PTS-STARTPTS+T/TB`** to shift the overlay's frame 0 to its window start. Otherwise you see the middle of the animation during the overlay window.
5. **Master SRT uses output-timeline offsets**: `output_time = word.start - segment_start + segment_offset`. Otherwise captions misalign after segment concat.
-6. **Never cut inside a word.** Snap every cut edge to a word boundary from the Scribe transcript.
-7. **Pad every cut edge.** Working window: 30–200ms. Scribe timestamps drift 50–100ms — padding absorbs the drift. Tighter for fast-paced, looser for cinematic.
+6. **Never cut inside a word.** Snap every cut edge to a word boundary from the word-level transcript.
+7. **Pad every cut edge.** Working window: 30–200ms. ASR timestamps drift tens of milliseconds — padding absorbs the drift. Tighter for fast-paced, looser for cinematic.
8. **Word-level verbatim ASR only.** Never SRT/phrase mode (loses sub-second gap data). Never normalized fillers (loses editorial signal).
9. **Cache transcripts per source.** Never re-transcribe unless the source file itself changed.
10. **Parallel sub-agents for multiple animations.** Never sequential. Spawn N at once via the `Agent` tool; total wall time ≈ slowest one.
@@ -45,7 +45,7 @@ The skill lives in `video-use/`. User footage lives wherever they put it. All se
├── project.md ← memory; appended every session
├── takes_packed.md ← phrase-level transcripts, the LLM's primary reading view
├── edl.json ← cut decisions
- ├── transcripts/.json ← cached raw Scribe JSON
+ ├── transcripts/.json ← cached raw STT JSON
├── animations/slot_/ ← per-animation source + render + reasoning
├── clips_graded/ ← per-segment extracts with grade + fades
├── master.srt ← output-timeline subtitles
@@ -59,7 +59,7 @@ The skill lives in `video-use/`. User footage lives wherever they put it. All se
First-time install lives in `install.md` (clone, deps, ffmpeg, skill registration, API key). Don't re-run it every session; on cold start just verify:
-- `ELEVENLABS_API_KEY` resolves — either in the environment or in `.env` at the video-use repo root. If missing, ask the user to paste one and write it to `.env` (never to the user's ``).
+- `XAI_API_KEY` resolves — either in the environment or in `.env` at the video-use repo root. If missing, ask the user to paste one and write it to `.env` (never to the user's ``). `ELEVENLABS_API_KEY` is an optional fallback (`--provider elevenlabs`).
- `ffmpeg` + `ffprobe` on PATH.
- Python deps installed (`uv sync` or `pip install -e .` inside the repo).
- Node.js + npm available if the session needs HyperFrames or Remotion slots. HyperFrames currently requires Node.js 22+.
@@ -67,11 +67,11 @@ First-time install lives in `install.md` (clone, deps, ffmpeg, skill registratio
- First-use animation setup happens inside the slot directory, never at the video-use repo root. HyperFrames can be invoked with `npx --yes hyperframes ...`; Remotion can be scaffolded with `npx create-video@latest` or installed as a project-local dependency before using its `remotion render` command.
- This skill vendors `skills/manim-video/`. Read its SKILL.md when building a Manim slot.
-Helpers (`helpers/transcribe.py`, `helpers/render.py`, etc.) live alongside this SKILL.md. Resolve their paths relative to the directory containing this file — the skill is typically symlinked at `~/.claude/skills/video-use/` or `~/.codex/skills/video-use/`.
+Helpers (`helpers/transcribe.py`, `helpers/render.py`, etc.) live alongside this SKILL.md. Resolve their paths relative to the directory containing this file — the skill is typically symlinked at `~/.grok/skills/video-use/`, `~/.claude/skills/video-use/`, or `~/.codex/skills/video-use/`.
## Helpers
-- **`transcribe.py