Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
@@ -1 +1,4 @@
# Either key is enough. If both are set, Grok STT is used unless you pass
# --provider elevenlabs.
XAI_API_KEY=
ELEVENLABS_API_KEY=
21 changes: 11 additions & 10 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,9 +4,9 @@

# video-use

Introducing **video-use** — edit videos with Claude Code. 100% open source.
Introducing **video-use** — edit videos with Grok, Claude Code, or any coding agent. 100% open source.

Drop raw footage in a folder, chat with Claude Code, get `final.mp4` back. Works for any content — talking heads, montages, tutorials, travel, interviews — without presets or menus.
Drop raw footage in a folder, chat with your agent, get `final.mp4` back. Works for any content — talking heads, montages, tutorials, travel, interviews — without presets or menus.

Try video-use in [Browser Use Cloud](https://cloud.browser-use.com/v4?utm_campaign=video-use-use-in-cloud&utm_source=github).

Expand All @@ -22,21 +22,21 @@ Try video-use in [Browser Use Cloud](https://cloud.browser-use.com/v4?utm_campai

## Setup prompt

Paste into Claude Code, Codex, Hermes, Openclaw, or any agent with shell access:
Paste into Grok, Claude Code, Codex, Hermes, Openclaw, or any agent with shell access:

```text
Set up https://github.com/browser-use/video-use for me.

Read install.md first to install this repo, wire up ffmpeg, register the skill with whichever agent you're running under, and set up the ElevenLabs API key — ask me to paste it when you need it. Then read SKILL.md for daily usage, and always read helpers/ because that's where the editing scripts live. After install, don't transcribe anything on your own — just tell me it's ready and wait for me to drop footage into a folder.
Read install.md first to install this repo, wire up ffmpeg, register the skill with whichever agent you're running under, and set up an xAI or ElevenLabs API key — ask me to paste it when you need it. Then read SKILL.md for daily usage, and always read helpers/ because that's where the editing scripts live. After install, don't transcribe anything on your own — just tell me it's ready and wait for me to drop footage into a folder.
```

The agent handles the clone, dependencies, skill registration, and prompts you once for your ElevenLabs API key (grab one at [elevenlabs.io/app/settings/api-keys](https://elevenlabs.io/app/settings/api-keys)).
The agent handles the clone, dependencies, skill registration, and prompts you once for a key — [xAI](https://console.x.ai/team/default/api-keys) or [ElevenLabs](https://elevenlabs.io/app/settings/api-keys). Either works. Existing `ELEVENLABS_API_KEY` setups keep working.

Then point your agent at a folder of raw takes:

```bash
cd /path/to/your/videos
claude # or codex, hermes, etc.
grok # or claude, codex, hermes, etc.
```

For always-on editing from your own VPS or Telegram, run the agent through [Browser Use Box](https://browser-use.com/bux). [Watch the 15-second demo](https://www.tiktok.com/@browser_use/video/7639824093721758989).
Expand All @@ -54,7 +54,8 @@ If you'd rather do it by hand:
```bash
# 1. Clone and symlink into your agent's skills directory
git clone https://github.com/browser-use/video-use ~/Developer/video-use
ln -sfn ~/Developer/video-use ~/.claude/skills/video-use # Claude Code
ln -sfn ~/Developer/video-use ~/.grok/skills/video-use # Grok
# ln -sfn ~/Developer/video-use ~/.claude/skills/video-use # Claude Code
# ln -sfn ~/Developer/video-use ~/.codex/skills/video-use # Codex

# 2. Install deps
Expand All @@ -63,9 +64,9 @@ uv sync # or: pip install -e .
brew install ffmpeg # required
brew install yt-dlp # optional, for downloading online sources

# 3. Add your ElevenLabs API key
# 3. Add an xAI and/or ElevenLabs API key (either is enough)
cp .env.example .env
$EDITOR .env # ELEVENLABS_API_KEY=...
$EDITOR .env # XAI_API_KEY=... and/or ELEVENLABS_API_KEY=...
```

## How it works
Expand All @@ -76,7 +77,7 @@ The LLM never watches the video. It **reads** it — through two layers that tog
<img src="static/timeline-view.svg" alt="timeline_view composite — filmstrip + speaker track + waveform + word labels + silence-gap cut candidates" width="100%">
</p>

**Layer 1 — Audio transcript (always loaded).** One ElevenLabs Scribe call per source gives word-level timestamps, speaker diarization, and audio events (`(laughter)`, `(applause)`, `(sigh)`). All takes pack into a single ~12KB `takes_packed.md` — the LLM's primary reading view.
**Layer 1 — Audio transcript (always loaded).** One Grok STT or ElevenLabs Scribe call per source gives word-level timestamps, speaker diarization, and filler-word retention. If both keys are set, Grok is used unless you pass `--provider elevenlabs`. All takes pack into a single ~12KB `takes_packed.md` — the LLM's primary reading view.

```
## C0103 (duration: 43.0s, 8 phrases)
Expand Down
14 changes: 7 additions & 7 deletions SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,8 +24,8 @@ These are the things where deviation produces silent failures or broken output.
3. **30ms audio fades at every segment boundary** (`afade=t=in:st=0:d=0.03,afade=t=out:st={dur-0.03}:d=0.03`). Otherwise audible pops at every cut.
4. **Overlays use `setpts=PTS-STARTPTS+T/TB`** to shift the overlay's frame 0 to its window start. Otherwise you see the middle of the animation during the overlay window.
5. **Master SRT uses output-timeline offsets**: `output_time = word.start - segment_start + segment_offset`. Otherwise captions misalign after segment concat.
6. **Never cut inside a word.** Snap every cut edge to a word boundary from the Scribe transcript.
7. **Pad every cut edge.** Working window: 30–200ms. Scribe timestamps drift 50–100ms — padding absorbs the drift. Tighter for fast-paced, looser for cinematic.
6. **Never cut inside a word.** Snap every cut edge to a word boundary from the word-level transcript.
7. **Pad every cut edge.** Working window: 30–200ms. ASR timestamps drift tens of milliseconds — padding absorbs the drift. Tighter for fast-paced, looser for cinematic.
8. **Word-level verbatim ASR only.** Never SRT/phrase mode (loses sub-second gap data). Never normalized fillers (loses editorial signal).
9. **Cache transcripts per source.** Never re-transcribe unless the source file itself changed.
10. **Parallel sub-agents for multiple animations.** Never sequential. Spawn N at once via the `Agent` tool; total wall time ≈ slowest one.
Expand All @@ -45,7 +45,7 @@ The skill lives in `video-use/`. User footage lives wherever they put it. All se
├── project.md ← memory; appended every session
├── takes_packed.md ← phrase-level transcripts, the LLM's primary reading view
├── edl.json ← cut decisions
├── transcripts/<name>.json ← cached raw Scribe JSON
├── transcripts/<name>.json ← cached raw STT JSON
├── animations/slot_<id>/ ← per-animation source + render + reasoning
├── clips_graded/ ← per-segment extracts with grade + fades
├── master.srt ← output-timeline subtitles
Expand All @@ -59,19 +59,19 @@ The skill lives in `video-use/`. User footage lives wherever they put it. All se

First-time install lives in `install.md` (clone, deps, ffmpeg, skill registration, API key). Don't re-run it every session; on cold start just verify:

- `ELEVENLABS_API_KEY` resolves — either in the environment or in `.env` at the video-use repo root. If missing, ask the user to paste one and write it to `.env` (never to the user's `<videos_dir>`).
- `XAI_API_KEY` or `ELEVENLABS_API_KEY` resolves — environment or `.env` at the video-use repo root. Either is enough. If both are set, Grok STT is used unless you pass `--provider elevenlabs`. An existing ElevenLabs-only `.env` keeps working with no flags. If neither is set, ask the user to paste one and write it to `.env` (never to the user's `<videos_dir>`). Do not overwrite the other key if `.env` already has it.
- `ffmpeg` + `ffprobe` on PATH.
- Python deps installed (`uv sync` or `pip install -e .` inside the repo).
- Node.js + npm available if the session needs HyperFrames or Remotion slots. HyperFrames currently requires Node.js 22+.
- `yt-dlp`, HyperFrames, Remotion, Manim installed only on first use.
- First-use animation setup happens inside the slot directory, never at the video-use repo root. HyperFrames can be invoked with `npx --yes hyperframes ...`; Remotion can be scaffolded with `npx create-video@latest` or installed as a project-local dependency before using its `remotion render` command.
- This skill vendors `skills/manim-video/`. Read its SKILL.md when building a Manim slot.

Helpers (`helpers/transcribe.py`, `helpers/render.py`, etc.) live alongside this SKILL.md. Resolve their paths relative to the directory containing this file — the skill is typically symlinked at `~/.claude/skills/video-use/` or `~/.codex/skills/video-use/`.
Helpers (`helpers/transcribe.py`, `helpers/render.py`, etc.) live alongside this SKILL.md. Resolve their paths relative to the directory containing this file — the skill is typically symlinked at `~/.grok/skills/video-use/`, `~/.claude/skills/video-use/`, or `~/.codex/skills/video-use/`.

## Helpers

- **`transcribe.py <video>`** — single-file Scribe call. `--num-speakers N` optional. Cached.
- **`transcribe.py <video>`** — single-file STT call. Auto: Grok if `XAI_API_KEY` is set, else ElevenLabs Scribe. Force with `--provider grok|elevenlabs`. `--num-speakers N` optional (Scribe only). Cached.
- **`transcribe_batch.py <videos_dir>`** — 4-worker parallel transcription. Use for multi-take.
- **`pack_transcripts.py --edit-dir <dir>`** — `transcripts/*.json` → `takes_packed.md` (phrase-level, break on silence ≥ 0.5s).
- **`timeline_view.py <video> <start> <end>`** — filmstrip + waveform PNG. On-demand visual drill-down. **Not a scan tool** — use it at decision points, not constantly.
Expand Down Expand Up @@ -310,7 +310,7 @@ Things that consistently fail regardless of style:
- **Hierarchical pre-computed codec formats** with USABILITY / tone tags / shot layers. Over-engineering. Derive from the transcript at decision time.
- **Hand-tuned moment-scoring functions.** The LLM picks better than any heuristic you'll write.
- **Whisper SRT / phrase-level output.** Loses sub-second gap data. Always word-level verbatim.
- **Running Whisper locally on CPU.** Slow and it normalizes fillers. Use hosted Scribe.
- **Running Whisper locally on CPU.** Slow and it normalizes fillers. Use hosted Grok STT (or ElevenLabs Scribe).
- **Burning subtitles into base before compositing overlays.** Overlays hide them. (Hard Rule 1.)
- **Single-pass filtergraph when you have overlays.** Double re-encodes. Use per-segment extract → concat.
- **Linear animation easing.** Looks robotic. Always cubic.
Expand Down
165 changes: 144 additions & 21 deletions helpers/transcribe.py
Original file line number Diff line number Diff line change
@@ -1,13 +1,20 @@
"""Transcribe a video with ElevenLabs Scribe.
"""Transcribe a video with Grok STT or ElevenLabs Scribe.

Extracts mono 16kHz audio via ffmpeg, uploads to Scribe with verbatim +
diarize + audio events + word-level timestamps, writes the full response
to <edit_dir>/transcripts/<video_stem>.json.
Accepts either XAI_API_KEY or ELEVENLABS_API_KEY. An existing ElevenLabs-only
.env keeps working with no flags. If both keys are set, Grok STT is used
unless you pass --provider elevenlabs.

Extracts mono 16kHz audio via ffmpeg, uploads with diarization + word-level
timestamps + filler-word retention, writes a Scribe-shaped transcript to
<edit_dir>/transcripts/<video_stem>.json so pack/render/timeline_view stay
provider-agnostic.

Cached: if the output file already exists, the upload is skipped.

Usage:
python helpers/transcribe.py <video_path>
python helpers/transcribe.py <video_path> --provider grok
python helpers/transcribe.py <video_path> --provider elevenlabs
python helpers/transcribe.py <video_path> --edit-dir /custom/edit
python helpers/transcribe.py <video_path> --language en
python helpers/transcribe.py <video_path> --num-speakers 2
Expand All @@ -27,22 +34,53 @@
import requests


GROK_STT_URL = "https://api.x.ai/v1/stt"
SCRIBE_URL = "https://api.elevenlabs.io/v1/speech-to-text"
PROVIDERS = ("grok", "elevenlabs")
KEY_FOR_PROVIDER = {
"grok": "XAI_API_KEY",
"elevenlabs": "ELEVENLABS_API_KEY",
}


def load_api_key() -> str:
def _read_key(name: str) -> str:
for candidate in [Path(__file__).resolve().parent.parent / ".env", Path(".env")]:
if candidate.exists():
for line in candidate.read_text().splitlines():
line = line.strip()
if not line or line.startswith("#") or "=" not in line:
continue
k, v = line.split("=", 1)
if k.strip() == "ELEVENLABS_API_KEY":
return v.strip().strip('"').strip("'")
v = os.environ.get("ELEVENLABS_API_KEY", "")
if not candidate.exists():
continue
for line in candidate.read_text().splitlines():
line = line.strip()
if not line or line.startswith("#") or "=" not in line:
continue
k, v = line.split("=", 1)
if k.strip() != name:
continue
val = v.strip().strip('"').strip("'")
if val:
return val
return os.environ.get(name, "").strip()


def resolve_provider(explicit: str | None) -> str:
if explicit:
if explicit not in PROVIDERS:
sys.exit(f"unknown --provider {explicit!r} (want {'|'.join(PROVIDERS)})")
return explicit
if _read_key("XAI_API_KEY"):
return "grok"
if _read_key("ELEVENLABS_API_KEY"):
return "elevenlabs"
sys.exit(
"need XAI_API_KEY or ELEVENLABS_API_KEY in .env or the environment "
"(existing ElevenLabs-only setups keep working; pass --provider to force one)"
)


def load_api_key(provider: str | None = None) -> str:
provider = resolve_provider(provider)

@cubic-dev-ai cubic-dev-ai Bot Aug 12, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: When a caller supplies an ElevenLabs api_key but omits provider, this ambient-key lookup can select Grok and send the wrong credential, causing authentication failures. Preserve the existing direct-call behavior or require and validate the provider alongside the key.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At helpers/transcribe.py, line 75:

<comment>When a caller supplies an ElevenLabs `api_key` but omits `provider`, this ambient-key lookup can select Grok and send the wrong credential, causing authentication failures. Preserve the existing direct-call behavior or require and validate the provider alongside the key.</comment>

<file context>
@@ -27,22 +30,53 @@
+
+
+def load_api_key(provider: str | None = None) -> str:
+    provider = resolve_provider(provider)
+    name = KEY_FOR_PROVIDER[provider]
+    v = _read_key(name)
</file context>
Fix with cubic

@cubic-dev-ai cubic-dev-ai Bot Aug 12, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: Adding a second provider makes the existing filename-only transcript cache return stale cross-provider data. A user who previously transcribed with ElevenLabs and now runs the new grok default will get their old Scribe transcripts back (and vice versa) because the cache checks only that <video_stem>.json exists. Make the cache provider-aware: include the provider in the cached filename or verify the stored provider field on the cache hit, and skip/re-transcribe when it differs.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At helpers/transcribe.py, line 75:

<comment>Adding a second provider makes the existing filename-only transcript cache return stale cross-provider data. A user who previously transcribed with ElevenLabs and now runs the new grok default will get their old Scribe transcripts back (and vice versa) because the cache checks only that `<video_stem>.json` exists. Make the cache provider-aware: include the provider in the cached filename or verify the stored `provider` field on the cache hit, and skip/re-transcribe when it differs.</comment>

<file context>
@@ -27,22 +30,53 @@
+
+
+def load_api_key(provider: str | None = None) -> str:
+    provider = resolve_provider(provider)
+    name = KEY_FOR_PROVIDER[provider]
+    v = _read_key(name)
</file context>
Fix with cubic

name = KEY_FOR_PROVIDER[provider]
v = _read_key(name)
if not v:
sys.exit("ELEVENLABS_API_KEY not found in .env or environment")
sys.exit(f"{name} not found in .env or environment")
return v


Expand All @@ -55,6 +93,75 @@ def extract_audio(video_path: Path, dest: Path) -> None:
subprocess.run(cmd, check=True, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)


def grok_to_scribe(payload: dict) -> dict:
"""Normalize Grok STT words into the Scribe-shaped schema pack/render expect.

Grok returns {text, start, end, speaker?} with no type / spacing / audio_event
entries. We add type=word, map speaker→speaker_id, and synthesize spacing
tokens from inter-word gaps so silence-aware packing still works.
"""
words_out: list[dict] = []
prev_end: float | None = None
for w in payload.get("words") or []:
start = w.get("start")
if start is None:
continue
end = w.get("end", start)
if prev_end is not None and start > prev_end:
words_out.append({
"type": "spacing",
"text": " ",
"start": prev_end,
"end": start,
})
entry: dict = {
"type": "word",
"text": w.get("text") or "",
"start": start,
"end": end,
}
speaker = w.get("speaker")
if speaker is not None:
entry["speaker_id"] = f"speaker_{speaker}"
words_out.append(entry)
prev_end = end
return {
"text": payload.get("text", ""),
"language": payload.get("language"),
"duration": payload.get("duration"),
"provider": "grok",
"words": words_out,
}


def call_grok_stt(
audio_path: Path,
api_key: str,
language: str | None = None,
) -> dict:
# Verbatim + fillers: do not set format=true (that runs ITN).
form: list[tuple[str, str]] = [
("diarize", "true"),
("filler_words", "true"),
]
if language:
form.append(("language", language))

with open(audio_path, "rb") as f:
resp = requests.post(
GROK_STT_URL,
headers={"Authorization": f"Bearer {api_key}"},
data=form,
files={"file": (audio_path.name, f, "audio/wav")},
timeout=1800,
)

if resp.status_code != 200:
raise RuntimeError(f"Grok STT returned {resp.status_code}: {resp.text[:500]}")

return grok_to_scribe(resp.json())


def call_scribe(
audio_path: Path,
api_key: str,
Expand Down Expand Up @@ -84,7 +191,10 @@ def call_scribe(
if resp.status_code != 200:
raise RuntimeError(f"Scribe returned {resp.status_code}: {resp.text[:500]}")

return resp.json()
payload = resp.json()
if isinstance(payload, dict):
payload.setdefault("provider", "elevenlabs")
return payload


def transcribe_one(
Expand All @@ -94,11 +204,13 @@ def transcribe_one(
language: str | None = None,
num_speakers: int | None = None,
verbose: bool = True,
provider: str | None = None,
) -> Path:
"""Transcribe a single video. Returns path to transcript JSON.

Cached: returns existing path immediately if the transcript already exists.
"""
provider = resolve_provider(provider)
transcripts_dir = edit_dir / "transcripts"
transcripts_dir.mkdir(parents=True, exist_ok=True)
out_path = transcripts_dir / f"{video.stem}.json"
Expand All @@ -109,7 +221,7 @@ def transcribe_one(
return out_path

if verbose:
print(f" extracting audio from {video.name}", flush=True)
print(f" extracting audio from {video.name} [{provider}]", flush=True)

t0 = time.time()
with tempfile.TemporaryDirectory() as tmp:
Expand All @@ -118,7 +230,10 @@ def transcribe_one(
size_mb = audio.stat().st_size / (1024 * 1024)
if verbose:
print(f" uploading {video.stem}.wav ({size_mb:.1f} MB)", flush=True)
payload = call_scribe(audio, api_key, language, num_speakers)
if provider == "grok":
payload = call_grok_stt(audio, api_key, language)
else:
payload = call_scribe(audio, api_key, language, num_speakers)

out_path.write_text(json.dumps(payload, indent=2))
dt = time.time() - t0
Expand All @@ -133,14 +248,20 @@ def transcribe_one(


def main() -> None:
ap = argparse.ArgumentParser(description="Transcribe a video with ElevenLabs Scribe")
ap = argparse.ArgumentParser(description="Transcribe a video with Grok STT or ElevenLabs Scribe")
ap.add_argument("video", type=Path, help="Path to video file")
ap.add_argument(
"--edit-dir",
type=Path,
default=None,
help="Edit output directory (default: <video_parent>/edit)",
)
ap.add_argument(
"--provider",
choices=PROVIDERS,
default=None,
help="STT backend. Default: grok if XAI_API_KEY is set, else elevenlabs.",
)
ap.add_argument(
"--language",
type=str,
Expand All @@ -151,7 +272,7 @@ def main() -> None:
"--num-speakers",
type=int,
default=None,
help="Optional number of speakers when known. Improves diarization accuracy.",
help="Optional speaker count (ElevenLabs only). Improves diarization accuracy.",
)
args = ap.parse_args()

Expand All @@ -160,14 +281,16 @@ def main() -> None:
sys.exit(f"video not found: {video}")

edit_dir = (args.edit_dir or (video.parent / "edit")).resolve()
api_key = load_api_key()
provider = resolve_provider(args.provider)
api_key = load_api_key(provider)

transcribe_one(
video=video,
edit_dir=edit_dir,
api_key=api_key,
language=args.language,
num_speakers=args.num_speakers,
provider=provider,
)


Expand Down
Loading