anything2explainer

Give it a topic, get back a narrated, code-drawn explainer video with subtitles and a chapter progress bar.

Stars
★ 2.4k
Last updated
23d ago
License
NOASSERTION
Primary language
TypeScript

At a glance

How it runs
Agent plugin / skillCLI
Works with
Portable with changesCodex · Claude Code
Cost
Free software; you pay for model usage
Setup effort
Medium · a few setup steps
You'll need
Node.js 18+ffmpegPython 3zshRemotion 4.0edge-tts 7.2.8numpypillowscipykokoro + soundfile + espeak-ng (English TTS, optional)Chromium / REMOTION_BROWSER_EXECUTABLE (Linux ARM)Shell / CLINetwork accessLocal filesystem
Typical use
A technical creator who wants a 3–5 minute black-canvas motion-graphics explainer about an abstract topic such as vector databases or RAG, without shooting or licensing footage.
Not a fit if
  • Teams that need commercial use without author authorization (PolyForm Noncommercial)
  • Creators whose only target format is 9:16 vertical video

What does this agent do, and when should you use it?

anything2explainer is a skill for Claude Code and Codex that turns a topic or an article into a 1280×720 @30fps H.264 explainer video with TTS voiceover, word-boundary-aligned subtitles, chapter cards, a top HUD and a bottom chapter progress bar. It is not a CLI: what ships is the whole method an AI coding agent needs to finish the film — a compilable Remotion 4 template, a primitives and lighting library, voiceover/storyboard/render/quantitative-QC tooling, written style and motion specs, a multi-agent division-of-labour protocol, and one reference film as the quality bar. Every frame is drawn in code with Remotion (React + TypeScript); no stock footage and no generative video model. The run walks 9 stages — scaffold, research, narration and timeline, storyboard, overlays and primitives, pilot, parallel build, render, QC and fixes — and stops at four checkpoints: length and language, narration sign-off, voiceover engine, and the first 30 seconds. Chinese and English each have their own pacing model, subtitle budget and default voice.

The skill reads the 9-stage process in SKILL.md and drives agents through it. It scaffolds a Remotion project from template/scripts/new_project.sh; a research agent produces a sourced research doc (every number, year, organisation and English term must trace to a URL); narration goes into script/narration.txt and python3 scripts/tts_build.py generates the voiceover with edge-tts (Chinese zh-CN-YunxiNeural) or kokoro-82m (English am_liam), turning per-word boundaries into a frame-accurate timeline and subtitle table; python3 scripts/render_storyboard.py emits the storyboard; parallel build agents each take 5–7 shots and write pure-function Remotion components under src/shots/G1..Gn; src/config.ts controls lang, bg ('dots' dot-field wave or 'stars' star field with fog) and length. Rendering runs VER=v1 scripts/render.sh, quantitative checks run python3 scripts/frame_metrics.py, QC agents review per chapter, fix agents rework per group, and delivery notes close the run. A manual path renders the first 30 seconds with scripts/preview.sh 30. Every animation is a pure function of the frame number with seeded randomness, and text fitting is computed rather than measured in the DOM, so re-rendering produces identical frames.

  1. A technical creator who wants a 3–5 minute black-canvas motion-graphics explainer about an abstract topic such as vector databases or RAG, without shooting or licensing footage.
  2. An education team that needs to turn an article or document into a course short with voiceover and a chapter progress bar, while keeping the research sources and QC reports as an auditable paper trail.
  3. A developer already using Claude Code or Codex who wants to hand over a topic and get a finished cut in roughly 1–3 hours, steering it at four checkpoints.
  4. A publisher needing both Chinese and English versions: one storyboard can be re-timed to the English voiceover, sharing all 44 shots.
  5. A Raspberry Pi 5 or Linux ARM user running mostly-local TTS via kokoro_onnx or piper with a system Chromium build.
  6. A team that wants a quantitative QC stage: frame_metrics.py scores rendered frames objectively before QC agents judge them against written criteria.

How do you install or deploy this agent?

Clone the repository and symlink it into the Claude Code or Codex skills directory:

git clone https://github.com/Vincentwei1021/anything2explainer.git
ln -s "$PWD/anything2explainer" ~/.claude/skills/anything2explainer   # Claude Code
ln -s "$PWD/anything2explainer" ~/.codex/skills/anything2explainer    # Codex

Install dependencies (Node ≥18; the template's npm install pulls remotion 4.0.507 / react 19; ffmpeg is required):

brew install ffmpeg
python3 -m venv ~/.venvs/a2e && source ~/.venvs/a2e/bin/activate
pip install 'edge-tts==7.2.8' numpy pillow scipy

English narration additionally needs kokoro-82m running locally:

pip install kokoro soundfile && brew install espeak-ng

scipy is only used by the QC script frame_metrics.py. The shell scripts are zsh + Python 3, developed and verified on macOS; Linux should work, Windows is untested.

On Linux / Raspberry Pi (ARM64, verified on Python 3.13):

sudo apt install zsh espeak-ng
sudo apt install chromium                       # or chromium-browser
export REMOTION_BROWSER_EXECUTABLE=/usr/bin/chromium

When kokoro is hard to install on ARM, use kokoro_onnx or piper instead:

pip install kokoro-onnx
TTS_ENGINE=kokoro_onnx KOKORO_ONNX_MODEL=…/kokoro-v1.0.onnx KOKORO_ONNX_VOICES=…/voices-v1.0.bin \
  KOKORO_ONNX_VOICE=am_michael python3 scripts/tts_build.py
pip install piper-tts
TTS_ENGINE=piper PIPER_MODEL=…/en_US-ryan-medium.onnx python3 scripts/tts_build.py

The cloud edge engine also works on Linux and needs no local model:

TTS_ENGINE=edge VOICE=en-US-AndrewNeural python3 scripts/tts_build.py

How do you use this agent?

In Claude Code or Codex, say what you want and the skill triggers itself: "Make me an explainer video about vector databases" or "讲一下向量数据库,做成一条讲解视频". It then walks the 9 stages in SKILL.md — scaffold, research, narration and timeline, storyboard, overlays and primitives, pilot, parallel build, render, QC and fixes — pausing at four checkpoints for your sign-off.

You can also drive the template by hand:

template/scripts/new_project.sh ~/work/my-video myslug
cd ~/work/my-video
# 1. research/调研.md          2. script/narration.txt → python3 scripts/tts_build.py
# 3. script/storyboard_src.md → python3 scripts/render_storyboard.py     4. edit src/config.ts
# 5. src/shots/G1..Gn          6. scripts/preview.sh 30   (first 30 seconds)
# 7. VER=v1 scripts/render.sh + python3 scripts/frame_metrics.py         8. QC → fix → v2/v3

The visual style has a single switch, bg in src/config.ts:

bg: 'dots'   // dot-field wave, default
bg: 'stars'  // star field with fog gradient

To use your own voice, put the finished audio at public/assets/<slug>/audio.wav and fill src/common/timeline.ts and subs.ts by hand (format documented at the top of tts_build.py); everything downstream is unchanged.

What are this agent's strengths and limitations?

Pros
  • Deterministic code output: every animation is a pure function of the frame number with seeded randomness and text fitting is computed, so re-renders are identical and any frame can be fixed by editing one shot file.
  • Auditability is built in: every number, year, organisation and English term on screen must trace to a source URL in the film's research doc, and the run leaves QC reports and delivery notes behind.
  • Parallel-agent build plus quantitative QC: 8 build agents wrote shots in parallel for the 4′35″ reference cut in about 40 minutes, then two QC rounds rubber-stamped it against written criteria.
  • One storyboard yields both cuts — the Chinese and English films share 44 shots, with the English cut simply re-timed to the English voiceover.
  • No GPU required: Remotion renders through headless Chromium on CPU, and the default English voice (kokoro-82m) is an 82M-parameter model that runs locally on CPU.
Limitations
  • The toolkit is PolyForm Noncommercial 1.0.0 — free for noncommercial use, but commercial use requires prior authorisation, so it is not a standard permissive licence.
  • Narration is frozen once voiced: shot code hard-codes frame numbers, so a rewrite re-times the whole film and late changes are expensive.
  • Only Chinese and English are supported, and there is exactly one visual style with two backdrops; changing anything else means editing reference/style-guide.md and src/ui.tsx.
  • No 9:16 vertical output — the template and every safe-area rule assume 1280×720 landscape.
  • Parallel builds are resource-hungry: several agents bundle Remotion at once, so keep ≥5 GB free, tmux panes cap around 12 before you must dispatch in waves, and the shell scripts are verified on macOS with Windows untested.

How does this agent compare with similar options?

The README names four tool classes it positions itself against. Generative video models (Sora, Veo, Runway) synthesize footage from a prompt; anything2explainer is deterministic code rather than pixels, every number on screen traces to a source URL, and any frame is fixable by editing one shot file. Avatar/presenter tools (HeyGen, Synthesia) produce a digital presenter reading a script; there is no presenter here — motion-graphics diagrams show the mechanism while the narration drives the visuals. Using Remotion or Motion Canvas by hand gives you a programmable video canvas but not the method; this repo ships research → narration → storyboard → parallel build → QC on top of the canvas, plus style specs, a motion vocabulary and a reference film. Manim is Python mathematical animation; this is an agent-driven end-to-end pipeline with TTS-aligned subtitles, chapters and QC, in React/TypeScript rather than Python.

Key facts side by side with the most closely related agents.

Agent Source review Form / cost Stars Updated Language Full support on
anything2explainer This agent 47 · Major gaps Agent plugin / skillFree + model costs ★ 2.4k 23d ago TypeScript Codex · Claude Code
Motion Graphics Skills 39 · Major gaps Agent plugin / skillFree ★ 781 11d ago HTML Codex · Claude Code
Video Shotcraft 0 · Major gaps Agent plugin / skillFree ★ 11k today TypeScript Codex · Claude Code
Nano Banana Pro Prompt Recommender Skill 38 · Major gaps Agent plugin / skillFree ★ 1.9k today TypeScript Claude Code

How does FollowAgents rate this agent?

FollowAgents source review · FARS-2.1
Major gaps
47/ 100 5-point scale 2.4 / 5
Trust 12/29
Reliability 6/14
Adaptability 10/18
Convention 9/18
Effectiveness 7/13
Verifiability 3/8
Why each dimension lost points
Trust12 / 29 · 2.1/5

Evidence: the README defines exactly four user checkpoints (length/language, narration sign-off, voiceover engine, first 30 seconds), so user_confirmation scores 2; the Originality section requires every on-screen number/year/organisation to trace to a source URL and any B-roll to be logged with sha256/source/licence, so source_attribution scores 2. Deductions: no permission inventory or least-privilege statement exists — shell scripts run with the user's full rights and can write anywhere, so least_privilege is only 1; edge-tts sends narration text to a Microsoft cloud endpoint, disclosed only in one FAQ line with no data-flow diagram or retention note, so data_flow_transparency and sensitive_data_handling are 1 each; dependencies are given as bare pip installs with no lockfile, hash pinning or vulnerability scanning, so dependency_security is 1; cloning, rendering and disk writes have no dry-run or blast-radius description, so external_effects is 1; there is no rollback or artefact-cleanup path, so rollback is 1.

Reliability6 / 14 · 2.1/5

Evidence: the output spec (1280×720@30fps H.264, 44px subtitles, chapter progress bar), the length-tier table, chapter-count rules and frame gaps (10 frames intra-paragraph, 30 at paragraph end) are stated consistently, so self_consistency scores 2. Deductions: dependency availability rests on prose only — no lockfile or version matrix — and the README itself admits kokoro fails on ARM/Python 3.13 and that edge-tts breaks across upgrades, so dependency_availability is only 1; on failure messages the repo only points at lessons.md for traps hit, with no error codes, diagnostics or recovery steps shown, so failure_messages is 1.

Adaptability10 / 18 · 2.8/5

Evidence: it targets Claude Code and Codex, supports Chinese and English, 2–8 minute lengths, two backdrops and a bring-your-own-TTS path, so audience_and_scenarios and capability_boundaries score 2 each; environment fit is documented for verified macOS plus Linux/Raspberry Pi ARM with Chromium and TTS fallbacks, so environment_fit scores 2. Deduction: trigger precision relies on natural-language examples ('Make me an explainer video about…') with no trigger vocabulary, exclusion conditions or false-trigger guard, so trigger_precision is only 1.

Convention9 / 18 · 2.5/5

Evidence: repo layout, the 9-stage process, four checkpoints, FAQ, known limits and licensing (PolyForm Noncommercial plus separately licensed OFL fonts) are all clear, so information_architecture, install_notes, examples_and_faq, known_limitations and license score 2 each. Deductions: there is no CHANGELOG, version number or release tag, so versioning_changelog is 0; naming stability is weak — slug, VER=v1 and TTS_ENGINE conventions are scattered with no stable contract, so naming_stability is 1; maintenance responsibility appears only indirectly via the Required Notice and author attribution, with no issue-response or maintenance commitment, so maintenance_responsibility is 1.

Effectiveness7 / 13 · 2.7/5

Evidence: the deliverable is a finished MP4 plus a full paper trail (research, narration, storyboard, per-shot source, QC reports), so output_usability scores 2; against hand-written Remotion/Manim it ships the whole method plus a reference film, so marginal_value scores 2. Deduction: cost-benefit is only a rough 1–3 hours and ~2–3 GB estimate, with no token budget or parallel-agent retry cost, and parallel builds need ≥5 GB free with tmux capped around 12 panes, so cost_benefit is only 1.

Verifiability3 / 8 · 1.9/5

Evidence: the README claims every on-screen number traces to a URL in the research doc, and points to examples/rag/ for the full paper trail and examples/contrast/ for bad/good frame pairs, so claim_traceability and cross_source_corroboration score 1 each. Deduction: all quality conclusions ('two QC rounds', 'reproducible', reference-film quality) are author self-report that a static review cannot check against the rendered frames or QC reports, so fact_inference_separation is only 1 — the document blends design intent with verified fact, e.g. 'Are the renders reproducible? Yes.' is an assertion, not checkable evidence.

Risks and how to mitigate them
  • edge-tts sends narration text to a Microsoft cloud endpoint; use local TTS (kokoro/piper) or your own audio if the narration contains non-public material.
  • Scripts execute shell and rendering with the user's full rights; no permission inventory or sandbox guidance is provided, so run in an isolated environment or container.
  • Dependencies are given as bare pip commands with no lockfile or hash pinning, and the author admits edge-tts breaks across upgrades — pin versions and audit before installing.
  • PolyForm Noncommercial forbids unauthorised commercial use, and Remotion carries its own company licensing terms; confirm both before commercial use.
  • Quality and reproducibility claims in the README are author self-report that static review cannot verify; check the artefacts under examples/rag/ yourself.
  • No CHANGELOG or release tags exist, leaving no dependable version anchor for upgrades or rollback.
Evidence confidence: Low Reviewed Oct 11, 2026 Reviewed revision 735c79c8724e
Review evidence README.mdLICENSE
See the full review method →

FAQ

What does it cost to produce a video?
The toolkit itself is free for noncommercial use under PolyForm Noncommercial; there is no subscription. You still supply the AI coding agent account and its usage, and commercial use requires prior authorisation from the author. The English default voice is a cloud call to a Microsoft endpoint; the Chinese default runs locally.
Does it need a GPU?
No. Remotion renders through headless Chromium on the CPU. The Chinese default (edge-tts) is a cloud call to a Microsoft endpoint, while the English default (kokoro-82m) is an 82M-parameter model that runs locally on CPU.
Can I use my own audio or a different TTS engine?
Yes. Drop the finished audio at public/assets/<slug>/audio.wav and fill src/common/timeline.ts and subs.ts by hand — the format is documented at the top of tts_build.py. On Linux/ARM you can also pick TTS_ENGINE=kokoro_onnx, piper or edge.
Are the renders reproducible?
Yes. Every animation is a pure function of the frame number with seeded randomness, and text fitting is computed rather than measured in the DOM, so re-rendering produces identical frames.
Can I still edit the script after the voiceover is generated?
You can, but it is costly: shot code hard-codes frame numbers, so changing one word re-times the entire film. Narration sign-off is the cheapest checkpoint to intervene at; fixing style after the full render costs every build group rather than one.
View on GitHub ↗ Install ↓

Compare agents like this one

The same FARS review applied across the shortlist this agent qualifies for.

Related agents