anything2explainer
Give it a topic, get back a narrated, code-drawn explainer video with subtitles and a chapter progress bar.
- Source repo
- Vincentwei1021/anything2explainer
- Stars
- ★ 2.4k
- Last updated
- 23d ago
- License
- NOASSERTION
- Primary language
- TypeScript
- FA score
- 47/100 · Major gaps
At a glance
- How it runs
- Works with
- Portable with changesCodex · Claude Code
- Cost
- Free software; you pay for model usage
- Setup effort
- Medium · a few setup steps
- You'll need
- Typical use
- A technical creator who wants a 3–5 minute black-canvas motion-graphics explainer about an abstract topic such as vector databases or RAG, without shooting or licensing footage.
- Not a fit if
- Teams that need commercial use without author authorization (PolyForm Noncommercial)
- Creators whose only target format is 9:16 vertical video
- Source review
- 47/100 · Major gaps
What does this agent do, and when should you use it?
anything2explainer is a skill for Claude Code and Codex that turns a topic or an article into a 1280×720 @30fps H.264 explainer video with TTS voiceover, word-boundary-aligned subtitles, chapter cards, a top HUD and a bottom chapter progress bar. It is not a CLI: what ships is the whole method an AI coding agent needs to finish the film — a compilable Remotion 4 template, a primitives and lighting library, voiceover/storyboard/render/quantitative-QC tooling, written style and motion specs, a multi-agent division-of-labour protocol, and one reference film as the quality bar. Every frame is drawn in code with Remotion (React + TypeScript); no stock footage and no generative video model. The run walks 9 stages — scaffold, research, narration and timeline, storyboard, overlays and primitives, pilot, parallel build, render, QC and fixes — and stops at four checkpoints: length and language, narration sign-off, voiceover engine, and the first 30 seconds. Chinese and English each have their own pacing model, subtitle budget and default voice.
The skill reads the 9-stage process in SKILL.md and drives agents through it. It scaffolds a Remotion project from template/scripts/new_project.sh; a research agent produces a sourced research doc (every number, year, organisation and English term must trace to a URL); narration goes into script/narration.txt and python3 scripts/tts_build.py generates the voiceover with edge-tts (Chinese zh-CN-YunxiNeural) or kokoro-82m (English am_liam), turning per-word boundaries into a frame-accurate timeline and subtitle table; python3 scripts/render_storyboard.py emits the storyboard; parallel build agents each take 5–7 shots and write pure-function Remotion components under src/shots/G1..Gn; src/config.ts controls lang, bg ('dots' dot-field wave or 'stars' star field with fog) and length. Rendering runs VER=v1 scripts/render.sh, quantitative checks run python3 scripts/frame_metrics.py, QC agents review per chapter, fix agents rework per group, and delivery notes close the run. A manual path renders the first 30 seconds with scripts/preview.sh 30. Every animation is a pure function of the frame number with seeded randomness, and text fitting is computed rather than measured in the DOM, so re-rendering produces identical frames.
- A technical creator who wants a 3–5 minute black-canvas motion-graphics explainer about an abstract topic such as vector databases or RAG, without shooting or licensing footage.
- An education team that needs to turn an article or document into a course short with voiceover and a chapter progress bar, while keeping the research sources and QC reports as an auditable paper trail.
- A developer already using Claude Code or Codex who wants to hand over a topic and get a finished cut in roughly 1–3 hours, steering it at four checkpoints.
- A publisher needing both Chinese and English versions: one storyboard can be re-timed to the English voiceover, sharing all 44 shots.
- A Raspberry Pi 5 or Linux ARM user running mostly-local TTS via kokoro_onnx or piper with a system Chromium build.
- A team that wants a quantitative QC stage: frame_metrics.py scores rendered frames objectively before QC agents judge them against written criteria.
How do you install or deploy this agent?
Clone the repository and symlink it into the Claude Code or Codex skills directory:
git clone https://github.com/Vincentwei1021/anything2explainer.git
ln -s "$PWD/anything2explainer" ~/.claude/skills/anything2explainer # Claude Code
ln -s "$PWD/anything2explainer" ~/.codex/skills/anything2explainer # CodexInstall dependencies (Node ≥18; the template's npm install pulls remotion 4.0.507 / react 19; ffmpeg is required):
brew install ffmpeg
python3 -m venv ~/.venvs/a2e && source ~/.venvs/a2e/bin/activate
pip install 'edge-tts==7.2.8' numpy pillow scipyEnglish narration additionally needs kokoro-82m running locally:
pip install kokoro soundfile && brew install espeak-ngscipy is only used by the QC script frame_metrics.py. The shell scripts are zsh + Python 3, developed and verified on macOS; Linux should work, Windows is untested.
On Linux / Raspberry Pi (ARM64, verified on Python 3.13):
sudo apt install zsh espeak-ng
sudo apt install chromium # or chromium-browser
export REMOTION_BROWSER_EXECUTABLE=/usr/bin/chromiumWhen kokoro is hard to install on ARM, use kokoro_onnx or piper instead:
pip install kokoro-onnx
TTS_ENGINE=kokoro_onnx KOKORO_ONNX_MODEL=…/kokoro-v1.0.onnx KOKORO_ONNX_VOICES=…/voices-v1.0.bin \
KOKORO_ONNX_VOICE=am_michael python3 scripts/tts_build.pypip install piper-tts
TTS_ENGINE=piper PIPER_MODEL=…/en_US-ryan-medium.onnx python3 scripts/tts_build.pyThe cloud edge engine also works on Linux and needs no local model:
TTS_ENGINE=edge VOICE=en-US-AndrewNeural python3 scripts/tts_build.pyHow do you use this agent?
In Claude Code or Codex, say what you want and the skill triggers itself: "Make me an explainer video about vector databases" or "讲一下向量数据库,做成一条讲解视频". It then walks the 9 stages in SKILL.md — scaffold, research, narration and timeline, storyboard, overlays and primitives, pilot, parallel build, render, QC and fixes — pausing at four checkpoints for your sign-off.
You can also drive the template by hand:
template/scripts/new_project.sh ~/work/my-video myslug
cd ~/work/my-video
# 1. research/调研.md 2. script/narration.txt → python3 scripts/tts_build.py
# 3. script/storyboard_src.md → python3 scripts/render_storyboard.py 4. edit src/config.ts
# 5. src/shots/G1..Gn 6. scripts/preview.sh 30 (first 30 seconds)
# 7. VER=v1 scripts/render.sh + python3 scripts/frame_metrics.py 8. QC → fix → v2/v3The visual style has a single switch, bg in src/config.ts:
bg: 'dots' // dot-field wave, default
bg: 'stars' // star field with fog gradientTo use your own voice, put the finished audio at public/assets/<slug>/audio.wav and fill src/common/timeline.ts and subs.ts by hand (format documented at the top of tts_build.py); everything downstream is unchanged.
What are this agent's strengths and limitations?
- Deterministic code output: every animation is a pure function of the frame number with seeded randomness and text fitting is computed, so re-renders are identical and any frame can be fixed by editing one shot file.
- Auditability is built in: every number, year, organisation and English term on screen must trace to a source URL in the film's research doc, and the run leaves QC reports and delivery notes behind.
- Parallel-agent build plus quantitative QC: 8 build agents wrote shots in parallel for the 4′35″ reference cut in about 40 minutes, then two QC rounds rubber-stamped it against written criteria.
- One storyboard yields both cuts — the Chinese and English films share 44 shots, with the English cut simply re-timed to the English voiceover.
- No GPU required: Remotion renders through headless Chromium on CPU, and the default English voice (kokoro-82m) is an 82M-parameter model that runs locally on CPU.
- The toolkit is PolyForm Noncommercial 1.0.0 — free for noncommercial use, but commercial use requires prior authorisation, so it is not a standard permissive licence.
- Narration is frozen once voiced: shot code hard-codes frame numbers, so a rewrite re-times the whole film and late changes are expensive.
- Only Chinese and English are supported, and there is exactly one visual style with two backdrops; changing anything else means editing reference/style-guide.md and src/ui.tsx.
- No 9:16 vertical output — the template and every safe-area rule assume 1280×720 landscape.
- Parallel builds are resource-hungry: several agents bundle Remotion at once, so keep ≥5 GB free, tmux panes cap around 12 before you must dispatch in waves, and the shell scripts are verified on macOS with Windows untested.
How does this agent compare with similar options?
The README names four tool classes it positions itself against. Generative video models (Sora, Veo, Runway) synthesize footage from a prompt; anything2explainer is deterministic code rather than pixels, every number on screen traces to a source URL, and any frame is fixable by editing one shot file. Avatar/presenter tools (HeyGen, Synthesia) produce a digital presenter reading a script; there is no presenter here — motion-graphics diagrams show the mechanism while the narration drives the visuals. Using Remotion or Motion Canvas by hand gives you a programmable video canvas but not the method; this repo ships research → narration → storyboard → parallel build → QC on top of the canvas, plus style specs, a motion vocabulary and a reference film. Manim is Python mathematical animation; this is an agent-driven end-to-end pipeline with TTS-aligned subtitles, chapters and QC, in React/TypeScript rather than Python.
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| anything2explainer This agent | 47 · Major gaps | Agent plugin / skillFree + model costs | ★ 2.4k | 23d ago | TypeScript | Codex · Claude Code |
| Motion Graphics Skills | 39 · Major gaps | Agent plugin / skillFree | ★ 781 | 11d ago | HTML | Codex · Claude Code |
| Video Shotcraft | 0 · Major gaps | Agent plugin / skillFree | ★ 11k | today | TypeScript | Codex · Claude Code |
| Nano Banana Pro Prompt Recommender Skill | 38 · Major gaps | Agent plugin / skillFree | ★ 1.9k | today | TypeScript | Claude Code |
How does FollowAgents rate this agent?
Why each dimension lost points
Evidence: the README defines exactly four user checkpoints (length/language, narration sign-off, voiceover engine, first 30 seconds), so user_confirmation scores 2; the Originality section requires every on-screen number/year/organisation to trace to a source URL and any B-roll to be logged with sha256/source/licence, so source_attribution scores 2. Deductions: no permission inventory or least-privilege statement exists — shell scripts run with the user's full rights and can write anywhere, so least_privilege is only 1; edge-tts sends narration text to a Microsoft cloud endpoint, disclosed only in one FAQ line with no data-flow diagram or retention note, so data_flow_transparency and sensitive_data_handling are 1 each; dependencies are given as bare pip installs with no lockfile, hash pinning or vulnerability scanning, so dependency_security is 1; cloning, rendering and disk writes have no dry-run or blast-radius description, so external_effects is 1; there is no rollback or artefact-cleanup path, so rollback is 1.
Evidence: the output spec (1280×720@30fps H.264, 44px subtitles, chapter progress bar), the length-tier table, chapter-count rules and frame gaps (10 frames intra-paragraph, 30 at paragraph end) are stated consistently, so self_consistency scores 2. Deductions: dependency availability rests on prose only — no lockfile or version matrix — and the README itself admits kokoro fails on ARM/Python 3.13 and that edge-tts breaks across upgrades, so dependency_availability is only 1; on failure messages the repo only points at lessons.md for traps hit, with no error codes, diagnostics or recovery steps shown, so failure_messages is 1.
Evidence: it targets Claude Code and Codex, supports Chinese and English, 2–8 minute lengths, two backdrops and a bring-your-own-TTS path, so audience_and_scenarios and capability_boundaries score 2 each; environment fit is documented for verified macOS plus Linux/Raspberry Pi ARM with Chromium and TTS fallbacks, so environment_fit scores 2. Deduction: trigger precision relies on natural-language examples ('Make me an explainer video about…') with no trigger vocabulary, exclusion conditions or false-trigger guard, so trigger_precision is only 1.
Evidence: repo layout, the 9-stage process, four checkpoints, FAQ, known limits and licensing (PolyForm Noncommercial plus separately licensed OFL fonts) are all clear, so information_architecture, install_notes, examples_and_faq, known_limitations and license score 2 each. Deductions: there is no CHANGELOG, version number or release tag, so versioning_changelog is 0; naming stability is weak — slug, VER=v1 and TTS_ENGINE conventions are scattered with no stable contract, so naming_stability is 1; maintenance responsibility appears only indirectly via the Required Notice and author attribution, with no issue-response or maintenance commitment, so maintenance_responsibility is 1.
Evidence: the deliverable is a finished MP4 plus a full paper trail (research, narration, storyboard, per-shot source, QC reports), so output_usability scores 2; against hand-written Remotion/Manim it ships the whole method plus a reference film, so marginal_value scores 2. Deduction: cost-benefit is only a rough 1–3 hours and ~2–3 GB estimate, with no token budget or parallel-agent retry cost, and parallel builds need ≥5 GB free with tmux capped around 12 panes, so cost_benefit is only 1.
Evidence: the README claims every on-screen number traces to a URL in the research doc, and points to examples/rag/ for the full paper trail and examples/contrast/ for bad/good frame pairs, so claim_traceability and cross_source_corroboration score 1 each. Deduction: all quality conclusions ('two QC rounds', 'reproducible', reference-film quality) are author self-report that a static review cannot check against the rendered frames or QC reports, so fact_inference_separation is only 1 — the document blends design intent with verified fact, e.g. 'Are the renders reproducible? Yes.' is an assertion, not checkable evidence.
- edge-tts sends narration text to a Microsoft cloud endpoint; use local TTS (kokoro/piper) or your own audio if the narration contains non-public material.
- Scripts execute shell and rendering with the user's full rights; no permission inventory or sandbox guidance is provided, so run in an isolated environment or container.
- Dependencies are given as bare pip commands with no lockfile or hash pinning, and the author admits edge-tts breaks across upgrades — pin versions and audit before installing.
- PolyForm Noncommercial forbids unauthorised commercial use, and Remotion carries its own company licensing terms; confirm both before commercial use.
- Quality and reproducibility claims in the README are author self-report that static review cannot verify; check the artefacts under examples/rag/ yourself.
- No CHANGELOG or release tags exist, leaving no dependable version anchor for upgrades or rollback.