Material Learning Studio
Turn books, papers, and office documents into traceable, testable, offline learning pages.
- Source repo
- dmoshehun-prog/learn-from-materials
- Stars
- ★ 775
- Last updated
- 2d ago
- License
- MIT
- Primary language
- Python
- FA score
- 92/100 · Excellent
At a glance
- How it runs
- Works with
- Universal · cross-platformCodex · Claude Code
- Cost
- Free, no paid service needed
- Setup effort
- Low · running in minutes
- You'll need
- Typical use
- Graduate researchers studying long papers or technical PDFs who need every major claim, method, and source location to remain traceable.
- Not a fit if
- Users limited to pure chat without filesystem access or Python execution
- Users processing MOBI/AZW without Calibre installed
- Teams expecting automatic installation of parsing or OCR dependencies
- Source review
- 92/100 · Excellent
What does this agent do, and when should you use it?
learn-from-materials is a cross-agent skill built on the Agent Skills specification for books, PDFs, slides, Word files, web pages, and multi-material collections. It combines a SKILL.md workflow with Python extraction and validation scripts, content protocols, HTML templates, and example knowledge bases, and it relies on a host agent that can read files and run local commands. Users choose between Quick Overview, which scans the full structure and reviews displayed citations, and Systematic Study, which adds full coverage, summary mapping, action-rule review, and an independent content pass. Deliverables include a traceable `.learnkb/` knowledge base, a self-contained interactive HTML page, Markdown derived from the same data, bound page data, a delivery manifest, and a ZIP archive. The page includes framework maps, guided content, a glossary, action rules, dynamic self-checks, and local notes, while Apply This Methodology only creates a copyable prompt and does not run AI inside the page. Its scripts operate offline by default, install no dependencies automatically, and keep notes, learner profiles, and wrong-answer records in browser local storage.
The workflow starts with scripts/extract.py, which extracts text and source mappings from formats including PDF, EPUB, DOCX, PPTX/PPTM, HTML, Markdown, TXT, and RTF; MOBI/AZW/AZW3 requires Calibre's ebook-convert. After extraction, the agent reads the source and authors page.json, methods.json, and methodology.json; the scripts do not automatically create this learning content from the source text. scripts/methods.py bind and scripts/methodology.py bind then connect method cards, the integrated whole-material methodology, and page content through a shared data model. A Quick Overview uses scripts/prepare_quick.py to audit citations displayed in the final page, while Systematic Study additionally requires summary-ledger.json, action-rule-ledger.json, content-review.json, and a reviewed original-heading index for PDFs. Finally, scripts/finalize.py validates the selected depth's source and coverage records and produces HTML, Markdown, bound page JSON, the knowledge base, a delivery manifest, and a ZIP. Generated pages can show fine-grained PDF page, slide, or EPUB chapter locations and distinguish source-backed facts, uncovered areas, model-added content, and externally verified items.
- Graduate researchers studying long papers or technical PDFs who need every major claim, method, and source location to remain traceable.
- Educators turning slides, handouts, and supplemental documents into an offline page with a glossary, self-checks, and local notes.
- Analysts producing a quick but source-aware map of an industry report's arguments, limits, and action rules.
- Book learners converting EPUBs or extracted books into browsable knowledge bases and accumulating reusable method cards.
- Teams using Codex, Claude Code, WorkBuddy, or GitHub Copilot CLI that want a structured material-learning workflow inside an existing agent.
How do you install or deploy this agent?
You need Python 3.10 or later and an agent that can read the material and skill directory, run local commands, and recognize SKILL.md. For the documented 1.0 release, extract learn-from-materials-1.0-release.zip and place its complete learn-from-materials/ directory in the client's skill search path. SKILL.md, scripts/, references/, and templates/ must sit directly inside that directory. Back up an existing directory of the same name before replacing it, then reload or restart the client.
You may instead install the repository's default branch, although it can differ from the 1.0 ZIP release. For a user-level Codex installation:
git clone https://github.com/dmoshehun-prog/learn-from-materials.git ~/.agents/skills/learn-from-materialsFor Claude Code:
git clone https://github.com/dmoshehun-prog/learn-from-materials.git ~/.claude/skills/learn-from-materialsFor GitHub Copilot CLI, one documented location is:
git clone https://github.com/dmoshehun-prog/learn-from-materials.git ~/.copilot/skills/learn-from-materialsNo API credential is documented as required. The project does not automatically install optional parsers, OCR tools, or browser dependencies.
How do you use this agent?
After installation and client reload, attach or identify the material and state the desired learning depth. A first Systematic Study request is:
Please use learn-from-materials to systematically study this PDF and generate a traceable knowledge base and a learning page.For a Quick Overview of slides:
Give me a quick overview of this PPT, keep per-slide sources, and generate an offline learning HTML.If you omit the depth, the skill asks once. From the skill root, you can first inspect available parsing capabilities:
python3 scripts/extract.py --checkTo run extraction manually, replace the placeholders with real paths:
python3 scripts/extract.py <material-path> --mode text --ocr auto --output-dir <topic>.learnkbAfter the agent has authored page.json, methods.json, and methodology.json, bind the method data:
python3 scripts/methods.py bind --library <topic>.learnkb/methods.json --page page.json --knowledge-base <topic>.learnkb --output page-with-methods.jsonpython3 scripts/methodology.py bind --model <topic>.learnkb/methodology.json --page page-with-methods.json --knowledge-base <topic>.learnkb --output page-ready.jsonFor a Quick Overview, review displayed citations and then run:
python3 scripts/prepare_quick.py page-ready.json --knowledge-base <topic>.learnkbpython3 scripts/finalize.py page-ready.json --knowledge-base <topic>.learnkb --output-dir ./delivery --name learning-quickFor Systematic Study, prepare the required full knowledge base and review records before running:
python3 scripts/finalize.py page-ready.json --knowledge-base <topic>.learnkb --output-dir ./delivery --name learning-systematicWhat are this agent's strengths and limitations?
- A single content model drives the knowledge base, interactive HTML, and Markdown, reducing drift among delivery formats.
- Fine-grained mappings cover PDF pages, PPT slides, and EPUB chapters while explicitly separating source facts, model additions, and external verification.
- Systematic Study uses coverage, summary, action-rule, and independent content-review records; Quick Overview still scans the complete structure and reviews displayed citations.
- Method cards retain stable IDs and versions, while synthesized methods receive new IDs and explicit parent references instead of overwriting source methods.
- Scripts are offline-first and do not transmit materials; page notes, learner profiles, and wrong-answer records remain in local browser storage.
- The host agent must still read the source and author the page and method JSON files; the scripts do not automatically write the substantive learning content.
- Python 3.10, filesystem access, and local command execution are required, so pure chat environments cannot perform extraction, coverage validation, or HTML rendering.
- Scanned PDFs and image-heavy slides need separate OCR or vision capability; otherwise affected ranges remain explicitly unverified.
- MOBI/AZW/AZW3 has no built-in fallback and depends on Calibre's
ebook-convert. - Initial deliveries are
static-onlywithvisualReview: not-run; static and in-memory DOM checks do not constitute browser visual acceptance.
How does this agent compare with similar options?
The repository explicitly derives from two MIT-licensed projects: virgiliojr94/book-to-skill supplies material extraction, knowledge decomposition, and the base knowledge-base structure, while crayon-ai/book-to-webpage supplies the interactive page, themes, source display, and follow-up interaction design. learn-from-materials extends that foundation with a cross-agent Agent Skills workflow, two study depths, a versioned method library, an integrated whole-material methodology, coverage and content-review records, and an offline delivery pipeline. The upstream projects may be a narrower fit when only their foundational extraction or webpage conversion scope is needed; this repository is the broader option when one workflow must produce source-audited learning artifacts and reusable method records.
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| Material Learning Studio This agent | 92 · Excellent | Agent plugin / skillFree | ★ 775 | 2d ago | Python | Codex · Claude Code |
| LLM Wiki | 69 · Some gaps | CLIFree + model costs | ★ 1.7k | 1mo ago | Python | Codex · Claude Code · Claude.ai |
| LLM Wiki Agent | 52 · Major gaps | Agent plugin / skillFree + model costs | ★ 3.6k | 5d ago | Python | Codex · Claude Code |
| LLM Wiki | 85 · Good | Agent plugin / skillFree + model costs | ★ 1.4k | 18d ago | Python | Codex · Claude Code |
How does FollowAgents rate this agent?
Why each dimension lost points
The sources describe processing only user-designated materials and output directories, offline-by-default behavior, no automatic dependency installation, no access to unrelated credentials, and refusal to execute instructions embedded in materials. ZIP limits, browser request blocking, source locators, quotations, and path-redaction defaults are concrete. Deductions apply because write confirmation beyond the depth prompt is not comprehensively demonstrated, rollback depends largely on the user making backups, and dependency security offers isolation and pinning guidance without lockfiles, vulnerability scans, or supply-chain audit evidence.
The README, workflow contracts, and tests are consistent about binding, stale inputs, missing citations, heading indexes, relationship coverage, bilingual rendering, and publish-time rejection. Optional parsers have explicit fallback behavior, while tests assert actionable errors and absence of partial delivery. No deduction was made for lack of executed results because runtime reproduction is outside this static assessment; the score concerns the source-level handling shown.
The project documents supported formats, quick and systematic depths, bilingual behavior, multiple agent environments, and degraded behavior when OCR, browsers, or parsers are unavailable. It also states boundaries for pure-chat clients, scanned documents, MOBI-family files, and visual verification. Explicit depth requests trigger directly and an omitted depth prompts once, providing strong scenario and trigger coverage.
Repository structure, installation locations, commands, output naming, version metadata, examples, and limitations are well organized, and the complete MIT license is present. Deductions apply because no standalone FAQ content is shown, CHANGELOG is named but its contents are absent from the supplied evidence, and the Security Advisory route does not establish a clearly accountable maintainer, support commitment, or release-governance path; publisher identity remains unknown as instructed.
The proposed delivery combines interactive HTML, Markdown, canonical page data, a knowledge base, audit records, and an archive, while notes, quizzes, reusable methods, and traceable citations add substantial value over plain extraction. The cost-benefit score is reduced because systematic study requires agent-led reading and authoring of several structured records, and long or scanned materials may require high-reasoning models, long context, OCR, and browser tooling.
Claims can be tied to source blocks, quotations, page or chapter locations, hashes, coverage audits, rule ledgers, and content reviews. The design explicitly separates source-backed facts, uncovered areas, model additions, and externally verified material, while tests reject missing mappings, absent quotations, and stale inputs. Cross-source corroboration is reduced because the supplied support consists mainly of repository-authored documentation, synthetic fixtures, and internal tests, with no independent validation of semantic completeness or accuracy on real materials.
- This is a low-confidence static assessment: no scripts, tests, or browser checks were executed, so repository claims of passing behavior are not independently reproduced.
- Before processing private, copyrighted, or regulated material, verify the underlying AI platform's data policy and manually inspect outputs for personal data, credentials, absolute paths, and extensive source text.
- Install optional parsers only in an isolated environment with reviewed pinned versions; the supplied evidence shows neither a lockfile nor automated vulnerability auditing.
- Back up an existing skill and prior outputs before updating; no automatic rollback mechanism is demonstrated.
- Static checks, hashes, and structural ledgers cannot prove that summaries, relationships, or methods are semantically correct or complete; human source review remains necessary.