Productivity & Collaboration knowledge-baseparallel-researchsource-ingestionthesis-researchwiki-querysession-memoryobsidianartifact-generation

LLM Wiki

Compile multi-agent research, sources, and session memory into durable, queryable topic wikis.

FollowAgents review · FARS-2.1
Recommended
85/ 100 5-point scale 4.3 / 5
1 2 3 4 5 6
1Trust24 / 29 · 4.1/5

The evidence shows read-only query profiles with restricted tools, path-specific sandbox guidance, and a private-adapter design using read/write roots, exact remote-resource allowlists, plan-hash approval, expected revisions, idempotency keys, and receipts. Sensitive-data cleanup is described as dry-run-first, accepting hidden or stdin input and requiring explicit application, so least privilege, confirmation, data-flow disclosure, sensitive-data handling, and external effects are handled thoroughly. Dependency security is reduced to 1 because CI uses major-version GitHub Actions and executes promptfoo@latest, remote curl downloads, and global package installation without supplied evidence of commit pinning, integrity verification, lockfiles, or vulnerability scanning. Archive/restore, dry runs, and uninstall paths provide recovery, but no uniform transactional rollback is shown for research, compilation, or bulk writes, so rollback is 2. Source metadata and auditing are required, but the complete attribution implementation and representative outputs are absent, so source attribution is 2.

2Reliability11 / 14 · 3.9/5

Runtime mirrors have an identified source of truth and are covered by sync, structure, documentation/version, concurrency, and behavioral checks, providing strong self-consistency evidence. Multi-runtime installation, upgrades, permission diagnostics, and local mode improve availability, but the product still relies on several agent CLIs, GitHub, npm, search services, and model providers without a complete offline or degraded-mode guarantee, so dependency availability is 2. Assertions and sandbox diagnostics return concrete failures, but the supplied material does not show unified error classification, retry behavior, or recovery messages across all workflows, so failure messages is 2.

3Adaptability16 / 18 · 4.4/5

The README addresses Claude Code, Codex, OpenCode, Pi/DS4, and generic agents, with scenarios covering queries, research, collection, ingestion, audits, and artifact generation. Read-only query-lite and full write-capable profiles, tool restrictions, sandbox paths, and local mode establish strong capability and environment boundaries. Trigger precision is reduced to 2 because the read-only Codex path is explicit, while the full skill can activate from natural language and fuzzy routing without supplied ambiguity rules or false-trigger evaluation results.

4Convention16 / 18 · 4.4/5

The documentation has clear navigation, architecture, extensive cross-runtime install/upgrade/removal/troubleshooting instructions, and numerous usable examples. The MIT license is complete, and the changelog is detailed with CI consistency checks. Naming stability is 2 because compatibility aliases and generated mirrors help, but multiple invocation styles, runtime wrappers, and recent terminology changes still impose migration complexity. Known limitations is 2 because best-effort OpenCode behavior, cache issues, and permission constraints are disclosed, but there is no comprehensive centralized limitations register. The repository, copyright holder, website, social account, and update paths are identified, but no maintenance commitment or support responsibility is established and publisher identity is unverified, so maintenance responsibility is 2; unknown identity is not used to penalize unrelated criteria.

5Effectiveness12 / 13 · 4.6/5

The system produces queryable wikis, reports, slides, catalogs, dataset manifests, and audits, while structured frontmatter, indexes, and Obsidian compatibility support practical use. Parallel research, opposing-evidence investigation, ingestion, compilation, auditing, and multi-runtime reuse provide clear marginal value over a basic retrieval workflow. Cost-benefit is reduced to 2 because profile sizes, token benchmarks, and research scales are mentioned, but actual monetary cost, latency, resource ceilings, and measured quality gains are not supplied; deep modes may also run 8–10 agents for hours.

6Verifiability6 / 8 · 3.8/5

Static assertions require source or sources, timestamps, confidence, summaries, and valid index structure, providing strong foundations for claim traceability. Thesis research is described as collecting evidence for and against a claim, and audits span artifacts, wiki content, and fresh research, but no concrete corroboration algorithm, minimum independent-source rule, or representative result is supplied, so cross-source corroboration is 2. Confidence fields and evidence-oriented workflows help distinguish claims from judgment, but the full protocol and output examples needed to show consistent separation of fact, inference, and conclusion are absent, so fact-inference separation is 2.

Evidence confidence: Low Reviewed Aug 16, 2026 Reviewed revision b513c7007fc4
The upstream repository has new commits since this review. The score still applies to the reviewed revision shown and may not cover the latest changes.
Before you use it
  • Do not treat README security or quality claims as independently verified; this assessment did not execute code, tests, agent workflows, or network operations.
  • CI and installation paths include promptfoo@latest, major-version Actions, global npm installation, and remotely fetched curl content; pin versions and verify integrity before controlled deployment.
  • Full research mode can write to the wiki, access external directories, use the network, and launch multiple agents; prefer query-lite and grant write access only to explicitly required paths and operations.
  • Long-running parallel research may incur substantial model cost, latency, and stored output; set budgets, time limits, source scope, and human-review gates in advance.
  • Archive restoration exists, but the supplied evidence does not establish atomic rollback for every bulk ingestion, compilation, or research write; maintain separate backups or version control for important wikis.
  • Publisher identity is unknown; adopters should independently establish responsibility for maintenance, updates, and security response.
Review evidence [1][2][3][4][5]
See the full review method →

What does this agent do, and when should you use it?

LLM Wiki is a filesystem-based knowledge workflow for AI coding agents that turns rough ideas, external sources, and research runs into isolated topic wikis. Claude Code is its primary distribution and UX target, while documented packages also cover Codex, OpenCode, Pi, and portable file-capable agents. It supports parallel research, thesis testing, source ingestion, article compilation, inventory and dataset catalogs, audits, and artifact generation. Knowledge, immutable raw sources, outputs, logs, and optional session memory are stored locally as Markdown and can be opened as independent Obsidian vaults. It is a strong fit for users who want research to remain inspectable and maintainable, provided their agent can access the filesystem and, for online research, the network.

A workflow normally begins with /wiki:research, /wiki:thesis, /wiki:ingest, or /wiki:collect. The agent searches or reads web pages, files, PDFs, quoted text, Git documentation repositories, MediaWiki sources, message archives, and Wayback snapshots, then stores evidence inside a selected topic wiki. Research mode dispatches 5, 8, or 10 parallel agents across academic, technical, applied, news, contrarian, historical, adjacent, and data-oriented paths; thesis mode defines variables and falsification criteria, gathers evidence on both sides, and issues a verdict. /wiki:compile synthesizes raw material into cross-linked, confidence-rated articles, while /wiki:query offers quick, standard, deep, and list retrieval. /wiki:audit, /wiki:librarian, and /wiki:lint examine provenance, freshness, quality, and structural integrity. Ideas, Projects, inventory records, dataset manifests, archives, and the read-only portfolio view manage work beyond articles, while /wiki:output produces summaries, reports, study guides, slides, timelines, glossaries, and comparisons. An optional .sessions/ layer captures redacted operational digests and feedback candidates, which enter topic knowledge only after explicit promotion.

  1. A researcher entering a new field can run /wiki:research "topic" --new-topic --deep to create a dedicated wiki and investigate the subject from several parallel perspectives.
  2. An analyst evaluating a concrete claim can use /wiki:thesis to collect supporting, opposing, mechanistic, and review evidence before receiving a supported, mixed, contradicted, or insufficient-evidence verdict.
  3. A team with web pages, PDFs, Git documentation repositories, MediaWiki dumps, or message CSVs can ingest them into a consistent local corpus and compile them into searchable articles.
  4. A user moving between Claude Code, Codex, OpenCode, and Pi can keep one filesystem wiki as the shared knowledge layer and load the runtime-specific plugin or skill.
  5. An Obsidian user who wants agent-maintained notes can open each topic directory as a separate vault while retaining aliases, tags, graph links, and ordinary Markdown navigation.
  6. A long-running project can retain redacted session checkpoints and selected feedback without placing complete transcripts directly into its durable knowledge base.

What are this agent's strengths and limitations?

Pros
  • Covers the complete path from parallel research and source capture through compilation, querying, auditing, and deliverable generation rather than stopping at a chat answer.
  • Shares one behavior model across Claude Code, Codex, OpenCode, Pi, and portable agents, with a compact roughly 2.8 KB protocol for read-only queries.
  • Keeps each subject in an isolated Markdown wiki with separate raw evidence, synthesized articles, outputs, indexes, and append-only activity logs.
  • Thesis research deliberately assigns supporting and opposing paths and directs later rounds toward the weaker evidence side to counter confirmation bias.
  • Session digests and feedback remain separate candidates until a user explicitly promotes them into durable topic knowledge.
Limitations
  • The project is explicitly Claude-first; Codex and OpenCode packages are generated from the Claude source, so runtime UX and validation are not identical.
  • The OpenCode profile is documented as best effort and lacks a provider-specific live quality gate; query-only sessions should also disable its write and shell permissions.
  • The hub commonly resides outside the working project, requiring additional filesystem and sandbox configuration for Codex, OpenCode, or nono.
  • Online research depends on the host agent's network and search facilities; OpenCode web search specifically requires OPENCODE_ENABLE_EXA=1.
  • The repository recommends adding qmd for local search beyond roughly 100 articles, so larger corpora may need another retrieval component.
  • Automated session capture depends on trusted hooks; the main skill still works without them, but automatic memory capture does not.

How do you install or deploy this agent?

Choose a supported agent runtime and grant it access to the wiki directory; online research also requires network and search access. For Claude Code, run claude plugin install wiki@llm-wiki. For Codex, run codex plugin marketplace add nvk/llm-wiki, followed by codex plugin add wiki@llm-wiki, then start a new Codex thread. For OpenCode, add https://raw.githubusercontent.com/nvk/llm-wiki/master/plugins/llm-wiki-opencode/skills/wiki-manager/SKILL.md to the instructions array in opencode.json and allow ~/.config/llm-wiki/** plus the actual wiki path under permission.external_directory; use the wiki-query/SKILL.md URL for a smaller read-only setup. Pi can load the full workflow with pi --skill path/to/llm-wiki/plugins/llm-wiki-opencode/skills/wiki-manager/SKILL.md. Other file-capable agents can copy the repository's root AGENTS.md, or copy profiles/query-lite/SKILL.md for read-only queries. No separate API credential is specified, although the selected agent runtime and its search provider may require their own configuration.

How do you use this agent?

For a first working run, execute /wiki:research "nutrition" --new-topic; this creates a topic wiki and begins research. In an existing wiki, add a source with /wiki:ingest https://example.com/article, compile pending material with /wiki:compile, and ask /wiki:query "How does fiber affect mood?". Codex exposes equivalent entry points such as @wiki research "hardware wallet threat models" and the explicit read-only $wiki-query "What does the wiki say about hardware wallet threat models?". Add --deep or --min-time 1h for broader, multi-round research, and use /wiki:thesis "claim" when the task is to test a specific proposition. Review and trust the bundled hooks through Codex /hooks if automated session capture is desired; @wiki remains usable without hook trust. When the hub sits outside the project, configure its hub path and allow that exact path in the Codex, OpenCode, or nono sandbox.

FAQ

Do I have to use Claude Code?
No. Claude Code is the principal design and distribution target, but the repository also documents a Codex marketplace plugin, OpenCode and Pi skills, and portable instructions for other file-capable agents. Codex and Claude Code have explicit native installation paths.
Does it store complete chat transcripts in the wiki?
Not by default. Balanced capture records harness metadata, small redacted events, state, and Markdown digests rather than full transcript bodies. Digests and feedback become topic notes only after explicit promotion.
Can I query a wiki without allowing writes?
Yes. Codex provides the explicit $wiki-query skill, and other agents can load profiles/query-lite/SKILL.md. The query-lite protocol is read-only and requires only file reading and search capabilities.
Which filesystem permissions are required?
The agent needs access to ~/.config/llm-wiki and the actual wiki data directory; mutating workflows need read-write access to the wiki. Codex plugin installation additionally needs read-write access to $HOME/.codex, while OpenCode also requires an external_directory allowance.
How does it handle large datasets or media collections?
It can create external dataset manifests, profiles, samples, and query recipes, and its collector can build media catalogs with bounded caching. Large row-oriented data and large media collections are represented as datasets or corpus inventory records rather than being embedded wholesale in ordinary Markdown articles.

Compare agents like this one

The same FARS review applied across the shortlist this agent qualifies for.

Related agents