Gyoshu Research Lab

A professor-and-TA agent pair that turns a research goal into reproducible Jupyter notebooks and publication-style reports.

Stars
★ 241
Last updated
7mo ago
Primary language
TypeScript

At a glance

How it runs
Agent plugin / skillMCP serverCLI
Works with
Portable with changesClaude Code
Cost
Free software; you pay for model usage
Setup effort
Medium · a few setup steps
You'll need
OpenCode v0.1.0+ or Claude CodePython 3.10+Node.js 18+npm or bunproject .venv virtual environmentpsutil (optional)Shell / CLINetwork accessLocal filesystemMCP Server
Typical use
A data scientist wants a full exploratory analysis — multi-dimensional Binance USD-M futures visualizations, dendrogram correlation heatmaps, rolling statistics — preserved as a reproducible notebook rather than a throwaway script.
Not a fit if
  • Users on native Windows who won't set up WSL2
  • Researchers who want a browser UI instead of a CLI agent

What does this agent do, and when should you use it?

Gyoshu and Jogyo are an end-to-end research automation system that plugs into OpenCode, splitting the work between a planning orchestrator and a code-running executor. Gyoshu (교수, Professor) scopes the research plan and manages sessions; Jogyo (조교, Teaching Assistant) runs Python in a persistent REPL, executes experiments and produces outputs. Two more agents round it out: Baksa (박사, PhD Reviewer) adversarially challenges every claim and computes a trust score, and Jogyo Paper Writer turns raw results into narrative reports. Results are organized by structured markers such as [OBJECTIVE], [HYPOTHESIS], [FINDING], [STAT:ci] and [STAT:effect_size], which drive automatic capture into notebooks/*.ipynb and outputs under reports/ (figures, models, report.md). It installs three ways: as an OpenCode plugin listed in opencode.json, as an npm/bun CLI installer (bunx gyoshu install), or as an MCP server for Claude Code exposing 12 tools including python_repl, research_manager and notebook_writer. It runs against your project's own .venv and docs state it never touches system Python or existing project files, keeping runtime artifacts in the OS temp directory.

Given a research goal, Gyoshu first produces a structured plan with objectives and hypotheses, then Jogyo executes analysis in a persistent Python REPL where variables survive across sessions, similar to a live Jupyter kernel. Execution is instrumented through printed markers — [OBJECTIVE], [HYPOTHESIS], [DATA], [METRIC:accuracy], [FINDING], [CONCLUSION] — which the tooling parses to capture the research trail. Every experiment is written to notebooks/*.ipynb, while figures, trained models and a report.md land under reports/<project>/. A quality gate requires [STAT:ci] and [STAT:effect_size] within 10 lines before any [FINDING]; findings that fail are downgraded to Exploratory Observations, and ML work additionally needs [METRIC:baseline_*] and [METRIC:cv_*] or loses 20 and 25 trust points. The Baksa reviewer challenges each claim and assigns a trust score: 80 or above is Verified, 60–79 is accepted with caveats, below 60 is rejected as requiring rework. Alongside the command surface (/gyoshu, /gyoshu-auto, /gyoshu plan, /gyoshu continue, /gyoshu report, /gyoshu list, /gyoshu search, /gyoshu doctor) there are interactive, autonomous and REPL modes, plus session continue, replay and branching. The MCP build exposes python_repl, research_manager, gyoshu_snapshot, checkpoint_manager, notebook_writer and notebook_search, among 12 research tools in total.

  1. A data scientist wants a full exploratory analysis — multi-dimensional Binance USD-M futures visualizations, dendrogram correlation heatmaps, rolling statistics — preserved as a reproducible notebook rather than a throwaway script.
  2. An ML engineer needs a baseline-plus-cross-validation modeling run on a classic dataset (Titanic survival, Iris clustering) and wants to set the goal with /gyoshu-auto and walk away.
  3. A researcher who insists every conclusion carry a confidence interval and effect size, and who wants an adversarial reviewer to attack claims before they reach a report.
  4. A Claude Code user who wants python_repl and notebook_writer available as MCP tools inside an existing coding session instead of switching to OpenCode.
  5. Someone maintaining several long-running research threads who needs /gyoshu list, /gyoshu search and /gyoshu continue to find and resume work across sessions and notebooks.
  6. A quant or data-product team that produces insights in Gyoshu and hands them to a separate build tool (such as Oh-My-OpenCode) to ship the resulting feature.

How do you install or deploy this agent?

Start with a project-local Python 3.10+ virtual environment. Gyoshu uses the project's .venv and never modifies system Python.

python3 -m venv .venv
.venv/bin/pip install pandas numpy scikit-learn matplotlib seaborn

Option 1 — install as an MCP server for Claude Code (requires Node.js 18+).

git clone https://github.com/Yeachan-Heo/My-Jogyo.git
cd My-Jogyo/src/mcp
npm install && npm run build
claude mcp add gyoshu-mcp "$(pwd)/build/index.cjs"
claude mcp list
# Should show: gyoshu-mcp: ✓ Connected

Option 2 — install as an OpenCode plugin; OpenCode auto-installs it from npm on next startup.

{
  "plugin": ["gyoshu"]
}

Option 3 — use the CLI installer, which adds Gyoshu to your opencode.json for you.

bunx gyoshu install

Or install globally first, then run the installer.

npm install -g gyoshu
gyoshu install

Contributors can instead clone the repo, run bun install, and point opencode.json's plugin entry at file:///path/to/My-Jogyo. Verify afterwards with bunx gyoshu check or /gyoshu doctor inside OpenCode.

How do you use this agent?

Launch OpenCode and say hello, then check status.

opencode

Inside the OpenCode session, use the command surface:

/gyoshu
/gyoshu analyze customer churn patterns in the telecom dataset
/gyoshu-auto classify iris species using random forest
/gyoshu report
/gyoshu continue

Autonomous mode suits clear goals that need no mid-course steering; interactive mode suits iterative exploration; /gyoshu repl <query> is for quick checks and debugging. To produce findings that survive the quality gate, emit the markers in pairs from your code:

print("[OBJECTIVE] Predict wine quality from physicochemical properties")
print("[HYPOTHESIS] Alcohol content is the strongest predictor")
print("[STAT:ci] 95% CI [0.82, 0.94]")
print("[STAT:effect_size] Cohen's d = 0.75 (medium)")
print(f"[METRIC:accuracy] {accuracy:.3f}")
print("[FINDING] Alcohol shows r=0.47 correlation with quality")
print("[CONCLUSION] Hypothesis supported - alcohol is key predictor")

To hand context to another LLM, tell it to read AGENTS.md in the Gyoshu directory, or paste this prompt:

I've installed Gyoshu. Read AGENTS.md and help me run /gyoshu to analyze my data.

If something breaks, run /gyoshu doctor. A "Session locked" message can be cleared with /gyoshu unlock <sessionId> after confirming no process is running.

What are this agent's strengths and limitations?

Pros
  • Four distinct agent roles (professor plans, TA executes, PhD reviewer attacks, grad student writes) separate planning from execution instead of collapsing them into one loop.
  • The persistent Python REPL keeps variables alive across sessions and every experiment lands in notebooks/*.ipynb, which is the stated source of truth for reproducibility.
  • The quality gate is a hard, inspectable rule rather than advice: findings need [STAT:ci] and [STAT:effect_size], ML needs baseline and cross-validation metrics, and trust scores map to Verified / caveated / rejected.
  • Three documented entry points — OpenCode plugin, npm/bun CLI, and MCP server for Claude Code — let you adopt it in whichever coding agent you already use.
  • Runtime sockets and locks go to OS temp directories, and the docs explicitly promise that .venv, data/ and other existing project files are left untouched.
Limitations
  • It only runs inside a host coding agent (OpenCode or Claude Code); there is no standalone mode, so moving to another agent means writing the integration yourself.
  • The project must contain a .venv, and the runtime needs Python 3.10+ plus Node.js 18+; missing any of these produces a hard failure such as "No .venv found".
  • Native Windows is unsupported — Linux, macOS and WSL2 only — which is a real blocker for Windows-first teams.
  • Licensing is inconsistent: the README says MIT while the repository metadata lists the license as unknown, so reuse rights should be verified before adoption.
  • Supporting materials are incomplete: the demo GIF is still "coming soon" and referenced docs such as docs/user-guide.md and AGENTS.md are not included in the source, so behavior has to be validated hands-on.

How does this agent compare with similar options?

The repository names one adjacent tool: Oh-My-OpenCode (code-yeongyu/oh-my-opencode). Gyoshu focuses on research and analysis, Oh-My-OpenCode on product development and shipping features, and each is documented as fully standalone. Gyoshu does not require it, but the README describes an optional combined flow — analyze with Gyoshu, then implement the resulting insight with Oh-My-OpenCode.

Key facts side by side with the most closely related agents.

Agent Source review Form / cost Stars Updated Language Full support on
Gyoshu Research Lab This agent 39 · Major gaps Agent plugin / skillFree + model costs ★ 241 7mo ago TypeScript Claude Code
optim-agent 65 · Some gaps Library / SDKFree + model costs ★ 801 1mo ago Python Codex · Claude Code
MiroFlow Research Agent 53 · Major gaps CLIFree + model costs ★ 3.1k 8mo ago Python —
Metaflow 77 · Good Library / SDKFree ★ 10k 5d ago Python —

How does FollowAgents rate this agent?

FollowAgents source review · FARS-2.1
Major gaps
39/ 100 5-point scale 2.0 / 5
Trust 8/29
Reliability 5/14
Adaptability 9/18
Convention 8/18
Effectiveness 6/13
Verifiability 3/8
Why each dimension lost points
Trust8 / 29 · 1.4/5

Evidence shows the install script is executed via curl piped to bash, and the README states Gyoshu executes Python code and writes to notebooks/ and reports/ directories, which are local filesystem writes. However, the repository does not provide install.sh source, permission declarations, or sandboxing details, so least_privilege cannot be confirmed and scores 1. For user_confirmation, /gyoshu-auto is described as a hands-off autonomous mode, and the README does not state whether confirmation is required before destructive operations, so user_confirmation scores 1. For data_flow_transparency, the README mentions runtime files (sockets, locks) are stored in the OS temp directory but does not explain whether data is sent externally or how LLM calls transmit data, so data_flow_transparency scores 1. For sensitive_data_handling, the repository contains no mention of credentials, PII, data masking, or key management, so it scores 0. For dependency_security, package.json declares sanitize-html and zod with caret ranges, and pyproject.toml has no runtime dependencies, but no lockfile or vulnerability audit is provided, so dependency_security scores 1. For external_effects, the README states Gyoshu will not modify .venv/, data/, or other existing files, but the install script writes to system paths and no uninstall method is described, so external_effects scores 1. For rollback, atomic-write tests show temp files are cleaned up on failure, but no overall install or research workflow rollback mechanism is provided, so rollback scores 1. For source_attribution, the README references OpenCode and Oh-My-OpenCode with links but does not detail code provenance, third-party reuse, or acknowledgements, so source_attribution scores 1.

Reliability5 / 14 · 1.8/5

For self_consistency, the README's described features (REPL, notebook generation, adversarial verification) align with package.json name/description, but pyproject.toml version is 0.1.0 while package.json is 0.5.1, an inconsistency, so self_consistency scores 1. For dependency_availability, package.json declares peerDependencies @opencode-ai/plugin >=1.0.0 and engines bun >=1.0.0, but no lockfile or install verification is provided, so dependency_availability scores 1. For failure_messages, the README provides a troubleshooting table (No .venv found, Bridge failed to start, Session locked, etc.) but does not show actual error message formats or log examples, so failure_messages scores 1.

Adaptability9 / 18 · 2.5/5

For audience_and_scenarios, the README clearly distinguishes interactive, autonomous, and REPL modes and provides use cases for researchers, so audience_and_scenarios scores 2. For capability_boundaries, the README states Gyoshu is standalone and does not require Oh-My-OpenCode, but does not clearly state what it cannot do (e.g., no native Windows, no non-Python languages), so capability_boundaries scores 1. For trigger_precision, the command list (/gyoshu, /gyoshu-auto, /gyoshu plan, etc.) is clear, but command conflicts or mis-trigger handling are not described, so trigger_precision scores 1. For environment_fit, the README lists a Linux/macOS/Windows(WSL2) support matrix and Python 3.10+ requirement, so environment_fit scores 2.

Convention8 / 18 · 2.2/5

For information_architecture, the README is well-structured with installation, commands, workflow, project structure, and troubleshooting sections, so information_architecture scores 2. For install_notes, three installation methods (curl, clone, npm/bunx) and a verification command are provided, so install_notes scores 2. For naming_stability, the project is named Gyoshu/Jogyo in the README, gyoshu in package.json and pyproject.toml, but the repository is My-Jogyo, an inconsistency, so naming_stability scores 1. For examples_and_faq, the README provides code and command examples but no dedicated FAQ section, so examples_and_faq scores 1. For known_limitations, the README mentions Windows WSL2-only and Python 3.10+ but does not systematically list other limitations, so known_limitations scores 1. For license, the README and package.json both declare MIT, but no LICENSE file content is provided in the repository root, so license scores 2 (consistent declaration but missing file evidence). For versioning_changelog, the README references CHANGELOG.md but its content is not provided, and package.json and pyproject.toml versions are inconsistent, so versioning_changelog scores 1. For maintenance_responsibility, package.json has author and repository fields but no maintainer, support channel, or response time is described, so maintenance_responsibility scores 1.

Effectiveness6 / 13 · 2.3/5

For output_usability, the README states outputs are reproducible .ipynb files, reports, figures, and models, and provides a project structure, so output_usability scores 2. For marginal_value, the project offers research automation and notebook generation, but similar tools exist and the README provides no comparison or unique value proof, so marginal_value scores 1. For cost_benefit, installation requires OpenCode, Python 3.10+, and bun, and autonomous mode may consume many LLM calls, but the README does not describe resource consumption or cost, so cost_benefit scores 1.

Verifiability3 / 8 · 1.9/5

For claim_traceability, README feature claims (e.g., adversarial verification, trust scores) are partially reflected in test files (adversarial-flow.test.ts), but tests only validate data structures rather than end-to-end behavior, so claim_traceability scores 1. For cross_source_corroboration, README, package.json, pyproject.toml, and test files are partially consistent, but version numbers are inconsistent and install.sh source is missing, so cross_source_corroboration scores 1. For fact_inference_separation, some README statements (e.g., 'will not modify .venv/') are assertions without code or test evidence, so fact_inference_separation scores 1.

Risks and how to mitigate them
  • Not found in source: sensitive-data handlingUse dedicated, low-privilege, revocable API keys — never production credentials — and keep secrets out of logs.
  • The install script is executed via curl piped to bash, and install.sh source is not provided in the repository, so its behavior cannot be statically verified; users should download and review the script before executing.
  • The README does not explain whether data is sent externally or how LLM calls transmit data, and does not mention credential or sensitive data handling; exercise caution when processing sensitive data.
  • package.json version is 0.5.1 while pyproject.toml version is 0.1.0; this inconsistency may cause dependency resolution or update issues.
  • No LICENSE file content is provided; only the README and package.json declare MIT, so legal compliance requires further confirmation.
  • The autonomous mode /gyoshu-auto is described as hands-off, but it is not stated whether user confirmation is required before destructive operations; use in a controlled environment is recommended.
Evidence confidence: Low Reviewed Oct 11, 2026 Reviewed revision b3a3a37a8213
See the full review method →

FAQ

Is Gyoshu itself paid, and what else will I be billed for?
The README states MIT licensing, so the tool is free. It runs inside OpenCode or Claude Code, which means you supply and pay for your own coding-agent and model account; those costs are outside the project.
Do I have to use OpenCode, or does Claude Code work too?
Claude Code is a documented path. Clone the repo, run npm install && npm run build inside src/mcp, then register it with claude mcp add gyoshu-mcp "$(pwd)/build/index.cjs" to get the 12 MCP tools including python_repl, research_manager and notebook_writer.
Why was my conclusion downgraded to Exploratory in the report?
Every [FINDING] needs [STAT:ci] and [STAT:effect_size] within the 10 preceding lines or it loses 30 trust points and becomes an exploratory observation. ML runs additionally need [METRIC:baseline_*] and [METRIC:cv_*], costing 20 and 25 points if missing. Anything scoring below 60 is rejected and needs rework.
Will it touch my existing data or virtual environment?
The docs say no: it uses your existing project .venv, leaves system Python alone, and does not modify data/ or other existing files. Generated notebooks go to notebooks/ and reports, figures and models to reports/. Sockets and locks live in the OS temp directory.
Do I need a database or a GPU to run it?
No. The stated requirements are Python 3.10+, Node.js 18+, and OpenCode v0.1.0+ or Claude Code, with psutil optional for memory tracking. It is aimed at scikit-learn-scale analysis, but the repository says nothing about GPU-dependent workloads.
View on GitHub ↗ Install ↓

Compare agents like this one

The same FARS review applied across the shortlist this agent qualifies for.

Related agents