Data & Analysis scientific-computingbioinformaticspythonrstatsmcpremote-computereproducible-researchliterature-search

Wisp Science

A local-first workbench for literature search, scientific computing, and reproducible research.

FollowAgents review · FARS-2.1
Use with care
71/ 100 5-point scale 3.6 / 5
1 2 3 4 5 6
1Trust19 / 29 · 3.3/5

The README states that approval gates are enabled by default, Full Permission is opt-in, credentials use the OS keyring, encrypted synchronization is manual, turns can be undone, and explorations are isolated. UI tests also show diff or confirmation stages for promotion, discard, and abandonment, justifying full marks for user confirmation. The product can still read and write project files, execute shells, access SSH hosts, invoke MCP services, and contact model providers; no fine-grained sandbox or per-tool permission policy is shown, so least privilege is 2. Data-flow and sensitive-data claims are clear at a high level, but the evidence lacks a network destination inventory, provider-boundary explanation, log-redaction rules, or credential-lifecycle implementation. External effects and rollback benefit from approvals, diffs, branches, and undo, although recovery after permanent discard is unclear. Dependency controls include an exact ACP version, locked tool installation, and hashed vendored assets, but many Rust requirements are broad, Actions use mutable tags, and no vulnerability-scanning or remediation policy is shown, limiting dependency security to 1. The author, repository, contributors, DOI, signing services, and third-party-notice route are attributed, but the publisher is unverified and the accountable maintenance body is not fully identified, so source attribution is not complete.

2Reliability8 / 14 · 2.9/5

The README, Cargo workspace, build workflows, and UI tests describe a coherent desktop research workbench. Tests cover isolated build outputs, vendored-file hashes, and exploration and branch state transitions. Self-consistency is reduced to 2 because the citation advertises v1.2.0 while the Cargo workspace reports 1.3.0. Platform packages, Rust requirements, build tooling, and model configuration paths are reasonably documented, but no offline dependency mirror, compatibility matrix, or prerequisite diagnostic is shown, so dependency availability is 2. The supplied tests check script exit behavior and some warnings, but there is little evidence of user-facing errors, retries, or handling for network, SSH, MCP, and model-provider failures; failure messages score 1.

3Adaptability15 / 18 · 4.2/5

The audience and scenarios are unusually specific: literature research, Python and R computing, scientific databases, publication evidence, remote execution, and long-running research workflows, earning full marks. Capabilities, BYOM support, approval modes, and isolated explorations are described, but unsupported tasks, model constraints, and hazardous-operation boundaries are not comprehensively stated, so capability boundaries score 2. Explicit @, #, and / interactions, approval gates, and distinct branch and exploration states reduce accidental triggering, with tests confirming several state guards; however, the evidence does not establish natural-language router or tool-selection precision, so trigger precision is 2. Windows, macOS, Linux, WSL, SSH, GPU, Python, and R coverage is explicit and partly backed by build configuration, justifying full environment-fit marks.

4Convention12 / 18 · 3.3/5

The README has strong navigation across startup, research, projects, compute, extensions, and development, earning full information-architecture marks. It provides release packages, platform formats, and configuration links, but the supplied material lacks detailed system requirements, source-build instructions, and troubleshooting, so install notes score 2. Wisp/WISP, Runs, Skills, and Explorations are used consistently, although the citation/Cargo version mismatch prevents full naming-stability and versioning scores. A bundled demo, screenshot, case-study route, and extensive test scenarios provide useful examples, but no visible FAQ is included. Some operational limits are stated, including Full Permission and platform behavior, yet there is no consolidated known-limitations section, so that criterion scores 1. AGPL-3.0-only is consistent across README, Cargo, and the full license text, earning full marks. Releases and version metadata exist, but no changelog content is supplied. The repository owner and contributors are visible, while no explicit maintenance team, support commitment, security contact, or governance responsibility is established; maintenance responsibility therefore scores 1. Unknown publisher identity is not treated as suspicious.

5Effectiveness12 / 13 · 4.6/5

The workbench turns conversations and computation into reusable project files, persistent kernels, run logs, figures, branches, evidence capsules, and publication artifacts. Tests substantiate diff, merge-summary, and evidence-file workflows, supporting full output-usability marks. Combining scientific databases, Python/R, remote runtimes, MCP, tracked history, and publication evidence in one local-first workbench provides clear marginal value beyond a generic chat agent, also earning full marks. Open source, BYOM support, local storage, and a keyless demo reduce adoption cost, but users still bear model charges, environment setup, remote infrastructure, and AGPL compliance obligations; the evidence does not quantify resource use or cost, so cost-benefit scores 2.

6Verifiability5 / 8 · 3.1/5

Many feature claims point to named documentation, Cargo modules, build workflows, or concrete UI and script tests, giving reasonable traceability. The linked documentation bodies are not included here, however, and the roughly 80 databases, local-storage behavior, and encryption claims lack complete implementation evidence in the supplied files, so claim traceability is 2. README claims are partially corroborated by Cargo, CI, and tests, notably for platform builds, exploration behavior, and dependency integrity, but privacy and security statements remain chiefly README assertions; cross-source corroboration is therefore 2. Product claims, operational guidance, and test assertions are mostly distinguishable, yet broad statements such as data and credentials remaining on the user's machines are not qualified by implementation-level exceptions in this evidence, preventing full fact/inference separation.

Evidence confidence: Low Reviewed Aug 14, 2026 Reviewed revision 7b6b4cc0a36d
The upstream repository has new commits since this review. The score still applies to the reviewed revision shown and may not cover the latest changes.
Before you use it
  • The agent can modify project files, execute shells, connect to SSH hosts, and call external model and MCP services. Keep approval gates enabled and verify each provider's data destination before using sensitive research data.
  • Local-first and on-device credential claims are supported mainly by the README; the supplied evidence does not fully expose telemetry, log redaction, network endpoints, key rotation, or deletion policies.
  • The exploration tests describe discard as permanent removal. Confirm that an external backup exists, or that loss is acceptable, before approving it.
  • No dependency-vulnerability workflow, supply-chain threat model, or security-response contact is shown. Independently audit lockfiles, GitHub Actions references, and packaged artifacts before deployment.
  • The README citation says v1.2.0 while Cargo reports 1.3.0; verify that documentation, binaries, and source revision correspond before relying on release-specific behavior.
Review evidence [1][2][3][4][5][6][7][8]
See the full review method →

What does this agent do, and when should you use it?

Wisp Science is an open-source, local-first desktop workbench for scientific research, keeping project data, conversations, and credentials on the user's machines. Its agent can read and write project files, invoke the shell, and perform analysis through persistent Python and R kernels. Bundled MCP servers provide access to PubMed, GEO, and roughly 80 scientific databases, while computation can run locally, under WSL, or on SSH hosts. The workbench preserves sessions, figures, runs, decisions, and drafts, with isolated explorations, per-turn file-edit undo, and Evidence Capsules for frozen manuscript revisions. It runs on Windows, macOS, and Linux and supports OpenAI-compatible and Anthropic models, plus Codex and Claude Code through ACP.

After files, artifacts, and runtimes are attached to a project, the agent can inspect or modify project files, execute shell commands, and load reusable SKILL.md instructions on demand. Persistent Python and R kernels retain variables across cells, conversations, and restarts; longer jobs can be submitted as Runs with live logs. Users can register local, WSL, and SSH hosts, probe their hardware, and choose where computation executes. Bundled MCP servers query PubMed, GEO, and about 80 other scientific databases, while notebooks, PDFs, Office documents, and images can be previewed offline. Outputs such as figures, run records, decisions, and drafts remain associated with the project; explorations isolate experimental directions, and the Publication Workspace freezes manuscript revisions into verifiable Evidence Capsules.

  1. A bioinformatics researcher needs to query PubMed or GEO and run Python or R analysis without splitting the work across unrelated applications.
  2. A team with an SSH-accessible workstation or GPU server wants to submit long-running computations and follow their live logs.
  3. A scientist moving between a laptop and Windows WSL wants registered runtimes and persistent computational state across sessions.
  4. An author preparing a paper needs to retain analysis decisions, figures, runs, and manuscript revisions as verifiable Evidence Capsules.
  5. A researcher evaluating an alternative analysis path wants an isolated exploration that does not alter the project's mainline.
  6. A laboratory with data-control requirements wants project material and credentials kept on its own machines, with only explicit encrypted sync or transfer.

What are this agent's strengths and limitations?

Pros
  • Combines literature discovery, roughly 80 scientific databases, Python/R computation, and research records inside one desktop project.
  • Persistent Python and R kernels retain variables across cells, conversations, and application restarts.
  • Directly supports local, WSL, and SSH execution, including long-running Runs with live logs.
  • Has a concrete local-first boundary: project data, conversations, and credentials stay on user machines, and keys reside in the OS keyring rather than SQLite.
  • Explorations, per-turn file-edit undo, and Evidence Capsules address experimental isolation, traceability, and publication evidence.
  • Supports OpenAI-compatible and Anthropic models as well as Codex and Claude Code through ACP, avoiding a single documented model-provider dependency.
Limitations
  • Outside the bundled demo, model-backed work requires users to obtain and configure compatible model credentials, with any provider charges determined by their chosen service.
  • The documented delivery model is a desktop application; the supplied material does not establish a headless server deployment or embeddable library interface.
  • Encrypted project sync and transfer are manual, and nothing synchronizes in the background, adding steps for multi-device work.
  • SSH, WSL, GPU, and remote Runs depend on infrastructure that the user already operates; automatic provisioning is not documented.
  • The AGPL-3.0-only license may create compliance obligations for modification, distribution, or network-service use and should be reviewed before adoption.
  • The supplied material does not include complete source-build commands, resource benchmarks, or an itemized list of all supported scientific databases.

How do you install or deploy this agent?

Download the appropriate package from the repository's GitHub Releases. Windows is distributed as signed MSI or NSIS packages; macOS uses signed and notarized .dmg files for Apple Silicon and Intel; Linux provides .deb and AppImage packages for x86_64 and aarch64. The supplied material does not document a package-manager installation command. Building from source is mentioned, but complete build commands and toolchain prerequisites are not included in the supplied source.

How do you use this agent?

Install and launch the desktop application, then open the bundled RNA-seq demo to inspect a complete analysis trajectory without an API key. For a model-backed project, open Settings → Models and add an OpenAI-compatible or Anthropic model with its required credentials; Codex or Claude Code can instead be configured through ACP. Create a project, use @ to attach artifacts, files, and runtimes, use # to search saved sessions, and use / to apply a Skill. For remote computation, register a local, WSL, or SSH host, probe its hardware, and submit longer work as Runs to follow live logs. Beyond the bundled demo, the supplied material does not provide a specific first prompt or command invocation.

FAQ

Do I need an API key before I can evaluate it?
No. The bundled RNA-seq demo works without an API key. New model-backed projects require an OpenAI-compatible or Anthropic model configuration, or an ACP agent configuration.
Can the agent modify files and run commands?
Yes. It can read and write project files and invoke the shell. Approval gates remain enabled unless the user explicitly opts into Full Permission, and file edits made by a turn can be undone.
Does it automatically upload or synchronize my data?
The source states that data, conversations, and credentials stay on the user's machines, with keys stored in the OS keyring. Encrypted project sync and transfer are manual and do not run in the background; online model, database, and SSH operations still require network access.
Can it use a remote server or GPU?
Yes. Users can register SSH hosts, probe hardware, and submit long Runs with live logs; WSL and GPU runtimes are also explicitly described. The user must provide and configure the relevant host infrastructure.
Which desktop operating systems are supported?
Windows, macOS, and Linux are supported. Release formats include Windows MSI/NSIS, macOS .dmg packages for Apple Silicon and Intel, and Linux .deb/AppImage packages for x86_64 and aarch64.

Compare agents like this one

The same FARS review applied across the shortlist this agent qualifies for.

Related agents