Jelly: Local Multi-Bot Agent Workspace
Run persistent AI agents on your own Mac, Linux box or VM, with a real mobile web UI and a local Chrome you can take over.
- Source repo
- dctanner/jelly
- Stars
- ★ 55
- Last updated
- 2d ago
- License
- MIT
- Primary language
- TypeScript
- FA score
- 52/100 · Major gaps
At a glance
- Works with
- Platform-specificChatGPT · OpenAI API
- You'll need
- Typical use
- A solo founder running an entire business from one VM who wants separate agents per project directory to handle ops, scraping and routine tasks.
- Main limitation
- It is explicitly not a sandbox: agents execute shell commands and access files with your OS user's permissions and without per-tool approval, so anyone with network access can control its tools.
- Source review
- 52/100 · Major gaps
What does this agent do, and when should you use it?
Jelly is a local-first multi-agent system aimed at people running an entire business from a single VM. It is built with Bun, React, SQLite and the Pi agent harness; development serves on port 5173, production on port 3100, instance data lives in .jelly/ by default, and user-wide MCP configuration plus OAuth caches live in ~/.jelly/. You can create as many agents as you want, group them into project directories, and each agent gets its own browser window, profile and private desktop. Running agents requires a ChatGPT subscription or an OpenAI API key — there is no no-key demo mode — and model conversations are sent to your chosen provider. Jelly describes itself as a trusted-user tool rather than a sandbox: agents run shell commands and touch files with your OS user's permissions, while sudo requests go through a private authentication card whose password is never sent to the model or stored by Jelly.
Agents ship with shell tools, subagents and MCP connections, plus file uploads, inline images and downloadable artifacts. The structured browser tools let an agent inspect bounded visible text and element references, click and fill non-secret fields, manage session-local tabs, scroll, wait for text or readiness, and read redacted error diagnostics; these share the per-agent isolation and private-control gate, references expire on new observations, navigation or handoff, and snapshots cover the top-level DOM only. When an agent needs to sign in somewhere, you take over its browser via VNC inside the web UI; clipboard transfers are text-only, capped at 12,000 characters, and never reach agents, chat history or logs. The apply_patch tool adds Codex-style multi-file add/update/delete/move patches on both OpenAI API and ChatGPT, requiring exact unique context, rejecting symlinks and hardlinked sources, capping patches at 1 MiB per file/patch, 64 operations and 16 MiB combined working set, with non-transactional writes reported in a mutation ledger. Administrative work goes through request_sudo and needs Linux, sudo with askpass support and /usr/bin/python3. On the first message, Jelly makes one small tool-free OpenAI request to suggest a sea-themed agent name.
- A solo founder running an entire business from one VM who wants separate agents per project directory to handle ops, scraping and routine tasks.
- Someone who must let an agent sign into a third-party site but refuses to hand credentials to a model: the agent requests a login, then the human takes over via VNC in the web UI.
- A user with an existing ChatGPT subscription who wants to avoid buying separate API quota, connecting via the one-time device code in Settings.
- Sysadmin chores such as installing packages or managing services, where the agent calls request_sudo and the human approves through a private authentication card.
- Remote use from a phone or laptop: enable Tailscale and add the web UI to the iPhone home screen for near-native feel.
- Developers needing multi-file edits in one shot: apply_patch adds, updates, deletes and moves files over OpenAI API or ChatGPT.
How do you install or deploy this agent?
Requires Bun 1.3.9+ and Node.js 24+; the managed browser desktop and sudo handoffs require Linux.
git clone https://github.com/dctanner/jelly.git
cd jelly
bun install
bun run devOpen http://127.0.0.1:5173. Local environment values go in .env.local, which is Git-ignored; set FIRECRAWL_API_KEY there to enable web search and fetching.
For production:
bun run build
bun startThen open http://127.0.0.1:3100, stopping development first since both modes use API port 3100 by default.
On Debian/Ubuntu x86-64 with Google Chrome or Chromium already installed, install the optional browser desktop:
bun run setup:desktopSet JELLY_BROWSER_PATH if the browser lives outside its usual system location. To reach Jelly over a private tailnet:
bun run dev:tailscaleHow do you use this agent?
Before sending a message you must connect an account or add a key in Settings — ChatGPT sign-in uses a one-time device code, so open the approval link shown in Settings and enter the code. Automatic mode prefers ChatGPT, then an API key; there is no no-key demo mode.
Click the top-left + to immediately create and open New Agent, which inherits the current project or gets a private workspace; before the first message, click the centered name or avatar to edit in place. Editing and saving the name before the first message opts out of automatic naming.
For eligible ChatGPT subscriptions (Pro 500 or Enterprise) or OpenAI API access, select GPT-6 Astra Ultrafast in Settings → Model or the composer's model menu; it requests Astra with service_tier: "ultrafast" while standard Astra stays the default, and account-access errors are surfaced rather than silently falling back.
In the Agent computer panel choose Take control to drive a given agent's browser, then use Paste to remote and Copy from remote for text clipboard transfers; if control is stuck after a session expires, use Agent computer → Recover lost control… to disconnect the old controller and clear private tabs and the remote clipboard while stored website sign-ins remain.
Development checks:
bun run check
bun run test:desktopbun run check covers TypeScript, structured-browser Chromium tests, patch filesystem tests and local provider-transport tests; bun run test:desktop additionally requires the Linux desktop runtime. Tests use temporary data and deterministic fixtures and need no paid model calls.
What are this agent's strengths and limitations?
- Per-agent isolation is concrete: separate browser windows, profiles, tabs, clipboard and human-control state, with at most eight browser sessions allocated at once.
- Credential handling is specific: sudo passwords are entered in a private card, never sent to the model or saved by Jelly, and clipboard transfers are text-only, capped at 12,000 characters and excluded from history and logs.
- You can use an existing ChatGPT subscription instead of buying separate API quota, and the device-code login works over Tailscale without a localhost redirect.
- apply_patch gives Codex-style multi-file moves with explicit size and operation limits plus a mutation ledger for partial failures, alongside the unchanged Pi edit/write tools.
- The local data layout is clear and auditable: .jelly/ holds conversations, credentials, browser profiles and artifacts while ~/.jelly/ holds MCP configuration and OAuth caches.
- It is explicitly not a sandbox: agents execute shell commands and access files with your OS user's permissions and without per-tool approval, so anyone with network access can control its tools.
- Model-provider lock-in: running agents requires a ChatGPT subscription or an OpenAI API key with no no-key demo mode, and local storage does not mean local inference.
- Platform limits are real: the managed browser desktop and sudo handoffs need Linux, and the desktop runtime now also requires xclip and /usr/bin/python3.
- apply_patch writes are not transactional, so a failure can leave partial changes, and paths are not sandboxed; symlinks and hardlinked sources are rejected.
- Setup is multi-step: Bun 1.3.9+ and Node.js 24+ are both required, and development and production share API port 3100, so you must stop one before starting the other.
How does this agent compare with similar options?
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| Jelly: Local Multi-Bot Agent Workspace This agent | 52 · Major gaps | — | ★ 55 | 2d ago | TypeScript | ChatGPT · OpenAI API |
| Auto Browser | 62 · Some gaps | Self-hosted serviceFree + model costs | ★ 897 | 1d ago | Python | OpenAI API · Claude API |
| Aiden | 65 · Some gaps | CLIFree + model costs | ★ 844 | 20d ago | TypeScript | ChatGPT · OpenAI API · Claude API |
| JoySafeter | 59 · Major gaps | Self-hosted serviceFree + model costs | ★ 312 | 5mo ago | Python | — |
How does FollowAgents rate this agent?
Why each dimension lost points
The README explicitly states Jelly is a trusted-user tool, not a sandbox: agents run shell commands and access files with the OS user's permissions and without per-tool approval; apply_patch paths are not sandboxed; sudo commands run immediately when policy permits. This is a deliberate trade-off, so least_privilege and user_confirmation score only 1. Positives: sudo passwords go through a private auth card, are never sent to the model or stored; browser secret-field detection is heuristic with guidance to use private sign-in; clipboard transfers are text-only, capped at 12,000 characters, and excluded from chat/logs; .jelly/ and ~/.jelly/ are explicitly kept out of source control. Deductions: dependencies are only partially pinned (several ^ ranges) with no lockfile evidence, audit, or vulnerability mitigation, so dependency_security is 1; apply_patch is explicitly non-transactional and can leave partial changes, with rollback limited to a mutation ledger report, so rollback is 1; external_effects is 1 because browser clicks, navigation, and typing have irreversible website side effects; source_attribution is 1 because only MIT and a one-line noVNC MPL-2.0 note are given and the publisher is unverified.
Self-consistency is good: README descriptions of naming, browser isolation, and patch limits align with the test files (agent-naming.test.ts, apply-patch-transport.test.ts), and tests use local stubs and deterministic fixtures without paid model calls, so self_consistency is 2. Failure messages are deliberately designed: naming failures/timeouts keep 'New Agent' without retry or task failure; account-access errors are surfaced rather than silently falling back; patch failures are recorded in a mutation ledger, so failure_messages is 2. Deduction: dependency availability requires Bun 1.3.9+, Node 24+, a Linux desktop runtime, system Chrome/Chromium, xclip, /usr/bin/python3, and sudo askpass, with no documented fallback path, so dependency_availability is only 1.
Audience and scenarios are clear: users running a whole business from a single VM/Mac/Linux box who want multiple agents and mobile access, so audience_and_scenarios is 2. Capability boundaries are specific: browser snapshots cover only the top-level DOM, references expire, secret detection is heuristic, patches are capped at 1 MiB/64 operations/16 MiB, at most eight browser sessions, and Stop cannot undo already-sent actions, so capability_boundaries is 2. Deduction: trigger precision is weak because automatic naming fires an extra model request on the first message; editing the name opts out, but the default behavior is implicit, so trigger_precision is only 1. Environment fit is documented with a Linux requirement and Tailscale guidance, so environment_fit is 2.
Information architecture is solid: the README covers quick start, production, browser desktop, administrator commands, security and local data, and development, and points to docs/ARCHITECTURE.md, DESIGN.md, and ROADMAP.md, so information_architecture is 2. Install notes are concrete with commands and ports (bun install, bun run dev, 127.0.0.1:5173/3100), so install_notes is 2. Known limitations are unusually thorough (not a sandbox, immediate sudo, non-transactional patches, clipboard limits, Stop limits, session caps), so known_limitations is 3. The full MIT text is present and package.json declares MIT, so license is 3. Deductions: version is only 0.1.0 with no CHANGELOG, so versioning_changelog is 1; naming stability is weak because the README names specific models such as GPT-6 Astra Ultrafast and a service_tier whose availability varies by account, so naming_stability is 1; examples and FAQ are limited to quick start with no FAQ, so examples_and_faq is 1; maintenance responsibility is unstated (no maintainer, no contribution guide, unverified publisher), so maintenance_responsibility is 1.
Output usability is reasonable: a local web UI, mobile experience, file uploads, inline images, downloadable artifacts, and structured browser tools, so output_usability is 2. Marginal value is clear: local-first operation, multi-agent grouping, private browser handoff, and a sudo request flow differentiate it from generic agent frameworks, so marginal_value is 2. Deduction: cost-benefit is unclear because a ChatGPT subscription or OpenAI API key is mandatory with no no-key demo mode, Ultrafast is explicitly higher-priced with separate limits, and users must maintain their own VM, Chrome, VNC, and Tailscale infrastructure, so cost_benefit is only 1.
claim_traceability is 2: most README claims map to specific test files (naming, patch transport, browser tools) and doc paths, though no per-claim citation index is provided. cross_source_corroboration is 2: README, package.json scripts, and test files corroborate each other on the naming flow, patch tooling, and browser tool set. fact_inference_separation is 2: the README clearly distinguishes simulated content (demo data, simulated sudo and website sign-in) from real capabilities, and test comments state that isolated fixtures make no real model requests. Deduction: there is no independent third-party verification and no verified publisher, and this is a static review with no code executed, so scores do not exceed 2.
- Jelly is explicitly not a sandbox: agents execute shell commands and access files with your OS user's permissions and without per-tool approval by default; run it only on loopback or a trusted Tailscale network and never expose it to the public internet.
- apply_patch paths are not sandboxed and filesystem writes are non-transactional, so failures can leave partial changes; back up or use version control before running it against important repositories.
- Sudo commands run immediately as root when your policy permits, with no second confirmation; configure sudoers and askpass carefully.
- Dependencies are only partially pinned (several ^ ranges) with no lockfile, audit, or vulnerability mitigation evidence; pin and audit dependencies before deployment.
- Running requires Bun 1.3.9+, Node 24+, a Linux desktop runtime, system Chrome/Chromium, xclip, /usr/bin/python3, and sudo askpass, with no documented fallback path.
- A ChatGPT subscription or OpenAI API key is mandatory with no no-key demo mode; Ultrafast is explicitly higher-priced with separate limits, so evaluate cost yourself.
- The publisher is unverified, maintenance responsibility and update path are unstated, the version is only 0.1.0, and there is no CHANGELOG; long-term maintenance and upgrade risk is yours.
- This is a static source review with no code executed and no independent testing; all conclusions rest solely on the provided files.