Productivity & Collaboration multi-agent-orchestrationdesktop-wallpaperworkflow-automationtask-managementmodel-routingvoice-commandsmessaging-channelslocal-models

BagIdea Office

A visible, controllable multi-model agent team living on your desktop.

FollowAgents review · FARS-2.1
Use with care
71/ 100 5-point scale 3.6 / 5
1 2 3 4 5 6
Per-dimension scores and reasoning
1Trust18 / 29 · 3.1/5

The evidence shows several permission layers: registered projects, a permission broker, a unified approvals queue, plugin approval, workspace-write isolation by default, read-only container exposure, webhook HMAC, and contents:read CI permissions. Publishing, email, and reply workflows are also described as human-approved. Full marks are withheld because schedules, files, events, and messaging triggers can start unattended work, while agents may initiate meetings and proposals, so not every consequential action is individually confirmed. The README identifies major models, channels, memory features, and API-key purposes, but provides no complete data-flow map, provider-by-provider disclosure, credential-storage design, or retention/deletion policy. Dependency evidence includes locked Rust builds, a platform CI matrix, registry checks, and removal of deprecated packages, but no vulnerability scanning or comprehensive dependency audit, and Actions are tag-pinned rather than commit-pinned. Recovery covers retained skill versions, isolated worktrees, persisted runs, and approval history, but not general reversal of plugins, configuration, or external actions. Inspirations, a contributor, an MIT copyright holder, and community channels are attributed, although publisher identity and maintenance ownership remain unclear.

2Reliability9 / 14 · 3.2/5

Cross-platform Rust builds, Node 20/22 testing, Claude CLI and PATH verification, doctor diagnostics, and refusal to silently downgrade unsupported execution backends adequately address ordinary dependency availability. Abnormal runs now persist explicit reasons, while installer and desktop-connectivity failures receive targeted explanations and suggested fixes, earning full credit for failure messages. Self-consistency is weaker: the root package declares version 0.9.30 and private:true while the README advertises releases through 1.6.1, npm installation, and continuous publishing. CI expressly excludes the meeting test requiring a real Claude environment, and API tests skip when the daemon is absent or an endpoint is from an older build.

3Adaptability16 / 18 · 4.4/5

The README identifies concrete audiences and scenarios spanning development, research, content, support, personal assistance, marketing, mail, schedules, workflows, and multi-agent collaboration. Team templates, plugins, model swapping, remote execution, and multilingual support broaden adaptation. Trigger semantics cover schedules, webhooks, events, files, RSS, and channel keywords, with approval nodes and HMAC verification. Environment evidence includes Windows, macOS, and Linux CI plus Docker, SSH, Ollama, local or remote models, and a platform-aware folder picker. Capability boundaries lose a point because permissions, unavailable features, and provider differences are not enumerated for every advertised model, plugin, and channel; macOS/Linux installation is described as beta, and some features require additional API keys.

4Convention14 / 18 · 3.9/5

The README connects the website, getting-started material, ecosystem design, tool documentation, and changelog, while stable all-caps UI labels coordinate fourteen languages. Installation entry points, npx use, the Claude Code prerequisite, optional keys, and doctor are described, but the supplied material omits the full installation section and uninstall or failed-upgrade recovery instructions. Release notes are unusually detailed and the MIT license text is complete. Version convention is reduced by the conflict between package.json 0.9.30 and README 1.6.1. Limitations appear in beta notices, excluded meeting tests, and platform conditions rather than a consolidated limitations section. Discord, issue/PR references, and an update command provide an update path, but named maintainers, support expectations, a security-reporting route, and release ownership are not clear; unknown publisher identity is not itself treated as suspicious.

5Effectiveness10 / 13 · 3.8/5

The product combines multi-agent state, approvals, tasks, calendars, workflows, notifications, cost tracking, and memory in a desktop-office interface, offering substantial marginal value. Persistent history, task boards, approval queues, and diagnostics make results more actionable. Daily and project caps, 80% warnings, stopping new turns at 100%, agent/project attribution, and local Ollama options improve the cost-benefit case. Deductions remain because output usability and economics are largely README claims: the supplied evidence contains no representative end-to-end deliverables, quality benchmarks, resource measurements, or model-price estimation method. The full experience also requires Claude Code and may require several paid APIs.

6Verifiability4 / 8 · 2.5/5

Many claims are tied to releases, issues or pull requests, test counts, endpoints, and specific failure mechanisms. The supplied CI and test files corroborate platform builds, several API contracts, executable discovery, and AUTO status parsing. Full marks are not supported because most central capabilities appear only as README claims without corresponding implementation files or complete tests, while some API checks skip when the daemon or newer endpoint is unavailable. Marketing language blends observable behavior with interpretations such as agents having souls, growing, or developing a social life, so facts, inference, and aspiration are not cleanly separated. The root-version conflict further weakens claim traceability.

Evidence confidence: Low Reviewed Sep 11, 2026 Reviewed revision cdfc23c6dd62
Before you use it
  • The root package's 0.9.30/private:true metadata conflicts with the README's 1.6.1 and npm publication story; verify the actual release artifact and revision before installing or updating.
  • Before enabling schedule, webhook, file, RSS, or messaging triggers, review callable tools, network destinations, project directories, budgets, and human-approval boundaries individually.
  • The README does not fully specify storage, transmission, retention, or deletion behavior for API keys, conversations, memory, voice, media, and messaging-platform data.
  • CI omits the meeting test that needs a real Claude CLI/model, and some API tests can skip; this static evidence does not establish end-to-end agent behavior.
  • Plugins and self-correcting skills expand runtime capability; inspect proposals, diffs, dependencies, isolation, and recovery paths before approval.
Review evidence [1][2][3][4][5][6][7][8]
See the full review method →

What does this agent do, and when should you use it?

BagIdea Office is a 2.5D multi-agent workspace delivered as a desktop wallpaper, combining a Godot 4 world, a zero-dependency Node.js event daemon, WebView interfaces, and the `bagidea` CLI. Each employee is backed by a real headless Claude Code session that can use assigned tools, receive delegated work, and operate inside registered project folders. Individual employees can use Claude, OpenAI, Gemini, GLM, DeepSeek, Qwen, Kimi, Groq, Ollama, LM Studio, and other supported providers, although Claude Code remains the execution engine for every brain. The product also includes a task board, calendar, executable workflows, approvals, memory, skills, plugins, voice features, and Telegram, Discord, LINE, Slack, WhatsApp, and Messenger channels. It is best suited to people who want a persistent, observable local agent team and are prepared to manage a desktop runtime, permissions, provider connections, and usage budgets.

A request can enter through the desktop chat, a connected messaging channel, the bagidea CLI, or the HTTP API; POST /chat launches a real headless claude -p session with the employee's persona, assigned skills, and allowed tools. The Director can delegate work to employees or parallel ghost clones, while the WebSocket daemon turns session output into events shown by the Godot world and UI as working, meeting, blocked, completed, and cost states. The workflow engine executes graphs node by node, including parallel branches, joins, decisions, delays, notifications, and human approvals, and persists run state across restarts; schedules, webhooks, office events, arriving files, and channel keywords can trigger those graphs. Delegations, meeting actions, jobs, and workflow runs feed a shared task board, while the calendar supports recurrence, reminders, and .ics import and export. Memory is distilled into shared OFFICE.md, per-agent files, and project MEMORY.md files, retrieved through local BM25 search or an optional OpenAI-compatible embeddings endpoint. Plugins can add panels, server routes, triggers, workflow nodes, and agent-callable commands, while the File & Media Toolkit can invoke LibreOffice, pandoc, yt-dlp, ffmpeg, ImageMagick, and jq when those helpers are installed.

  1. A solo developer handling several projects can have the Director divide work among model-specific employees, follow progress on the task board, and request a Codex second opinion.
  2. An independent operator who needs unattended routines can launch persisted workflows from schedules, webhooks, files, or chat keywords while retaining human approval steps.
  3. A small team trying to contain model spend can reserve Claude for difficult decisions, route routine execution to cheaper or local models, and enforce office, employee, and project budgets.
  4. A user working away from the desktop can issue requests, receive notifications, and answer approvals through Telegram, Discord, LINE, Slack, WhatsApp, or Messenger.
  5. A knowledge worker handling PDFs, Office files, presentations, images, or recordings can ask employees to convert files, author documents, transcribe media, and retain relevant findings in searchable memory.

What are this agent's strengths and limitations?

Pros
  • It maps real Claude Code sessions, delegation, parallel ghost work, meetings, and blocked states into a live office instead of presenting only a chat window or static dashboard.
  • Each employee can use a different cloud or local model, with built-in protocol translation, live provider model lists, failover, context rollover, and provider-level cost reporting.
  • Its automation surface joins persisted graph execution, schedules, webhooks, task dependencies, recurring work, calendars, notifications, and cross-channel approvals.
  • Tool permissions, project hooks, proposals, and self-extension pass through human gates, while monetary caps can be applied to the office, individual employees, and projects.
  • Plugins, MCP servers, native skills, file-processing helpers, and local or semantic memory provide several concrete ways to extend the office without changing its core.
Limitations
  • Claude Code CLI remains a hard runtime dependency; multi-provider model routing does not make the core agent portable to an unrelated execution framework.
  • Platform maturity varies: Windows 11 is described as stable, macOS 13+ as beta, and Linux as experimental, with Wayland falling back to a fullscreen window pinned below other windows.
  • The Godot world, Node.js daemon, WebView shell, provider setup, permissions, and optional helper binaries create a larger deployment and troubleshooting footprint than a conventional chat client.
  • Transcription, TTS, realtime calls, image generation, and video introduce extra provider dependencies and charges; non-Claude costs are estimates, and video is described as costing about $2 per clip.
  • AUTO mode, proactive meetings, proposals, and event-triggered workflows increase the scope of unattended activity, so budgets, notifications, registered folders, and approval rules need deliberate configuration.

How do you install or deploy this agent?

Claude Code CLI is required because every employee runs through it, including employees assigned to third-party models. The documented npm entry point is npx bagidea; the project also advertises a one-shot installer for Windows and beta macOS/Linux paths, but the supplied material does not include the installer's complete shell command. After installation, use bagidea start to launch the office, bagidea status to inspect it, and bagidea doctor to diagnose daemon reachability, proxy or firewall interference, execution policy, and Claude Code installation. A Claude login is optional for users who exclusively select third-party providers, but those providers must still be connected in settings. OpenAI and Gemini keys are optional for basic operation and required for relevant voice, realtime calling, image, and other full-experience features.

How do you use this agent?

Run npx bagidea for the documented npm setup or launch path, then start the office with bagidea start. In settings, connect the desired model providers, assign a brain, skills, and tools to each employee, and register only the project folders in which they may work; project-supplied .claude command hooks are held for security approval. Submit the first job from the CEO chat, with bagidea ask or bagidea chat, or through POST /chat; inspect activity and spending with bagidea status, bagidea stats, and the STATS, TASKS, and RUNS panels. For unattended work, create a graph in Workflow Builder and attach a schedule, webhook, office-event, file, or channel-keyword trigger; pending decisions appear in APPROVALS and can also be answered from connected phone channels. To invoke Codex, enable it as a system tool and use bagidea codex "task", POST /codex/exec, a Codex workflow node, or the Director form DELEGATE: codex @ project :: task.

How does this agent compare with similar options?

The project explicitly credits openclaw for the agent-office idea and Hermes for self-learning skills, then describes BagIdea Office as combining those ideas with human-approved execution of real projects and agent-proposed plugins. Unlike systems where each provider supplies the complete agent runtime, BagIdea Office keeps Claude Code as the common tools, skills, and session engine and swaps only the model behind each employee. That produces a consistent execution layer but also makes Claude Code a central dependency.

FAQ

Can I run it without an Anthropic account?
Yes, an employee can use GLM, DeepSeek, Qwen, OpenAI, Gemini, Ollama, or another connected provider without a Claude login. Claude Code CLI must still be installed because it remains the execution engine.
Can agents execute untrusted project code automatically?
Projects must be registered, and project-owned .claude command hooks are displayed with their literal commands and parked for approval. Tool permissions, plugins, proposals, and blocked jobs can also use the central approval queue, though enabling auto-approval or AUTO mode changes the practical risk.
Can the office operate fully offline?
Ollama and LM Studio provide local model paths without API keys, and BM25 memory retrieval runs on-device. Installation, updates, remote channels, cloud models, and the full voice and media feature set still require network services.
How are unattended costs controlled?
You can set a daily office cap, a daily cap per employee, and a lifetime project cap. The office warns at 80% and refuses new turns in that scope at 100%, while allowing already-running turns to finish.
What survives a crash or restart?
The event journal replays state on reconnect, and workflow runs persist so delays and approvals can resume. An agent node interrupted mid-flight is marked failed rather than reported as completed; model overload can use a configured fallback brain, and long threads can be compacted into a continued session.

Compare agents like this one

The same FARS review applied across the shortlist this agent qualifies for.

Related agents