B4.run Agent Framework
Build testable, controlled, deployable TypeScript agent apps on LangGraph.js.
- Source repo
- cacheplane/b4run
- Stars
- ★ 311
- Last updated
- today
- License
- MIT
- Primary language
- TypeScript
- FA score
- 69/100 · Some gaps
At a glance
- How it runs
- Works with
- Portable with changesOpenAI API (Partial support)
- Cost
- Free software; you pay for model usage
- Setup effort
- Medium · a few setup steps
- You'll need
- Typical use
- A TypeScript team adopting an existing LangGraph.js graph can expose it incrementally as a raw
graphroute and validate checkpointer behavior per target. - Not a fit if
- Teams that require a native Python framework
- Teams seeking a hosted AI platform
- Source review
- 69/100 · Some gaps
What does this agent do, and when should you use it?
B4.run is a TypeScript framework built around LangGraph.js, adding file-system routes, generated types, local development and testing conventions, and persistence primitives for agents and workflows. Developers define an `agent`, `workflow`, `graph`, or `chain` in a route's `src/app/**/index.ts`, with shared tools, route-local tools, and runtime policies available. Tests can replay fixture responses offline; state defaults to SQLite-backed storage, whose persistence across restarts depends on retaining the app root and SQLite files. The default build emits a runnable Node server, a Dockerfile, and LangSmith graph artifacts. It requires Node.js and LangGraph.js; B4.run does not host the app, provision infrastructure, or manage secrets.
Create a TypeScript project with npm create b4-app@latest, then export agent() or another supported entry from a route's index.ts; B4.run discovers the route and generates route, parameter, state, and tool declarations. Put shared tools in src/tools/ or alongside a route, then restrict them with tools.deny or require human approval with tools.approve. Permissions can gate shell commands and file paths outside the workspace, and the optional sandbox isolates workspace tools. Develop with b4 dev and run offline tests with committed fixtures; builds can emit a Node server, Dockerfile, and LangSmith graph artifacts. You must validate the runtime, storage, authentication, and provider boundary for your chosen deployment target.
- A TypeScript team adopting an existing LangGraph.js graph can expose it incrementally as a raw
graphroute and validate checkpointer behavior per target. - A team that needs repeatable local agent tests can replay fixture responses and fail on unmatched interactions instead of silently calling a provider.
- Developers building agents with restricted tool access can deny tools, require human approval, and gate operations outside the workspace.
- An app that needs state to survive development server restarts can use the default SQLite-backed storage, provided the app root and SQLite files persist.
- A TypeScript project preparing a Node deployment can use the default build's server and Dockerfile outputs, then validate its own deployment environment.
How do you install or deploy this agent?
Requires Node.js 24 or later and npm 11. Scaffold and install the starter:
npm create b4-app@latest my-agent
cd my-agent
npm install
npm testThe starter test runs offline with fixtures and needs no API key or model calls. To run the README's live OpenAI navlog example, set OPENAI_API_KEY in server/.env.
How do you use this agent?
Define an agent in a route file, for example:
import { agent } from "@b4run/sdk"
export default agent({
model: "gpt-5-mini",
recursionLimit: 100,
description: "A VFR flight planner for a Cessna 172N…",
tools: { deny: ["runBash"], approve: [{ tool: "fileFlightPlan", allowAlways: false }] },
systemPrompt: "You are a VFR flight-planning assistant for a Cessna 172N…",
})Run the basic starter with npm run dev; set OPENAI_API_KEY for its live model path. The navlog template uses a two-package workspace, with separate commands for the server and Workbench:
npm run dev:server
npm run dev:webThe navlog server listens on port 3002 and the Workbench on port 3010. Its production build and start commands are:
npm run build
npm startModel credentials depend on the provider; the README also says a local Ollama route needs no provider key.
What are this agent's strengths and limitations?
- File-system routing discovers agent and graph entries, and type generation emits route, parameter, state, and tool declarations during typegen and build.
- Fixture-backed tests offer deterministic offline replay and fail on unmatched interactions instead of silently calling a model provider.
- Route policies can deny tools or require human approval; permissions and the optional sandbox put boundaries around shell commands, file paths, and workspace tools.
- The default build turns application source into a Node server, Dockerfile, and LangSmith graph artifacts.
- The framework requires LangGraph.js and targets TypeScript and Node.js; Python projects need a different approach.
- B4.run is pre-1.0 and its API surface is moving, so teams need to pin versions and track upgrades.
- B4.run does not host apps, provision infrastructure, or manage secrets; operators must validate storage, authentication, and provider boundaries for their target.
- Live model calls may require provider credentials and usage charges; the README's live OpenAI navlog path requires
OPENAI_API_KEY.
How does this agent compare with similar options?
Compared with raw LangGraph.js, B4.run adds file-system routing, generated types, local development and testing conventions, persistence primitives, and build targets. Teams that do not want those application conventions or require Python can use raw LangGraph.js directly. Existing LangGraph.js graphs can migrate incrementally as graph routes, with checkpointer behavior validated on each deployment target.
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| B4.run Agent Framework This agent | 69 · Some gaps | FrameworkFree + model costs | ★ 311 | today | TypeScript | — |
| TrueForge | 67 · Some gaps | CLIFree + model costs | ★ 6.1k | today | TypeScript | OpenAI API · Claude API |
| Sandbox Agent | 67 · Some gaps | Self-hosted serviceFree + model costs | ★ 1.6k | 1d ago | TypeScript | Codex · Claude Code |
| FastClaw | 43 · Major gaps | CLIFree + model costs | ★ 1.4k | today | Go | OpenAI API · Claude API |
How does FollowAgents rate this agent?
Why each dimension lost points
The README describes tool denial and approval, approval grants, path and shell permissions, and an optional sandbox. These provide visible controls for least privilege and consequential actions, but the supplied material does not fully describe all runtime permissions or user data flows. The README says users configure provider keys and the platform does not manage secrets; SECURITY.md provides a private vulnerability reporting route. Evidence for detailed sensitive-data handling and comprehensive dependency-risk mitigation is limited. Workflow permissions are narrow, dependency overrides are present, and chart tests cover selected safeguards; rollback and source attribution evidence is limited, so full marks are not justified.
The README describes offline fixture tests that fail on unmatched interactions, alongside build, typecheck, and several CI verification scripts. The workflow explicitly handles Docker skips and failure cases. The package manager and Node version requirements are stated, and dependency locking and overrides are visible. Scores are limited because this static review cannot establish that CI actually passed, and the README requires users to validate runtime, storage, authentication, and provider boundaries for their deployment target.
The README clearly identifies TypeScript/Node.js and LangGraph.js teams as the audience, explains when raw LangGraph.js or Python may be a better fit, and gives an incremental migration path. Multiple examples cover different tasks. Deductions reflect that route discovery and some triggering behavior are described at a high level, capabilities vary by build target, and users must validate their deployment environment.
The README organizes quickstart, architecture, examples, deployment, maturity, and support information. It provides installation commands, an MIT license, migration and upgrade guides, and a release-notes entry point. It clearly identifies the pre-1.0 status and moving API surface, and gives support and security reporting routes. The supplied material lacks concrete changelog content, a FAQ, and fuller maintainer identity and responsibility commitments; examples and documentation do not cover every case.
The framework combines file-system routes, type generation, fixture replay, persistence primitives, and build artifacts into a usable workflow. Examples show approval, tools, and subagents. The README also states that it does not host applications, provision infrastructure, or manage secrets. Deductions reflect the lack of quantitative support for the speed claim in the supplied material and that benefits depend on adopting the framework conventions and validating deployment independently.
README claims about features and boundaries can be compared with package scripts, CI workflows, chart tests, and the security policy; the supplied tests contain concrete assertions and failure handling. Deductions reflect the absence of execution results, revision-level release records, and complete implementation files in the supplied evidence. CI configuration does not prove results for this revision, and performance and some security claims need more direct evidence.
- The project is pre-1.0 and its API is moving; pin versions and consult the upgrade guidance.
- The project provides a framework and build artifacts but does not host applications or manage secrets. Validate authentication, storage, provider, and deployment boundaries for your target environment.
- The supplied material does not quantify the performance claim. This is a low-confidence static review; no code was run and CI results were not verified.
FAQ
Do I need to pay for a model API to try B4.run?
npm test runs offline. Live credentials depend on the provider; the README's OpenAI navlog path requires OPENAI_API_KEY, while a local Ollama route needs no provider key.Does B4.run host my app or provision production infrastructure?
Can tools require approval or restrict file and shell access?
tools.approve can pause a tool call for a human response, with each answer spending a single-use approval grant. Permissions can gate shell commands and paths outside the workspace, and the sandbox is opt-in.Will SQLite state survive a restart?
b4 dev restart when the app root and SQLite files persist. Validate storage and checkpointer behavior for each deployment target.