B4.run Agent Framework

Build testable, controlled, deployable TypeScript agent apps on LangGraph.js.

Source repo
cacheplane/b4run
Stars
★ 311
Last updated
today
License
MIT
Primary language
TypeScript

At a glance

How it runs
FrameworkSelf-hosted service
Works with
Portable with changesOpenAI API (Partial support)
Cost
Free software; you pay for model usage
Setup effort
Medium · a few setup steps
You'll need
Node.js 24 or laternpm 11LangGraph.jsShell / CLINetwork accessLocal filesystem
Typical use
A TypeScript team adopting an existing LangGraph.js graph can expose it incrementally as a raw graph route and validate checkpointer behavior per target.
Not a fit if
  • Teams that require a native Python framework
  • Teams seeking a hosted AI platform
Source review
69/100 · Some gaps

What does this agent do, and when should you use it?

B4.run is a TypeScript framework built around LangGraph.js, adding file-system routes, generated types, local development and testing conventions, and persistence primitives for agents and workflows. Developers define an `agent`, `workflow`, `graph`, or `chain` in a route's `src/app/**/index.ts`, with shared tools, route-local tools, and runtime policies available. Tests can replay fixture responses offline; state defaults to SQLite-backed storage, whose persistence across restarts depends on retaining the app root and SQLite files. The default build emits a runnable Node server, a Dockerfile, and LangSmith graph artifacts. It requires Node.js and LangGraph.js; B4.run does not host the app, provision infrastructure, or manage secrets.

Create a TypeScript project with npm create b4-app@latest, then export agent() or another supported entry from a route's index.ts; B4.run discovers the route and generates route, parameter, state, and tool declarations. Put shared tools in src/tools/ or alongside a route, then restrict them with tools.deny or require human approval with tools.approve. Permissions can gate shell commands and file paths outside the workspace, and the optional sandbox isolates workspace tools. Develop with b4 dev and run offline tests with committed fixtures; builds can emit a Node server, Dockerfile, and LangSmith graph artifacts. You must validate the runtime, storage, authentication, and provider boundary for your chosen deployment target.

  1. A TypeScript team adopting an existing LangGraph.js graph can expose it incrementally as a raw graph route and validate checkpointer behavior per target.
  2. A team that needs repeatable local agent tests can replay fixture responses and fail on unmatched interactions instead of silently calling a provider.
  3. Developers building agents with restricted tool access can deny tools, require human approval, and gate operations outside the workspace.
  4. An app that needs state to survive development server restarts can use the default SQLite-backed storage, provided the app root and SQLite files persist.
  5. A TypeScript project preparing a Node deployment can use the default build's server and Dockerfile outputs, then validate its own deployment environment.

How do you install or deploy this agent?

Requires Node.js 24 or later and npm 11. Scaffold and install the starter:

npm create b4-app@latest my-agent
cd my-agent
npm install
npm test

The starter test runs offline with fixtures and needs no API key or model calls. To run the README's live OpenAI navlog example, set OPENAI_API_KEY in server/.env.

How do you use this agent?

Define an agent in a route file, for example:

import { agent } from "@b4run/sdk"

export default agent({
  model: "gpt-5-mini",
  recursionLimit: 100,
  description: "A VFR flight planner for a Cessna 172N…",
  tools: { deny: ["runBash"], approve: [{ tool: "fileFlightPlan", allowAlways: false }] },
  systemPrompt: "You are a VFR flight-planning assistant for a Cessna 172N…",
})

Run the basic starter with npm run dev; set OPENAI_API_KEY for its live model path. The navlog template uses a two-package workspace, with separate commands for the server and Workbench:

npm run dev:server
npm run dev:web

The navlog server listens on port 3002 and the Workbench on port 3010. Its production build and start commands are:

npm run build
npm start

Model credentials depend on the provider; the README also says a local Ollama route needs no provider key.

What are this agent's strengths and limitations?

Pros
  • File-system routing discovers agent and graph entries, and type generation emits route, parameter, state, and tool declarations during typegen and build.
  • Fixture-backed tests offer deterministic offline replay and fail on unmatched interactions instead of silently calling a model provider.
  • Route policies can deny tools or require human approval; permissions and the optional sandbox put boundaries around shell commands, file paths, and workspace tools.
  • The default build turns application source into a Node server, Dockerfile, and LangSmith graph artifacts.
Limitations
  • The framework requires LangGraph.js and targets TypeScript and Node.js; Python projects need a different approach.
  • B4.run is pre-1.0 and its API surface is moving, so teams need to pin versions and track upgrades.
  • B4.run does not host apps, provision infrastructure, or manage secrets; operators must validate storage, authentication, and provider boundaries for their target.
  • Live model calls may require provider credentials and usage charges; the README's live OpenAI navlog path requires OPENAI_API_KEY.

How does this agent compare with similar options?

Compared with raw LangGraph.js, B4.run adds file-system routing, generated types, local development and testing conventions, persistence primitives, and build targets. Teams that do not want those application conventions or require Python can use raw LangGraph.js directly. Existing LangGraph.js graphs can migrate incrementally as graph routes, with checkpointer behavior validated on each deployment target.

Key facts side by side with the most closely related agents.

Agent Source review Form / cost Stars Updated Language Full support on
B4.run Agent Framework This agent 69 · Some gaps FrameworkFree + model costs ★ 311 today TypeScript —
TrueForge 67 · Some gaps CLIFree + model costs ★ 6.1k today TypeScript OpenAI API · Claude API
Sandbox Agent 67 · Some gaps Self-hosted serviceFree + model costs ★ 1.6k 1d ago TypeScript Codex · Claude Code
FastClaw 43 · Major gaps CLIFree + model costs ★ 1.4k today Go OpenAI API · Claude API

How does FollowAgents rate this agent?

FollowAgents source review · FARS-2.1
Some gaps
69/ 100 5-point scale 3.5 / 5
Trust 17/29
Reliability 9/14
Adaptability 15/18
Convention 14/18
Effectiveness 9/13
Verifiability 5/8
Why each dimension lost points
Trust17 / 29 · 2.9/5

The README describes tool denial and approval, approval grants, path and shell permissions, and an optional sandbox. These provide visible controls for least privilege and consequential actions, but the supplied material does not fully describe all runtime permissions or user data flows. The README says users configure provider keys and the platform does not manage secrets; SECURITY.md provides a private vulnerability reporting route. Evidence for detailed sensitive-data handling and comprehensive dependency-risk mitigation is limited. Workflow permissions are narrow, dependency overrides are present, and chart tests cover selected safeguards; rollback and source attribution evidence is limited, so full marks are not justified.

Reliability9 / 14 · 3.2/5

The README describes offline fixture tests that fail on unmatched interactions, alongside build, typecheck, and several CI verification scripts. The workflow explicitly handles Docker skips and failure cases. The package manager and Node version requirements are stated, and dependency locking and overrides are visible. Scores are limited because this static review cannot establish that CI actually passed, and the README requires users to validate runtime, storage, authentication, and provider boundaries for their deployment target.

Adaptability15 / 18 · 4.2/5

The README clearly identifies TypeScript/Node.js and LangGraph.js teams as the audience, explains when raw LangGraph.js or Python may be a better fit, and gives an incremental migration path. Multiple examples cover different tasks. Deductions reflect that route discovery and some triggering behavior are described at a high level, capabilities vary by build target, and users must validate their deployment environment.

Convention14 / 18 · 3.9/5

The README organizes quickstart, architecture, examples, deployment, maturity, and support information. It provides installation commands, an MIT license, migration and upgrade guides, and a release-notes entry point. It clearly identifies the pre-1.0 status and moving API surface, and gives support and security reporting routes. The supplied material lacks concrete changelog content, a FAQ, and fuller maintainer identity and responsibility commitments; examples and documentation do not cover every case.

Effectiveness9 / 13 · 3.5/5

The framework combines file-system routes, type generation, fixture replay, persistence primitives, and build artifacts into a usable workflow. Examples show approval, tools, and subagents. The README also states that it does not host applications, provision infrastructure, or manage secrets. Deductions reflect the lack of quantitative support for the speed claim in the supplied material and that benefits depend on adopting the framework conventions and validating deployment independently.

Verifiability5 / 8 · 3.1/5

README claims about features and boundaries can be compared with package scripts, CI workflows, chart tests, and the security policy; the supplied tests contain concrete assertions and failure handling. Deductions reflect the absence of execution results, revision-level release records, and complete implementation files in the supplied evidence. CI configuration does not prove results for this revision, and performance and some security claims need more direct evidence.

Risks and how to mitigate them
  • The project is pre-1.0 and its API is moving; pin versions and consult the upgrade guidance.
  • The project provides a framework and build artifacts but does not host applications or manage secrets. Validate authentication, storage, provider, and deployment boundaries for your target environment.
  • The supplied material does not quantify the performance claim. This is a low-confidence static review; no code was run and CI results were not verified.
Evidence confidence: Low Reviewed Oct 09, 2026 Reviewed revision 24bc5c7c8464 New commits since this review; the score may not cover them
See the full review method →

FAQ

Do I need to pay for a model API to try B4.run?
No API key is needed for the fixture-backed starter test: npm test runs offline. Live credentials depend on the provider; the README's OpenAI navlog path requires OPENAI_API_KEY, while a local Ollama route needs no provider key.
Does B4.run host my app or provision production infrastructure?
No. The default build emits a Node server, Dockerfile, and LangSmith graph artifacts, but you operate the app and manage its infrastructure and secrets.
Can tools require approval or restrict file and shell access?
Yes. tools.approve can pause a tool call for a human response, with each answer spending a single-use approval grant. Permissions can gate shell commands and paths outside the workspace, and the sandbox is opt-in.
Will SQLite state survive a restart?
It can survive a b4 dev restart when the app root and SQLite files persist. Validate storage and checkpointer behavior for each deployment target.
Is the API stable at 1.0?
No. B4.run is pre-1.0 and its API surface is moving; the README recommends pinning versions and reading release notes and the upgrade guide.
View on GitHub ↗ Install ↓

Compare agents like this one

The same FARS review applied across the shortlist this agent qualifies for.

Related agents