Opensteer

A compact Python interface for agents to operate local and cloud browsers.

Stars
★ 203
Last updated
2mo ago
License
MIT
Primary language
Python

At a glance

How it runs
CLILibrary / SDKFramework
Works with
Universal · cross-platform
Cost
Free, no paid service needed
Setup effort
Low · running in minutes
You'll need
Chrome or EdgeOPENSTEER_API_KEY (for Opensteer Cloud)Shell / CLINetwork accessLocal filesystem
Typical use
Agent developers who need Chrome or Edge navigation, clicking, typing, and page inspection from Python commands.
Not a fit if
  • Teams that cannot allow agents to execute local Python or JavaScript
  • Users expecting Opensteer Cloud access without an API key
  • Workflows that require read operations to provision cloud browsers automatically

What does this agent do, and when should you use it?

Opensteer is a browser automation framework that gives AI agents a small Python command surface and a set of importable helpers. It handles CDP transport, local and cloud attachment, named sessions, tabs, screenshots, JavaScript execution, clicks, typing, and scrolling. Users can run short snippets through `opensteer -c` or import functions from `opensteer.helpers` into reusable Python code. A project directory acts as a specialized harness, supplying `AGENTS.md`, selectors, data, API clients, and custom functions for a particular workflow. Opensteer can attach to local Chrome or Edge and, when `OPENSTEER_API_KEY` is configured, to Opensteer Cloud browsers with explicit session and transport controls.

An automation run begins with a Python snippet passed to opensteer -c. The code can create, select, or enumerate browsers through open_browser(), use_browser(), and list_browsers(), then open pages with new_tab() and wait using wait_for_load(). It can inspect a page with page_info(), navigate with goto_url(), click with click_at_xy(), enter content through type_text(), scroll, manage tabs, capture screenshots, execute JavaScript through js(), or invoke raw CDP methods. When several browsers are active, the caller retains each returned browser handle and distinguishes sessions with readable labels. Workflow-specific modules such as tools.py and selectors.py are imported from the working directory, allowing a command to produce page information, screenshots, or custom function results. Local App execution can request auto, inapp, local, or cloud, while Cloud execution accepts only auto and cloud.

  1. Agent developers who need Chrome or Edge navigation, clicking, typing, and page inspection from Python commands.
  2. Support teams building a dedicated harness that wraps customer-page actions in functions inside tools.py.
  3. QA teams that need JavaScript evaluation, screenshots, and raw CDP access while investigating web behavior.
  4. Developers running concurrent research, work, or site-specific browser contexts through named sessions and explicit handles.
  5. Sales, data-entry, or research teams that want separate selectors and workflow data while sharing one browser-control layer.

How do you install or deploy this agent?

Run the documented installer from a shell:

curl -fsSL https://opensteer.com/install.sh | sh

Verify the installation and attempt the first page inspection:

opensteer --doctor
opensteer -c "print(page_info())"

If local attachment is not ready, start the setup flow. It guides Chrome or Edge through enabling remote debugging for the current profile:

opensteer --setup

OPENSTEER_API_KEY is required only for Opensteer Cloud browser access. The source does not provide a specific shell command for assigning the key.

How do you use this agent?

Open a page, wait for it to load, and print its page information:

opensteer -c "new_tab('https://example.com'); wait_for_load(); print(page_info())"

Import browser helpers into reusable Python code:

from opensteer.helpers import goto_url, js, click_at_xy, type_text, wait_for_load

Keep the returned handle when operating more than one browser:

opensteer -c "linkedin = open_browser(label='linkedin', profile='fresh'); linkedin.new_tab('https://linkedin.com')"

After configuring OPENSTEER_API_KEY, create or attach to a cloud browser:

opensteer -c "open_browser(label='research', profile='fresh'); new_tab('https://example.com')"
opensteer -c "browser = open_browser(label='work', profileId='bp_...'); print(browser.page_info())"

List and select existing browsers without creating another one:

opensteer -c "print(list_browsers())"
opensteer -c "research = use_browser('research'); research.new_tab('https://example.com')"

A locally executing agent can explicitly request a cloud browser:

opensteer -c "open_browser(label='research', profile='fresh', location='cloud')"

A specialized harness can define a local workflow function and import it directly from the CLI:

# tools.py
from opensteer.helpers import goto_url, js, type_text, wait_for_load


def open_customer(customer_id):
    goto_url(f"https://support.example.com/customers/{customer_id}")
    wait_for_load()
    return js("document.title")
cd support-harness
opensteer -c "from tools import open_customer; print(open_customer('cus_123'))"

What are this agent's strengths and limitations?

Pros
  • One helper surface covers CDP, local and cloud attachment, sessions, screenshots, JavaScript, and common page interactions, reducing the need for a custom browser bridge.
  • The opensteer -c interface suits short operations, while opensteer.helpers supports reusable Python modules as workflows mature.
  • Browser infrastructure is separated from workflow assets: Opensteer owns transport and generic interaction, while each project retains its selectors, data, instructions, and business functions.
  • Named browsers and explicit handles support concurrent browser contexts, with mode controlling whether a logical session is reused or intentionally created.
Limitations
  • Local attachment depends on Chrome or Edge remote-debugging setup and may require an additional opensteer --setup step.
  • Cloud browsers require OPENSTEER_API_KEY, and Cloud execution accepts only auto and cloud, rejecting device-only locations.
  • Domain-specific automation is not supplied out of the box; adopters must maintain their own instructions, selectors, data files, and Python functions.
  • The interface can execute Python, in-page JavaScript, and raw CDP calls, so adopters must define an appropriate security boundary for agent execution.
  • The source documents no native adapter for ChatGPT, Codex, Claude, or their APIs.

How does this agent compare with similar options?

Key facts side by side with the most closely related agents.

Agent Source review Form / cost Stars Updated Language Full support on
Opensteer This agent 49 · Major gaps CLIFree ★ 203 2mo ago Python —
Obscura Headless Browser 75 · Good CLIFree ★ 28k 2d ago Rust Claude.ai
Open Browser 47 · Major gaps CLIFree + model costs ★ 9.6k 6mo ago TypeScript ChatGPT · Claude.ai · OpenAI API · Claude API
Browser4 Agentic Browser 47 · Major gaps CLIFree + model costs ★ 1.1k 2d ago Kotlin OpenAI API

How does FollowAgents rate this agent?

FollowAgents source review · FARS-2.1
Major gaps
49/ 100 5-point scale 2.5 / 5
Trust 8/29
Reliability 8/14
Adaptability 12/18
Convention 10/18
Effectiveness 7/13
Verifiability 4/8
Why each dimension lost points
Trust8 / 29 · 1.4/5

Transport location and session creation are explicit in some paths, and read helpers reportedly avoid implicit provisioning in strict cloud-agent environments, giving limited support for least privilege. However, the framework exposes arbitrary Python, JavaScript, raw CDP, clicks, and typing without granular permissions, domain restrictions, or pre-action confirmation. The README distinguishes local and cloud browsers and mentions API-key configuration, but it does not fully document the flow, storage, retention, or telemetry handling of page content, screenshots, credentials, and other sensitive data, nor provide redaction or secret-management guidance. Runtime dependencies are exactly pinned and the publishing workflow uses narrow permissions plus OIDC, but no vulnerability scanning, dependency audit, or update policy is shown. External write capabilities are disclosed but lack effect classification, confirmation gates, and rollback. The MIT file supplies project-level attribution, while publisher identity remains unverifiable from the supplied material.

Reliability8 / 14 · 2.9/5

The README, package metadata, CLI entry point, and version information are broadly consistent and describe a coherent product, but no implementation code or test results are supplied, preventing full credit. Dependency names and versions, the Python requirement, and build backend are declared, providing an ordinary installation basis; the curl-delivered installer and service availability cannot be verified from these files. `--doctor`, `--setup`, and rejection of unsupported locations provide basic failure paths, but there is no error catalog, detailed recovery guidance, or representative failure-message documentation.

Adaptability12 / 18 · 3.3/5

The material identifies browser-control agents and research, sales, QA, support, and data-entry harnesses as scenarios, with a clear project-specific extension pattern; coverage is useful but not a complete operational guide. Browser versus workflow responsibilities, local versus cloud transport, and supported locations are distinguished, although arbitrary Python, JavaScript, and CDP remain broadly unconstrained. Actions are triggered through explicit CLI or Python calls, `reuse` is the stated default, and `new` intentionally creates a session; finer action policies and ambiguity handling are absent. Python 3.11, local Chrome/Edge, cloud credentials, and setup commands are documented, but an OS compatibility matrix, headless requirements, and network prerequisites are missing.

Convention10 / 18 · 2.8/5

The README is sensibly organized around installation, capabilities, specialized harnesses, ownership boundaries, and verification, but lacks a complete API reference and troubleshooting section. Installation is only a remote script piped to a shell, without integrity verification, uninstall guidance, virtual-environment advice, or platform prerequisites. Package, CLI, entry point, and example API naming are consistent and a version is declared, but there is no compatibility or deprecation policy. Examples cover common browser and extension workflows, while no FAQ is present. Some capability boundaries are documented, but known limitations, safety constraints, and cloud restrictions are not systematically collected. The complete MIT text justifies full license credit. A semantic-looking version and release-tag validation exist, but there is no changelog or migration guidance. A copyright holder and release workflow are visible, while named maintainers, support channels, ownership responsibilities, and update commitments are unspecified.

Effectiveness7 / 13 · 2.7/5

Short Python commands, importable helpers, named browsers, and the harness pattern support directly usable automation, and verification commands improve usability; structured output conventions, end-to-end outcome examples, and recovery behavior are missing. Combining local/cloud attachment, sessions, screenshots, and interaction helpers plausibly adds value over building a CDP bridge from scratch, but no comparative evidence or implementation detail is supplied. The command surface appears compact and reusable, yet cloud pricing, resource use, latency, installer risk, and maintenance cost are undisclosed, limiting the cost-benefit score.

Verifiability4 / 8 · 2.5/5

Most capability claims are illustrated by commands, but implementation files, API documentation, and tests are absent, so claims cannot be traced comprehensively. Package identity, Python requirements, and the CLI align across the README and pyproject, while the workflow corroborates version-aware publishing; cloud and browser behavior lack independent source corroboration. Product claims are generally stated as facts without distinguishing design intent, constraints, assumptions, or verified observations, so fact-inference separation is limited.

Risks and how to mitigate them
  • Not found in source: confirmation before actingTurn on (or add) a confirmation step before it acts, and try it in a sandbox or test environment before real data.
  • Not found in source: rollback or recovery pathBack up first, or work on a git branch or snapshot, so its changes can be undone.
  • Installation pipes a network-fetched script directly into a shell; pin and inspect the installer or use a verifiable package path before deployment.
  • The framework can execute arbitrary Python, page JavaScript, raw CDP, clicks, and typing; run it with an isolated browser profile and low privileges, and add human confirmation for consequential actions.
  • Do not assume safe storage, transmission, or retention of page data, screenshots, browser profiles, or `OPENSTEER_API_KEY` without additional data-handling documentation.
  • No rollback mechanism, systematic limitations document, changelog, vulnerability-scanning evidence, or comprehensive failure-recovery guide is supplied.
Evidence confidence: Low Reviewed Oct 05, 2026 Reviewed revision e8ebf8c3ce00
See the full review method →

FAQ

Is Opensteer Cloud mandatory?
No. Opensteer can attach to local Chrome or Edge. Cloud browsers are an additional path available when OPENSTEER_API_KEY is configured.
Will a read helper automatically create a browser?
Not in strict cloud-agent environments. Read helpers do not provision browsers implicitly there; call open_browser() first or select an existing browser with use_browser().
What is the difference between mode='reuse' and mode='new'?
reuse is the safe default for reusing a logical session. new deliberately creates another logical session. Labels are human-readable, while cloud session IDs and run scope are platform-owned.
How should separate sites or business workflows be organized?
Create a dedicated harness for each workflow, keeping its Markdown instructions, selectors, data, API clients, and Python functions in that project directory while sharing Opensteer's browser primitives.
Is hosted browser pricing documented?
No. The repository is MIT-licensed and supports local browsers, but the supplied source does not state pricing for Opensteer Cloud.
View on GitHub ↗ Install ↓

Related agents