SurfSense Open-Web Research
Give agents structured live-web research through one REST API or MCP server, with cited outputs and a searchable knowledge base.
The README identifies cloud and self-hosted modes, API keys, OAuth connectors, external sources, and automation write targets, while the test setup forces an isolated test database. It does not document least-privilege scopes, per-action confirmation, retention or deletion, secret storage, privacy boundaries, or connector-level transfers. Automations may write to Notion, Slack, Linear, and Jira without shown preview, approval, or impact controls. Daily Watchtower updates can be disabled, but pinning, rollback, and recovery are not described. CI runs detect-secrets and Bandit and uses lockfiles and versioned Actions, supporting ordinary dependency hygiene, but scanning is limited to PR changes and no vulnerability-response or supply-chain policy is shown. Cited answers are claimed without implementation evidence or source-integrity guarantees in the supplied files.
Separate unit and integration CI jobs, a PostgreSQL/pgvector service, aggregate gates, and backend/frontend quality jobs support basic consistency and dependency availability. No results, coverage data, or Agent behavior contract are supplied, and the gates treat cancelled or skipped jobs as non-failures. The README claims fast, deterministic connectors and no billing for failed calls while also stating that the product is not production-ready; the supplied source cannot resolve those claims. Failure-message evidence is limited to generic CI output, with no demonstrated structured user-facing or tool-call errors and recovery guidance.
The README thoroughly identifies developers, researchers, teams, self-hosters, and Claude/Cursor-style Agent users, with REST, MCP, cloud, desktop, and broad model-provider options. It distinguishes the product from browser agents, generic scraping APIs, search APIs, and scraper marketplaces, and acknowledges weaker video/podcast capabilities and production immaturity. Deductions reflect the absence of tool schemas, trigger conditions, scope constraints, rate-limit behavior, or ambiguity handling; scheduled and event triggers are described only at feature level.
The README has a table of contents, coherent feature sections, comparison tables, roadmap and contribution paths, multilingual navigation, and copyable REST/MCP examples. Installation covers Docker prerequisites, Linux/macOS and Windows commands, a Watchtower opt-out, and a documentation path, but executes remote scripts directly and does not include manual verification, upgrade, or troubleshooting details in the evidence. Naming and interface examples are reasonably consistent, though there is no compatibility policy or real FAQ. Production immaturity and weaker capabilities are disclosed. The root license clearly separates Apache-2.0 content from a proprietary-directory BSL 1.1 exception, but the subsidiary license is not supplied and the broad open-source wording needs that qualification. Changelog, releases, and roadmap routes exist without a shown versioning policy. Maintenance channels exist, but the publisher identity and concrete ownership responsibilities remain unclear.
The material describes structured JSON, cited answers, reports, many export formats, podcasts, slides, alerts, and collaboration, giving outputs broad practical utility. A unified REST/MCP surface and integrated knowledge base plausibly add value over building separate site scrapers. Cost information covers per-successful-item/page billing, unbilled failures, disabled billing for self-hosting, and a free credit. Full marks are withheld because no actual output artifacts, quality metrics, latency measurements, detailed credit prices, or substantiated comparison methodology are provided.
The README points to connector pages, documentation, a changelog, releases, a roadmap, and named dependencies, and it claims citation-bearing answers; CI definitions and fixtures also corroborate that a test structure exists. Most evidence supplied is nevertheless a marketing-oriented README, and many numerical, competitive, performance, determinism, and platform-coverage claims are not traced into the provided code or tests. The comparisons mix facts, evaluations, and inferences; a few caveats are explicit, but systematic sourcing, cross-checking, and fact-versus-inference labels are absent.
- The project explicitly says it is not production-ready; independently review authentication, tenant isolation, retention, deletion, and audit controls before using sensitive, regulated, or business-critical data.
- Automations can write to Notion, Slack, Linear, and Jira, but the evidence shows no per-action confirmation, preview, idempotency, compensation, or rollback controls.
- The one-command installers execute remote scripts and enable daily automatic updates by default; pin revisions, inspect scripts and images, and establish a tested rollback path before deployment.
- Scraping social, search, maps, employment, and commerce sites may be constrained by platform terms, privacy rules, and local law; the supplied material does not define compliance boundaries.
- Licensing mixes Apache-2.0 with BSL 1.1 for a designated directory; inspect the omitted subsidiary license before distribution, hosted use, or commercialization.
- Performance, determinism, competitive comparisons, and citation quality are principally README claims and were neither executed nor independently verified in this static assessment.
What does this agent do, and when should you use it?
SurfSense is a self-hostable open-web research platform and knowledge workspace positioned as an open-source NotebookLM alternative. Its REST API and MCP server expose connectors for Reddit, YouTube, Instagram, TikTok, Google Search, Google Maps, Indeed, Amazon, Walmart, and arbitrary web pages. Findings can flow into a built-in knowledge base for hybrid semantic and full-text retrieval, cited answers, reports, podcasts, presentations, and video overviews. Scheduled and event-triggered agent runs can also write results back to Notion, Slack, Linear, and Jira. Teams can use the metered cloud service or deploy the stack with Docker; self-hosting disables SurfSense billing but requires their own infrastructure and model keys. The project is actively developed and explicitly described as not yet production-ready.
A user or agent starts from the natural-language workspace, a REST endpoint, or the MCP server. For example, Reddit research is submitted to POST /workspaces/$WORKSPACE_ID/scrapers/reddit/scrape, while MCP exposes native tools such as surfsense_reddit_scrape and surfsense_google_search. The connectors retrieve public posts, comments, video transcripts, search results, place reviews, job listings, and product records; Web Crawl turns arbitrary open-web pages into structured content. The agent harness adds retries, structured output, and credit metering, then can turn findings into cited briefs or alerts. Retrieved material can enter the knowledge base alongside uploaded files, Google Drive, OneDrive, Dropbox, and synchronized local folders for hybrid search and cited question answering. The Deliverable studio exports reports as PDF, DOCX, HTML, LaTeX, EPUB, ODT, or plain text and can also create podcasts, editable slide decks, video overviews, and AI images. Automations run full agent turns on schedules or events and can write their results to Notion, Slack, Linear, or Jira.
- A brand researcher tracking what Reddit posts and comments have said about a product since launch and delivering a cited weekly brief.
- A local-market analyst comparing recurring complaints across Google Maps ratings and reviews for multiple locations.
- A recruiting or market-intelligence team monitoring public Indeed listings, salaries, and full job descriptions and sending updates into collaboration systems.
- A research group combining PDFs, office files, cloud-drive documents, and live web findings in one searchable, citation-backed workspace.
- A developer exposing Reddit, YouTube, search, and crawling connectors as native MCP tools to Claude, Cursor, or a custom agent framework.
- A privacy-conscious organization self-hosting the platform with Docker and using local models through Ollama or vLLM.
What are this agent's strengths and limitations?
- One REST API and MCP server cover social networks, search, maps, jobs, commerce, and generic pages while returning structured posts, comments, transcripts, and reviews.
- The system extends beyond scraping with retries, structured output, metering, a citation-backed knowledge base, deliverable generation, and automation.
- Docker self-hosting is documented, with multi-provider and local-model paths through the OpenAI specification, LiteLLM, vLLM, and Ollama.
- Live-web findings can be searched alongside uploaded files, cloud drives, and local folders, then exported into several document and media formats.
- Collaboration includes real-time chats, comments, mentions, and Owner, Admin, Editor, and Viewer roles.
- The project explicitly says it is not yet production-ready, creating stability, upgrade, and operational risk for critical deployments.
- Hosted connectors charge per returned item or successfully fetched page, so high-volume monitoring can create ongoing usage costs.
- Self-hosting removes SurfSense billing but shifts Docker operations, hardware capacity, maintenance, and model-key management to the adopter.
- The repository license is reported as NOASSERTION, and no definitive license terms appear in the supplied material; organizations must verify redistribution and deployment rights.
- The project acknowledges that NotebookLM currently produces better video and podcast outputs.
How do you install or deploy this agent?
Self-hosting requires Docker Desktop to be installed and running. On Linux or macOS, run:
curl -fsSL https://raw.githubusercontent.com/MODSetter/SurfSense/main/docker/scripts/install.sh | bashOn Windows PowerShell, run:
irm https://raw.githubusercontent.com/MODSetter/SurfSense/main/docker/scripts/install.ps1 | iexThe installer configures Watchtower for daily automatic updates. Add --no-watchtower to skip it. Billing is disabled in self-hosted installations, but agent execution still requires model keys supplied by the operator. Docker Compose and manual deployment are mentioned, although their complete commands are not included in the supplied material.
How do you use this agent?
For a REST connector call, obtain a SurfSense API key and workspace ID and set the API base URL. A first Reddit request is:
curl -X POST "$SURFSENSE_API_URL/workspaces/$WORKSPACE_ID/scrapers/reddit/scrape" \
-H "Authorization: Bearer $SURFSENSE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"search_queries":["your brand"],"community":"SaaS","sort":"top","time_filter":"week"}'To expose the connectors over MCP, add this configuration to Claude, Cursor, or a custom framework:
{"mcpServers":{"surfsense":{"url":"https://mcp.surfsense.com/mcp","headers":{"Authorization":"Bearer ${SURFSENSE_API_KEY}"}}}}
The agent can then invoke each connector as a native tool. For the hosted product, sign in at surfsense.com and request live web data in natural language; new accounts receive $5 in free credit with no subscription.
How does this agent compare with similar options?
Against browser agents such as Browserbase and Browser Use, SurfSense uses structured connector calls for read-only retrieval instead of spending model time navigating each page; browser agents remain better suited to login, clicking, and form-filling tasks. Against Firecrawl, SurfSense emphasizes platform-native records such as posts, comments, transcripts, and reviews rather than a generic Markdown page. Unlike search APIs such as Exa, Tavily, and Parallel, it can retrieve Reddit comments, TikTok reactions, YouTube transcripts, and Google Maps reviews rather than only finding indexed pages. Compared with Apify's marketplace of actors with varying schemas and prices, SurfSense provides one typed API, one MCP server, an agent harness, and a research workspace. Compared with Google NotebookLM, SurfSense offers live connectors, MCP, self-hosting, multi-model support, local models, automations, and real-time chat, while conceding that NotebookLM currently has stronger video and podcast generation.