Aurora

Open-source AI-agent incident management and root cause analysis: it investigates your stack automatically when an alert fires and delivers a structured RCA before you open your laptop.

Source repo
Arvo-AI/aurora
Stars
★ 426
Last updated
today
License
Apache-2.0
Primary language
Python

At a glance

How it runs
Self-hosted serviceWeb appMCP server
Works with
Universal · cross-platformOpenAI API · Claude API
Cost
Free software; you pay for model usage
Setup effort
High · needs real infrastructure
You'll need
DockerKubernetes (for sandboxed agent execution and Helm deployment)HashiCorp VaultPostgreSQLRedisMemgraphLLM API key (OpenRouter, OpenAI, etc.)Shell / CLINetwork accessLocal filesystemMCP Server
Typical use
An SRE team paged by PagerDuty or Datadog at 3 AM wants the alert auto-triaged and investigated in the background instead of waking an engineer for a 30-60 minute manual investigation.
Not a fit if
  • Teams that cannot self-host Docker/Kubernetes infrastructure
  • Organizations unwilling to grant LLM agents access to cloud and cluster credentials
  • Users who want a vendor-hosted SaaS that works out of the box

What does this agent do, and when should you use it?

Aurora is an open-source, AI-agent-powered incident management platform from Arvo-AI, built for SRE and DevOps teams. When an alert arrives, LangGraph agents dynamically pick from 30+ tools and run kubectl, aws, az, and gcloud inside sandboxed Kubernetes pods, querying logs, checking deployments, and correlating data across AWS, Azure, GCP, and Kubernetes to produce a structured root cause analysis. The system consists of a Python/Flask backend with Celery workers, a Next.js frontend, Memgraph for the infrastructure knowledge graph, Vault for secrets, and PostgreSQL/Redis storage, deployed via Docker Compose or Helm entirely on your own infrastructure with zero telemetry. It also auto-generates postmortems exportable to Confluence, Notion, or SharePoint, traces blast radius across services through a dependency graph, and can generate pull requests with remediation fixes.

Aurora ingests alerts from PagerDuty, Datadog, Grafana, New Relic, OpsGenie, incident.io, and CloudWatch alarm webhooks; every alert auto-triggers a background investigation. LangGraph agents select from 30+ tools and run kubectl, aws, az, and gcloud in isolated Kubernetes pods (with NetworkPolicy), query logs, check deployments, and correlate data across services and providers; investigations draw on Knowledge Base RAG and the Memgraph-backed infrastructure graph for blast radius analysis. On completion, Aurora produces an RCA report with timeline, root cause, impact assessment, and remediation steps, can auto-generate a postmortem and export it to Confluence/Notion/SharePoint, generate a fix pull request, and run automated Actions (open PRs, notify Slack). Persistent Artifacts documents are continuously updated as the investigation progresses; the platform also exposes an MCP Server (for Cursor, Claude Desktop, Windsurf), Terraform/IaC analysis, SigmaHQ command guardrails (37 threat detection signatures), and a NeMo input rail.

  1. An SRE team paged by PagerDuty or Datadog at 3 AM wants the alert auto-triaged and investigated in the background instead of waking an engineer for a 30-60 minute manual investigation.
  2. Multi-cloud teams need one agent to investigate incidents spanning AWS, Azure, GCP, Kubernetes, or even Fly.io and trace blast radius across services.
  3. Engineering leads want auto-generated postmortems with timeline and impact assessment exported straight to Confluence, Notion, or SharePoint.
  4. Platform teams want to capture investigation reasoning so knowledge stops living in individual engineers' heads.
  5. Security-sensitive organizations require incident data to never leave their infrastructure, with Ollama support for fully air-gapped operation.
  6. MCP users want to invoke Aurora's investigation capabilities directly from Cursor or Claude Desktop.

How do you install or deploy this agent?

Local evaluation in under 5 minutes (Docker required):

bash

git clone https://github.com/arvo-ai/aurora.git && cd aurora
make init                # Generate secure secrets
nano .env                # Add your LLM API key (OpenRouter, OpenAI, etc.)
make prod-prebuilt       # Pull prebuilt images and start

Open http://localhost:3000; the first registered user becomes admin. Vault setup is required after first start:

bash

# Get the auto-generated root token
docker logs vault-init 2>&1 | grep "Root Token:"
# Add it to .env
echo "VAULT_TOKEN=hvs.your-token-here" >> .env
# Restart to connect services to Vault
make down && make prod-prebuilt

Other options: pin a version with make prod-prebuilt VERSION=v1.2.3; build from source with make prod-local; production Kubernetes via Helm:

bash

helm repo add aurora https://raw.githubusercontent.com/Arvo-AI/aurora/gh-pages
helm repo update
helm show values aurora/aurora-oss > my-values.yaml

# Edit my-values.yaml, then:

helm install aurora-oss aurora/aurora-oss -n aurora --create-namespace -f my-values.yaml

Also available via OCI: oci://ghcr.io/arvo-ai/charts/aurora-oss. An air-tight bundle covers air-gapped / restricted networks.

How do you use this agent?

After deployment, open http://localhost:3000; the first registered user is admin. The only required external configuration is an LLM API key in .env (OpenRouter, OpenAI, Anthropic, Gemini, Vertex AI, AWS Bedrock, or self-hosted Ollama); all cloud connectors are optional and Aurora runs without any cloud provider accounts. Day-to-day flow: connect PagerDuty, Datadog, Grafana, New Relic, OpsGenie, incident.io, or CloudWatch alarm webhooks; alerts automatically trigger background investigations where agents run kubectl/aws/az/gcloud in sandboxed pods and produce an RCA. Admins can explore the infrastructure knowledge graph, export investigations as postmortems, and configure Actions to open fix PRs or notify Slack automatically when investigations complete.

What are this agent's strengths and limitations?

Pros
  • Alerts auto-trigger full-stack investigations: agents run kubectl/aws/az/gcloud in sandboxed Kubernetes pods with NetworkPolicy, off the control plane, backed by 37 SigmaHQ threat signatures and a NeMo prompt-injection rail.
  • End-to-end closure: RCA reports, auto-generated postmortems (export to Confluence/Notion/SharePoint), fix PR generation, and post-RCA Actions workflows.
  • Model- and cloud-agnostic: OpenAI, Anthropic, Gemini, Vertex AI, AWS Bedrock, OpenRouter, and Ollama (air-gapped); connectors for AWS/Azure/GCP/OVH/Scaleway/Cloudflare.
  • 100% self-hosted with zero telemetry, secrets encrypted at rest in Vault or AWS Secrets Manager, Apache 2.0 with no per-seat or per-incident pricing.
Limitations
  • High deployment overhead: Docker Compose or Kubernetes/Helm plus Vault, PostgreSQL, Redis, and Memgraph, with a manual Vault root-token configuration step after first start.
  • Ongoing LLM costs: agents call many tools and models in parallel per alert; you supply paid API keys and the source gives no usage estimates.
  • Security sensitivity: agents need credentials to execute cloud CLIs and kubectl, so routing an LLM into your cloud credential path is itself a new attack surface to evaluate.
  • As a young project, public production case studies or third-party audits beyond README/docs/Discord are not evidenced.

How does this agent compare with similar options?

Key facts side by side with the most closely related agents.

Agent Source review Form / cost Stars Updated Language Full support on
Aurora This agent 47 · Major gaps Self-hosted serviceFree + model costs ★ 426 today Python OpenAI API · Claude API
HolmesGPT SRE Agent 57 · Major gaps CLIFree + model costs ★ 3.5k 4d ago Python OpenAI API
Ongrid 53 · Major gaps Self-hosted serviceFree + model costs ★ 1.1k 4d ago Go OpenAI API · Claude API
Archestra 71 · Some gaps Self-hosted serviceFreemium ★ 4.3k today TypeScript Codex · Claude Code · OpenAI API · Claude API

How does FollowAgents rate this agent?

FollowAgents source review · FARS-2.1
Major gaps
47/ 100 5-point scale 2.4 / 5
Trust 8/29
Reliability 5/14
Adaptability 12/18
Convention 12/18
Effectiveness 7/13
Verifiability 3/8
Why each dimension lost points
Trust8 / 29 · 1.4/5

README asserts sandboxed Kubernetes pod execution, NetworkPolicy, 37 SigmaHQ signatures, NeMo input rails, and Casbin RBAC, but no implementation code is present in the provided files to substantiate any of these; deduction: all core security claims are unevidenced assertions. Additional deductions: the Vault root token is printed to docker logs and copied into .env, a plaintext sensitive-data exposure path; 'Actions' auto-trigger PR creation and Slack notification with no visible user-confirmation gate; no rollback mechanism evidence. LICENSE copyright year is 2026 (anomalous) and publisher identity is unverified.

Reliability5 / 14 · 1.8/5

The build.yml CI build validation is explicitly 'Temporarily paused' (triggers commented out), inconsistent with the maturity the README projects; the sole e2e test asserts the URL matches /(\/chat|\/)/ — it passes regardless of outcome and verifies nothing. Deduction: test substance contradicts implied rigor. Version pinning (VERSION=v1.2.3) and prebuilt images exist, but no dependency manifests are provided.

Adaptability12 / 18 · 3.3/5

Audience and scenario are clearly defined (SRE on-call incident investigation), and deployment spans local Docker, Helm, and air-gapped Ollama operation — the strongest area. Deductions: capability boundaries are entirely undocumented ('30+ tools', 'any LLM' are promotional and unbounded); trigger precision is only a generic webhook-ingestion description.

Convention12 / 18 · 3.3/5

Install notes are thorough (Quick Start, Vault setup, version pinning, Helm/OCI, air-gapped table) — full marks. Repo layout, docs site, CHANGELOG reference, and a SECURITY.md with a 72-hour acknowledgment commitment exist. Deductions: no known-limitations section (only an implicit manual Vault step); no examples/FAQ; CHANGELOG content and release history not visible in evidence; paused CI weakens the maintenance-activity case.

Effectiveness7 / 13 · 2.7/5

The README describes usable outputs (structured RCA, postmortem exports, knowledge graph) and a clear differentiation (replacing 30-60 minutes of manual investigation), but these are narrative claims with no inspectable artifacts, and LLM/token/operating costs are not discussed at all — cost-benefit analysis is absent.

Verifiability3 / 8 · 1.9/5

Key quantitative claims ('30+ tools', '37 threat detection signatures', sandbox isolation, 'zero data sent') are untraceable in the provided files; the code in evidence (one API proxy route, one no-op e2e test) provides almost no cross-corroboration with the README's broad scope; facts and marketing inference are interleaved ('Aurora does all of that autonomously') with no fact/inference separation. Deduction: the claim-to-evidence chain is broadly broken.

Risks and how to mitigate them
  • Not found in source: rollback or recovery pathBack up first, or work on a git branch or snapshot, so its changes can be undone.
  • All security claims (sandboxing, guardrails, RBAC) are README assertions with no implementation visible in the provided files; audit the actual isolation and permission code under server/ and deploy/ before deployment.
  • The Vault root token workflow (printed to docker logs, written into .env) is a plaintext secret-exposure risk; use controlled secret injection in production.
  • 'Actions' such as auto-created PRs and Slack notifications show no visible confirmation gate; assess misfire risk before connecting real cloud accounts.
  • CI build validation is currently paused and the only e2e test is a tautological assertion — weak code-quality signals.
  • The agent requires credentials spanning AWS/Azure/GCP/Kubernetes, far beyond least privilege; configure read-only roles per connector and review each.
  • LICENSE copyright year reads 2026 and publisher identity is unverified; independently confirm the organization's authenticity.
Evidence confidence: Low Reviewed Sep 28, 2026 Reviewed revision 14177281f457
See the full review method →

FAQ

Do I need cloud provider accounts to use it?
No. The README states the LLM API key is the only external requirement; all cloud/Kubernetes connectors are optional. It deploys and runs without connectors, though investigation scope will be limited.
Is it safe to let agents run commands?
Commands run in isolated Kubernetes pods with NetworkPolicy, not on your control plane; SigmaHQ provides 37 threat-detection signatures on agent command execution, a NeMo input rail detects prompt injection, and org-level command policies let you restrict behavior.
Does my data go to Arvo AI?
No. Aurora is 100% self-hosted with zero telemetry; LLM calls go directly from your infrastructure to your chosen provider, and Ollama enables fully air-gapped operation.
Which LLMs are supported, and what does it cost?
OpenAI, Anthropic, Gemini, Vertex AI, AWS Bedrock, OpenRouter, and Ollama are supported. The software is free under Apache 2.0, but you pay your chosen model provider for API usage.
How do I deploy to production?
Use the official Helm chart (GKE/EKS/AKS, also installable from the OCI registry) or the air-tight bundle for air-gapped networks; local evaluation uses make prod-prebuilt.
View on GitHub ↗ Install ↓

Related agents