Aurora
Open-source AI-agent incident management and root cause analysis: it investigates your stack automatically when an alert fires and delivers a structured RCA before you open your laptop.
- Source repo
- Arvo-AI/aurora
- Stars
- ★ 426
- Last updated
- today
- License
- Apache-2.0
- Primary language
- Python
- FA score
- 47/100 · Major gaps
At a glance
- How it runs
- Works with
- Universal · cross-platformOpenAI API · Claude API
- Cost
- Free software; you pay for model usage
- Setup effort
- High · needs real infrastructure
- You'll need
- Typical use
- An SRE team paged by PagerDuty or Datadog at 3 AM wants the alert auto-triaged and investigated in the background instead of waking an engineer for a 30-60 minute manual investigation.
- Not a fit if
- Teams that cannot self-host Docker/Kubernetes infrastructure
- Organizations unwilling to grant LLM agents access to cloud and cluster credentials
- Users who want a vendor-hosted SaaS that works out of the box
- Source review
- 47/100 · Major gaps 1 safety controls not found
What does this agent do, and when should you use it?
Aurora is an open-source, AI-agent-powered incident management platform from Arvo-AI, built for SRE and DevOps teams. When an alert arrives, LangGraph agents dynamically pick from 30+ tools and run kubectl, aws, az, and gcloud inside sandboxed Kubernetes pods, querying logs, checking deployments, and correlating data across AWS, Azure, GCP, and Kubernetes to produce a structured root cause analysis. The system consists of a Python/Flask backend with Celery workers, a Next.js frontend, Memgraph for the infrastructure knowledge graph, Vault for secrets, and PostgreSQL/Redis storage, deployed via Docker Compose or Helm entirely on your own infrastructure with zero telemetry. It also auto-generates postmortems exportable to Confluence, Notion, or SharePoint, traces blast radius across services through a dependency graph, and can generate pull requests with remediation fixes.
Aurora ingests alerts from PagerDuty, Datadog, Grafana, New Relic, OpsGenie, incident.io, and CloudWatch alarm webhooks; every alert auto-triggers a background investigation. LangGraph agents select from 30+ tools and run kubectl, aws, az, and gcloud in isolated Kubernetes pods (with NetworkPolicy), query logs, check deployments, and correlate data across services and providers; investigations draw on Knowledge Base RAG and the Memgraph-backed infrastructure graph for blast radius analysis. On completion, Aurora produces an RCA report with timeline, root cause, impact assessment, and remediation steps, can auto-generate a postmortem and export it to Confluence/Notion/SharePoint, generate a fix pull request, and run automated Actions (open PRs, notify Slack). Persistent Artifacts documents are continuously updated as the investigation progresses; the platform also exposes an MCP Server (for Cursor, Claude Desktop, Windsurf), Terraform/IaC analysis, SigmaHQ command guardrails (37 threat detection signatures), and a NeMo input rail.
- An SRE team paged by PagerDuty or Datadog at 3 AM wants the alert auto-triaged and investigated in the background instead of waking an engineer for a 30-60 minute manual investigation.
- Multi-cloud teams need one agent to investigate incidents spanning AWS, Azure, GCP, Kubernetes, or even Fly.io and trace blast radius across services.
- Engineering leads want auto-generated postmortems with timeline and impact assessment exported straight to Confluence, Notion, or SharePoint.
- Platform teams want to capture investigation reasoning so knowledge stops living in individual engineers' heads.
- Security-sensitive organizations require incident data to never leave their infrastructure, with Ollama support for fully air-gapped operation.
- MCP users want to invoke Aurora's investigation capabilities directly from Cursor or Claude Desktop.
How do you install or deploy this agent?
Local evaluation in under 5 minutes (Docker required):
bash
git clone https://github.com/arvo-ai/aurora.git && cd auroramake init # Generate secure secrets
nano .env # Add your LLM API key (OpenRouter, OpenAI, etc.)
make prod-prebuilt # Pull prebuilt images and startOpen http://localhost:3000; the first registered user becomes admin. Vault setup is required after first start:
bash
# Get the auto-generated root token
docker logs vault-init 2>&1 | grep "Root Token:"# Add it to .env
echo "VAULT_TOKEN=hvs.your-token-here" >> .env# Restart to connect services to Vault
make down && make prod-prebuiltOther options: pin a version with make prod-prebuilt VERSION=v1.2.3; build from source with make prod-local; production Kubernetes via Helm:
bash
helm repo add aurora https://raw.githubusercontent.com/Arvo-AI/aurora/gh-pages
helm repo update
helm show values aurora/aurora-oss > my-values.yaml# Edit my-values.yaml, then:
helm install aurora-oss aurora/aurora-oss -n aurora --create-namespace -f my-values.yamlAlso available via OCI: oci://ghcr.io/arvo-ai/charts/aurora-oss. An air-tight bundle covers air-gapped / restricted networks.
How do you use this agent?
After deployment, open http://localhost:3000; the first registered user is admin. The only required external configuration is an LLM API key in .env (OpenRouter, OpenAI, Anthropic, Gemini, Vertex AI, AWS Bedrock, or self-hosted Ollama); all cloud connectors are optional and Aurora runs without any cloud provider accounts. Day-to-day flow: connect PagerDuty, Datadog, Grafana, New Relic, OpsGenie, incident.io, or CloudWatch alarm webhooks; alerts automatically trigger background investigations where agents run kubectl/aws/az/gcloud in sandboxed pods and produce an RCA. Admins can explore the infrastructure knowledge graph, export investigations as postmortems, and configure Actions to open fix PRs or notify Slack automatically when investigations complete.
What are this agent's strengths and limitations?
- Alerts auto-trigger full-stack investigations: agents run kubectl/aws/az/gcloud in sandboxed Kubernetes pods with NetworkPolicy, off the control plane, backed by 37 SigmaHQ threat signatures and a NeMo prompt-injection rail.
- End-to-end closure: RCA reports, auto-generated postmortems (export to Confluence/Notion/SharePoint), fix PR generation, and post-RCA Actions workflows.
- Model- and cloud-agnostic: OpenAI, Anthropic, Gemini, Vertex AI, AWS Bedrock, OpenRouter, and Ollama (air-gapped); connectors for AWS/Azure/GCP/OVH/Scaleway/Cloudflare.
- 100% self-hosted with zero telemetry, secrets encrypted at rest in Vault or AWS Secrets Manager, Apache 2.0 with no per-seat or per-incident pricing.
- High deployment overhead: Docker Compose or Kubernetes/Helm plus Vault, PostgreSQL, Redis, and Memgraph, with a manual Vault root-token configuration step after first start.
- Ongoing LLM costs: agents call many tools and models in parallel per alert; you supply paid API keys and the source gives no usage estimates.
- Security sensitivity: agents need credentials to execute cloud CLIs and kubectl, so routing an LLM into your cloud credential path is itself a new attack surface to evaluate.
- As a young project, public production case studies or third-party audits beyond README/docs/Discord are not evidenced.
How does this agent compare with similar options?
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| Aurora This agent | 47 · Major gaps | Self-hosted serviceFree + model costs | ★ 426 | today | Python | OpenAI API · Claude API |
| HolmesGPT SRE Agent | 57 · Major gaps | CLIFree + model costs | ★ 3.5k | 4d ago | Python | OpenAI API |
| Ongrid | 53 · Major gaps | Self-hosted serviceFree + model costs | ★ 1.1k | 4d ago | Go | OpenAI API · Claude API |
| Archestra | 71 · Some gaps | Self-hosted serviceFreemium | ★ 4.3k | today | TypeScript | Codex · Claude Code · OpenAI API · Claude API |
How does FollowAgents rate this agent?
Why each dimension lost points
README asserts sandboxed Kubernetes pod execution, NetworkPolicy, 37 SigmaHQ signatures, NeMo input rails, and Casbin RBAC, but no implementation code is present in the provided files to substantiate any of these; deduction: all core security claims are unevidenced assertions. Additional deductions: the Vault root token is printed to docker logs and copied into .env, a plaintext sensitive-data exposure path; 'Actions' auto-trigger PR creation and Slack notification with no visible user-confirmation gate; no rollback mechanism evidence. LICENSE copyright year is 2026 (anomalous) and publisher identity is unverified.
The build.yml CI build validation is explicitly 'Temporarily paused' (triggers commented out), inconsistent with the maturity the README projects; the sole e2e test asserts the URL matches /(\/chat|\/)/ — it passes regardless of outcome and verifies nothing. Deduction: test substance contradicts implied rigor. Version pinning (VERSION=v1.2.3) and prebuilt images exist, but no dependency manifests are provided.
Audience and scenario are clearly defined (SRE on-call incident investigation), and deployment spans local Docker, Helm, and air-gapped Ollama operation — the strongest area. Deductions: capability boundaries are entirely undocumented ('30+ tools', 'any LLM' are promotional and unbounded); trigger precision is only a generic webhook-ingestion description.
Install notes are thorough (Quick Start, Vault setup, version pinning, Helm/OCI, air-gapped table) — full marks. Repo layout, docs site, CHANGELOG reference, and a SECURITY.md with a 72-hour acknowledgment commitment exist. Deductions: no known-limitations section (only an implicit manual Vault step); no examples/FAQ; CHANGELOG content and release history not visible in evidence; paused CI weakens the maintenance-activity case.
The README describes usable outputs (structured RCA, postmortem exports, knowledge graph) and a clear differentiation (replacing 30-60 minutes of manual investigation), but these are narrative claims with no inspectable artifacts, and LLM/token/operating costs are not discussed at all — cost-benefit analysis is absent.
Key quantitative claims ('30+ tools', '37 threat detection signatures', sandbox isolation, 'zero data sent') are untraceable in the provided files; the code in evidence (one API proxy route, one no-op e2e test) provides almost no cross-corroboration with the README's broad scope; facts and marketing inference are interleaved ('Aurora does all of that autonomously') with no fact/inference separation. Deduction: the claim-to-evidence chain is broadly broken.
- Not found in source: rollback or recovery pathBack up first, or work on a git branch or snapshot, so its changes can be undone.
- All security claims (sandboxing, guardrails, RBAC) are README assertions with no implementation visible in the provided files; audit the actual isolation and permission code under server/ and deploy/ before deployment.
- The Vault root token workflow (printed to docker logs, written into .env) is a plaintext secret-exposure risk; use controlled secret injection in production.
- 'Actions' such as auto-created PRs and Slack notifications show no visible confirmation gate; assess misfire risk before connecting real cloud accounts.
- CI build validation is currently paused and the only e2e test is a tautological assertion — weak code-quality signals.
- The agent requires credentials spanning AWS/Azure/GCP/Kubernetes, far beyond least privilege; configure read-only roles per connector and review each.
- LICENSE copyright year reads 2026 and publisher identity is unverified; independently confirm the organization's authenticity.