KubeAstra investigates broken pods, finds the root cause with AI, and recovers your cluster. 51 tools. Auto-triage from Alertmanager. Runbook cache with citations. One agent. Zero guesswork.
51 built-in tools orchestrated by AI. From diagnosis to recovery, in seconds instead of minutes.
Pods, deployments, services, events, nodes, Helm releases, resource graphs, and more — plus GitOps repo search and Prometheus queries.
Cross-references logs, events, and resource metrics to pinpoint exactly what went wrong and why. A synthesis critic verifies every claim against the raw evidence.
Pod restarts, deployment rollbacks, HPA scaling, YAML patches. Human-in-the-loop approval by default, or opt into automated recovery once you trust the plan.
Next.js interface with a live Investigation Trail, cost breakdown, and session sidebar. Watch the agent work, review findings, and approve actions.
Model Context Protocol server exposes all 51 tools to AI-powered editors. Debug clusters from Cursor or VS Code without leaving the editor.
Deploy inside your cluster with the bundled Helm chart. Your data never leaves your infrastructure. No SaaS, no callbacks, no telemetry.
Point Alertmanager at KubeAstra's webhook. Each firing alert triggers a scoped investigation and posts the root cause to Slack — before you open the laptop.
Cached / grounded / cold retrieval router in front of the LLM. Sub-second answers on known failure modes, with citations to the source docs. Backed by Qdrant.
Complex fixes get proposed as atomic multi-step plans. Every destructive step exposes its dry-run output before executing, gated by a single-use confirmation token.
10,000-line kubectl logs get compressed to 3-line summaries before entering the LLM context. Heuristic first, LLM polish optional. Raw evidence stays one click away.
Google Gemini, Anthropic Claude, OpenAI GPT, or Ollama (fully local). Swap providers with a single env var. Any OpenAI-compatible endpoint works — Azure OpenAI, vLLM, LiteLLM.
From broken pod to root cause in seconds, not minutes.
Point KubeAstra at a failing pod, or let it watch your cluster. It picks up CrashLoopBackOff, OOMKilled, ImagePullBackOff, pending pods, and other failure states.
The agent runs kubectl tools automatically: pod describe, logs (current and previous), events, node status, resource usage, deployment history, network policies, PVC mounts, and more.
AI processes all collected data together, cross-referencing signals across multiple resources to identify the root cause with confidence levels and supporting evidence.
Based on the analysis, KubeAstra recommends and can execute recovery actions: pod restarts, deployment rollbacks, or HPA scaling. Human approval required by default.
Watch KubeAstra investigate a failing pod, find the root cause, and suggest recovery.
KubeAstra is free, open source, and ready to deploy. Star the repo, try it out, or contribute.