Building AI SRE in Public #1: Why AI SRE
AI kubectl is not enough. Production needs investigate, why, blast radius, and verify — still behind an approval boundary. Why the AI SRE category exists, what failed in classic AIOps, and what we ship first.
This is episode 1 of Building AI SRE in Public. The series hub lists the full arc — from intent compiler to why we still refuse unsupervised auto-remediation.
Most “AI for Kubernetes” products today are AI kubectl: natural language that emits or runs kubectl-shaped actions. That is useful. It is not SRE. SRE is the craft of keeping systems reliable under change — detecting symptoms, forming hypotheses, bounding blast radius, changing one thing at a time, and verifying the goal. An AI that only shortens the typing does not change that craft. An AI that participates in that craft — still under your credentials and your approval — is the category we call AI SRE.
The sentence that defines the category
Imagine an assistant that can say: “Error rate on payment rose after yesterday’s rollout; the Service still selects the old pods; here is a rollback plan with affected namespaces.” That sentence requires investigation graph, timeline, and a typed plan — not a chat transcript. Dashboards show charts. Fleet scanners dump findings. Chat CLIs race to the next kubectl. AI SRE is the system that proposes a reviewable next step.
Why classic AIOps struggled
AIOps promised correlation and auto-remediation years before LLMs. Many deployments stalled for boring reasons: brittle rules, noisy alerts, opaque black boxes, and remediation that operators did not trust. The models were weak at intent; the systems were strong at false confidence.
- Rules and ML that could not explain themselves in operator language
- Auto-remediation that skipped human judgment on shared clusters
- Tools that lived beside kubectl instead of composing with GitOps and metrics
- No shared artifact — only tickets, runbooks, and tribal memory
LLMs change the input side: they parse messy human intent and narrate evidence. They do not magically make unsupervised mutate safe. What changes is the chance to build an intentional loop — compile intent into a typed plan, attach risk and denies, require approval, then verify — instead of a chatbot that “just ran it.”
Approval boundary is the product
Every production AI agent that can change state needs an approval boundary. Human-in-the-loop is not theater; it is how you keep blast radius conscious. In kprompt the boundary is concrete: PlanResult on stdout (and JSON for CI), safety scoring, hard denies for wipe-class intents, interactive y/N or explicit --approve, and no silent multi-context apply from one flag.
The contract does not disappear for “smart” features
kprompt "scale api to 3" -n staging
# → Plan + risk → Apply? [y/N]
kprompt "optimize my cluster"
# → Report first; mutate follow-ups still need approveWhat we ship first (wedge, not wish)
AI SRE is the destination. The wedge is already usable: day-2 ops and investigation-shaped reads under the same plan → safety → approve → apply loop, plus integrations (Helm, Prom, OTel, GitOps, …) that keep one approval surface. Optimize reports and dependency graphs are early “think about the cluster” features — still not auto-remediation.
- Shipped: intent → PlanResult → approve; investigate / why / timeline / impact; blast-radius + --wait verify; Observe; namespace memory; service + topology graph MVP (v0.10)
- Building: deeper multi-agent reasoning; interactive Team topology UI; richer Prom/OTel/mesh hops
- Exploring: ADR/docs knowledge nodes — never silent mutate
What this episode is not
- Not a claim that kprompt is a finished AI SRE product
- Not a pitch for unsupervised auto-remediation
- Not “chat replaces kubectl forever” — compilers need escape hatches
Next
Episode 2 digs into the Intent Compiler — why Kubernetes deserves a compiler, not a chatbot, and how Intent → Action → PlanResult becomes the IR. Read it next, or revisit the hub for the full arc.
Related posts
kprompt + kagent: PlanResult as an MCP tool under a CNCF agent platform
How to compose kprompt with kagent without collapsing the layers: kagent hosts Agents-as-CRDs via MCPServer / RemoteMCPServer; kprompt ships read/plan-only MCP tools that return a typed PlanResult and never auto-apply. Validated against kagent quickstart + first MCP tool docs.
Read articlekprompt on Google Cloud: GKE day-2 with Gemini, without a new control plane
Use kprompt against GKE the same way you use kubectl — get-credentials, aliases, plan-before-apply — plus Gemini BYOK. Optional Observe agent on the cluster; no Marketplace SaaS, no kubeconfig upload.
Read articlekprompt on AWS: EKS day-2 with BYOK, without a new control plane
Use kprompt against Amazon EKS the same way you use kubectl — update-kubeconfig, aliases, plan-before-apply — plus Ollama or cloud BYOK. Optional Observe agent on the cluster; no Marketplace SaaS, no kubeconfig upload. Native Bedrock preset still deferred.
Read article