optimize my cluster: idle workloads, rightsizing, and HPA hints — without auto-apply
How kprompt’s shipped optimize report works: inventory, Prometheus-backed idle and rightsizing findings, HPA hints, JSON output, and optional follow-up scale/patch plans that still require their own approval. What it is not.
“Optimize my cluster” is one of those prompts every platform team wants to type — and every vendor wants to auto-remediate. kprompt ships a different shape: a read-only report (inventory, idle signals, rightsizing deltas, HPA hints), then optional suggested fixes that still go through plan → approve. Passing --approve on the optimize prompt does not silently patch your Deployments.
What the report covers (shipped)
- Inventory — workloads with replicas and request/limit sketches
- Idle / underutilized — when Prometheus usage vs requests supports it
- Rightsizing — CPU/memory request/limit deltas from usage percentiles
- HPA hints — present, maxed, or static-replica narration
- Suggestions — human-readable follow-ups with action hints
Without Prometheus configured, inventory and structural HPA notes still help; idle and rightsizing degrade honestly instead of inventing savings numbers.
Run it
Read-only optimize
brew install kprompt/tap/kprompt
export KPROMPT_GEMINI_API_KEY="..."
# Optional but recommended for idle / rightsizing
kprompt config set tools.prometheus.url http://prometheus.monitoring:9090
kprompt "optimize my cluster"
kprompt "optimize my cluster" -n production
kprompt "optimize my cluster" -o json # CI / jqIllustrative terminal shape
Optimize: cluster (1h)
Summary: 42 workloads scanned; 3 idle; 5 rightsizing candidates
Findings:
- [medium] Idle replicas: payment-worker under 5% CPU vs requests
- [low] HPA maxed: checkout at maxReplicas
Idle:
- Deployment/payment-worker: low CPU vs requests over window
Rightsizing:
- Deployment/api: memory request 512Mi → suggest 384Mi
HPA:
- Deployment/checkout: HPA present, currently at max
Suggestions:
- Scale down idle worker: review replicas (optional approved scale)
- Lower memory request on api: review patch (optional approved patch)Optional fixes still need their own approval
Top findings can become a separate scale or patch plan. That plan is risk-evaluated and needs TTY y/N or an explicit --approve on that follow-up — not the parent optimize flag. This is intentional: optimize is a report, not a bot that rightsizes production while you get coffee.
Report first, mutate second
kprompt "optimize my cluster" --approve
# Still read-only report — does not auto-apply suggestions
# Later, if you agree with a suggestion:
kprompt "scale payment-worker to 1" -n jobs
# → Plan → Apply? [y/N]JSON for pipelines
With --output json, the PlanResult carries the optimize report (idle, rightsizing, HPA, findings). Use it to fail a nightly job when high-severity idle findings appear — or to open a ticket — without applying anything.
What it is not
- Not a FinOps bill exporter (dollar / carbon notes are a separate backlog item)
- Not a guarantee of “safe” rightsizing without your review
- Not a replacement for Kubecost, OpenCost, or capacity planning reviews
- Not multi-cluster fleet rollup yet (local report per context; fleet is later)
- Not advice to --approve optimize in production and walk away
How it fits the AI SRE path
Optimize is cluster-level thinking under the intent-compiler contract: evidence in, structured report out, mutate only with a second gate. Pair it with the service dependency graph when you need consumers before you shrink a Deployment, and with the AI SRE roadmap when you want investigate / blast-radius next.
Try on staging first
Safe starting point
kprompt tools
kprompt "optimize my cluster" -n staging
kprompt "show service dependency graph" -n stagingkprompt remains experimental. Prefer non-production while you learn how findings look for your charted workloads and Prom labels. Star issues if a rightsizing heuristic is wrong for your stack — the report should stay honest, not optimistic. For the JSON shape of reports and plans, see the PlanResult JSON deep dive.
Related posts
Building AI SRE in Public #4: Safety Engine
Policy is code, not LLM vibes. How kprompt’s safety engine hard-denies wipe-class intents, scores risk, forces approval, and why fail-closed is the load-bearing wall of AI SRE.
Read articleBuilding AI SRE in Public #3: PlanResult
PlanResult is the IR of AI SRE: one typed document for humans and CI. Why JSON, what applied means vs verify, how blastRadius attaches, what never gets stored, and how investigate/why must extend the same artifact.
Read articleBuilding AI SRE in Public #2: Intent Compiler
Kubernetes deserves a compiler, not a chatbot. How kprompt turns natural language into Intent → Actions → PlanResult, why Go owns planning and safety, and why the IR must stay reviewable for AI SRE.
Read article