optimize my cluster: idle workloads, rightsizing, and HPA hints — without auto-apply
How kprompt’s shipped optimize report works: inventory, Prometheus-backed idle and rightsizing findings, HPA hints, JSON output, and optional follow-up scale/patch plans that still require their own approval. What it is not.
“Optimize my cluster” is one of those prompts every platform team wants to type — and every vendor wants to auto-remediate. kprompt ships a different shape: a read-only report (inventory, idle signals, rightsizing deltas, HPA hints), then optional suggested fixes that still go through plan → approve. Passing --approve on the optimize prompt does not silently patch your Deployments.
What the report covers (shipped)
- Inventory — workloads with replicas and request/limit sketches
- Idle / underutilized — when Prometheus usage vs requests supports it
- Rightsizing — CPU/memory request/limit deltas from usage percentiles
- HPA hints — present, maxed, or static-replica narration
- Suggestions — human-readable follow-ups with action hints
Without Prometheus configured, inventory and structural HPA notes still help; idle and rightsizing degrade honestly instead of inventing savings numbers.
Run it
Read-only optimize
brew install kprompt/tap/kprompt
export KPROMPT_GEMINI_API_KEY="..."
# Optional but recommended for idle / rightsizing
kprompt config set tools.prometheus.url http://prometheus.monitoring:9090
kprompt "optimize my cluster"
kprompt "optimize my cluster" -n production
kprompt "optimize my cluster" -o json # CI / jqIllustrative terminal shape
Optimize: cluster (1h)
Summary: 42 workloads scanned; 3 idle; 5 rightsizing candidates
Findings:
- [medium] Idle replicas: payment-worker under 5% CPU vs requests
- [low] HPA maxed: checkout at maxReplicas
Idle:
- Deployment/payment-worker: low CPU vs requests over window
Rightsizing:
- Deployment/api: memory request 512Mi → suggest 384Mi
HPA:
- Deployment/checkout: HPA present, currently at max
Suggestions:
- Scale down idle worker: review replicas (optional approved scale)
- Lower memory request on api: review patch (optional approved patch)Optional fixes still need their own approval
Top findings can become a separate scale or patch plan. That plan is risk-evaluated and needs TTY y/N or an explicit --approve on that follow-up — not the parent optimize flag. This is intentional: optimize is a report, not a bot that rightsizes production while you get coffee.
Report first, mutate second
kprompt "optimize my cluster" --approve
# Still read-only report — does not auto-apply suggestions
# Later, if you agree with a suggestion:
kprompt "scale payment-worker to 1" -n jobs
# → Plan → Apply? [y/N]JSON for pipelines
With --output json, the PlanResult carries the optimize report (idle, rightsizing, HPA, findings). Use it to fail a nightly job when high-severity idle findings appear — or to open a ticket — without applying anything.
What it is not
- Not a FinOps bill exporter (dollar / carbon notes are a separate backlog item)
- Not a guarantee of “safe” rightsizing without your review
- Not a replacement for Kubecost, OpenCost, or capacity planning reviews
- Not multi-cluster fleet rollup yet (local report per context; fleet is later)
- Not advice to --approve optimize in production and walk away
How it fits the AI SRE path
Optimize is cluster-level thinking under the intent-compiler contract: evidence in, structured report out, mutate only with a second gate. Pair it with the service dependency graph when you need consumers before you shrink a Deployment, and with the AI SRE roadmap when you want investigate / blast-radius next.
Try on staging first
Safe starting point
kprompt tools
kprompt "optimize my cluster" -n staging
kprompt "show service dependency graph" -n stagingkprompt remains experimental. Prefer non-production while you learn how findings look for your charted workloads and Prom labels. Star issues if a rightsizing heuristic is wrong for your stack — the report should stay honest, not optimistic. For the JSON shape of reports and plans, see the PlanResult JSON deep dive.
Related posts
kprompt + kagent: PlanResult as an MCP tool under a CNCF agent platform
How to compose kprompt with kagent without collapsing the layers: kagent hosts Agents-as-CRDs via MCPServer / RemoteMCPServer; kprompt ships read/plan-only MCP tools that return a typed PlanResult and never auto-apply. Validated against kagent quickstart + first MCP tool docs.
Read articlekprompt on Google Cloud: GKE day-2 with Gemini, without a new control plane
Use kprompt against GKE the same way you use kubectl — get-credentials, aliases, plan-before-apply — plus Gemini BYOK. Optional Observe agent on the cluster; no Marketplace SaaS, no kubeconfig upload.
Read articlekprompt on AWS: EKS day-2 with BYOK, without a new control plane
Use kprompt against Amazon EKS the same way you use kubectl — update-kubeconfig, aliases, plan-before-apply — plus Ollama or cloud BYOK. Optional Observe agent on the cluster; no Marketplace SaaS, no kubeconfig upload. Native Bedrock preset still deferred.
Read article