All posts
Muhtalip Dede profile photoMuhtalip Dede · Founder of kprompt4 min read

kprompt on Google Cloud: GKE day-2 with Gemini, without a new control plane

Use kprompt against GKE the same way you use kubectl — get-credentials, aliases, plan-before-apply — plus Gemini BYOK. Optional Observe agent on the cluster; no Marketplace SaaS, no kubeconfig upload.

Teams on Google Cloud often ask: “How do we run kprompt on GCP?” The honest answer is not a Marketplace listing or a managed fleet SaaS. kprompt is a laptop CLI (and an optional in-cluster Observe agent) that speaks Kubernetes. On GCP that means GKE in your kubeconfig, Gemini as your bring-your-own-key model if you want Google’s stack end-to-end — and the same plan → safety → approve contract you get on kind or EKS.

If you already operate GKE with gcloud and kubectl, you already have the hard parts. This post is the GCP-shaped path: credentials, aliases, Gemini, useful day-2 prompts, and when (not) to install the Observe agent.

What “on Google Cloud” actually means

LayerWhat you useWhat kprompt does
ClusterGKE (Autopilot or Standard)Read / plan / apply via your kubeconfig — never uploads credentials
LLMGemini (AI Studio key) or Ollama / other BYOKIntent → PlanResult; keys stay in env vars
Optional agentHelm chart in a namespaceWatch → Incident → gated notify; propose-only by default
Not in scopegcloud / Terraform / Config ConnectorDoes not provision GKE clusters or GCP projects

That last row matters. “Create me a GKE cluster in us-central1” is still gcloud or your IaC. kprompt’s lane is day-2: investigate CrashLoop, scale a Deployment, open a reviewable plan — after the cluster exists.

1. Point kubeconfig at GKE

Same muscle memory as kubectl. Authenticate the Google Cloud SDK, then fetch credentials for the cluster you care about. Prefer a non-production cluster for the first session.

GKE credentials into kubeconfig

gcloud auth login
gcloud config set project PROJECT_ID

gcloud container clusters get-credentials CLUSTER_NAME \
  --region REGION \
  # or: --zone ZONE
  --project PROJECT_ID

kubectl config current-context
# → gke_PROJECT_REGION_CLUSTER_NAME (typical shape)

kprompt doctor

doctor checks kube reachability and LLM readiness. If the API server is unreachable, fix gcloud / network / VPC access first — kprompt will not invent a tunnel.

2. Alias the ugly GKE context name

GKE context names are long on purpose. Aliases keep blast radius mental: prod means one string, staging means another. require_alias_match refuses a mutate when kubectl’s current-context does not match the alias you asked for — fat-finger insurance when three GKE contexts sit in one file.

Short names → GKE contexts

kprompt contexts
kprompt contexts --check

kprompt config alias set prod gke_myproj_us-central1_prod
kprompt config alias set staging gke_myproj_us-central1_staging
kprompt config set require_alias_match true

kprompt --context staging "list deployments"
kprompt --contexts staging,prod "list pods"

Read fan-out across staging and prod is explicit. Mutate fan-out never rides on a lone --approve — you need --approve-each-context if you truly meant every listed context. Credentials still never leave the laptop.

3. Wire Gemini (or stay on Ollama)

Natural-language plans need a model. On a Google-heavy stack, Gemini is the natural BYOK choice: AI Studio key in the environment, provider set once, no secrets in ~/.kprompt/config.yaml. Prefer Ollama when you want $0 inference and no cloud quota.

Gemini BYOK

# Key from https://aistudio.google.com/apikey
export KPROMPT_GEMINI_API_KEY=...

kprompt init --provider gemini
# or:
kprompt config set provider gemini
kprompt config set model gemini-3.6-flash

kprompt --context staging "list pods"

Honesty on quotas: AI Studio free tier returns HTTP 429 when you burn daily or per-minute limits — that is Google’s meter, not a kprompt bug. Enable billing on the Google project, wait for reset, or switch to Ollama. Native Vertex AI SDK is not a shipped preset yet; if you already expose an OpenAI-compatible Vertex gateway, point openai-compatible + base_url at it.

4. Day-2 on GKE — read first, then plan

Brownfield rule still applies: first value is a read. GKE Autopilot vs Standard does not change the contract — PlanResult before apply, wipe-class intents hard-denied.

Useful GKE session shape

# Read / investigate (risk = 0)
kprompt --context staging "explain why checkout is failing" -n payments
kprompt --context staging "investigate CrashLoopBackoff" -n payments
kprompt --context staging "optimize my cluster"

# Mutate — plan only by default; TTY y/N or --approve
kprompt --context staging "scale api to 3" -n payments
kprompt --context staging "scale api to 3" -n payments --approve

# Optional: bind existing Prometheus / Grafana URLs instead of installing a second stack
kprompt tools
kprompt config set tools.prometheus.url http://prometheus.monitoring:9090

Workload Identity, NetworkPolicy, and GKE-specific CRDs still obey your RBAC. If the ServiceAccount behind your kubeconfig cannot list Pods in payments, neither can kprompt. That is a feature.

5. Optional: Observe agent inside GKE

The CLI is reactive. The Observe agent is always-on watch in one namespace: correlate Pods/Events into an Incident, optionally analyze, then gate Discord/Slack/webhook. Default mode never patches or deletes. Put LLM keys in a Secret (envFrom) — never plaintext in a ConfigMap or CR.

Namespace-scoped Helm install

helm upgrade --install kprompt-agent ./charts/kprompt-agent -n payments \
  --create-namespace \
  # LLM / Slack / Discord via Secret + values — see chart README

# Laptop smoke before you Helm:
kprompt agent run -n payments --emit-initial --analyze --fetch-logs --heuristic

Start heuristic for demos ($0). Turn on LLM analysis when you accept token spend and have tightened --min-severity / --min-confidence. Autopilot apply stays gated — propose is not silent heal.

GCP checklist

StepCommand / move
0gcloud container clusters get-credentials …
1kprompt config alias set prod <gke_context>
2export KPROMPT_GEMINI_API_KEY=… + init --provider gemini
3kprompt doctor && contexts --check
4Read prompts on staging; one plan-only mutate
5Optional: Helm Observe agent in one namespace

What we are not claiming

  • Not a Google Cloud Marketplace app or managed “kprompt on GCP” control plane
  • Not a GKE / Anthos cluster provisioner
  • Not uploading kubeconfigs to api.kprompt.ai
  • Not native Vertex SDK parity (AI Studio Gemini or OpenAI-compat gateway today)
  • Not silent remediations from the in-cluster agent

Experimental software. Prefer staging. Read every plan. On Google Cloud the win is the same as everywhere else: intentional day-2 ops on the GKE you already run — with your keys, your RBAC, and approval still on the human side of the boundary.