All posts
Muhtalip Dede profile photoMuhtalip Dede · Founder of kprompt3 min read

Building AI SRE in Public #8: Cluster Memory

Namespace dependency facts that bias Observe without becoming root-cause proof. Local or in-cluster stores — never cloud dumps as fake authority — with AG-034 confidence caps.

This is episode 8 of Building AI SRE in Public. Episode 7 ordered what already happened. Memory answers a quieter question: what should the agent already know about this namespace before the next incident? Chat RAG that forgets between sessions is not institutional knowledge. A typed fact store that never pretends to be proof is.

The claim

AI SRE needs durable, namespace-scoped priors — “payments uses Redis for sessions” — that bias Explain and Observe. Those priors must stay local or in-cluster by default, inject as labeled evidence-not-proof blocks, and never raise confidence when Events, logs, metrics, and traces are empty. Memory is a prior. EvidenceRef is proof. PlanResult still needs approval.

Chat memory is not cluster memory

A model that “remembers” last Tuesday’s Slack thread inside one session is not the same as a store an on-call can list, edit, and audit. Session memory dies with the process. Cloud dumps of cluster state as “memory” invent a fake authority. Namespace memory is boring on purpose: dependency and note facts with ids, sources, and updatedAt.

Chat / cloud “memory”Namespace memory
Folklore in the prompt windowTyped Fact[] (dependency | note)
Often leaves the cluster boundaryFile (~/.config/kprompt/memory) or ConfigMap kprompt-namespace-memory
Sounds sure without EventsAG-034: memory alone caps confidence (≤0.35)
Hard to correct after a wrong prioragent memory set / list / discover

What ships (AG-015)

  • Fact kinds: dependency (redis, kafka, postgres, …) and note (operator free-form)
  • Discover: read-only scan of Services / Deployments for known dependency signals
  • Inject: relevant facts filtered into AgentContext when incident text mentions the key or infra failure patterns
  • Backends: file by default; ConfigMap in-cluster (Helm agent.memoryBackend=configmap)
  • Privacy: never uploaded to api.kprompt.ai by default

Priors you can audit

# Discover + inject while watching (Observe)
kprompt agent run -n payments --analyze --heuristic --memory

# Manual facts
kprompt agent memory set -n payments --kind dependency --key redis --value "cache for sessions"
kprompt agent memory discover -n payments
kprompt agent memory list -n payments

# In-cluster store
kprompt agent memory list -n payments --memory-backend configmap

Evidence, not proof (AG-034)

The prompt block is explicit: namespace_memory (evidence, not proof). Confidence calibration refuses to treat memory as root cause: if Memory is non-empty and live evidence count is zero, confidence is capped and the note is memory is not proof. Patterns (AG-016) can annotate Seen before (N×) the same way — boost explainability, never auto-mutate.

That rule is a reality anchor: humans own the store and the cap; the LLM may cite facts, not waive them into certainty. Soft-agree in the same session does not override an empty Events/logs/metrics/traces set.

Where memory sits in the graph

ArtifactJobMemory’s role
Investigation / timelineWhat happened / what’s brokenHint which deps to probe — not a substitute for EvidenceRef
PlanResultProposed mutateMay bias Explain; never skips approve
Knowledge Graph (ep.9)Service graph + impact + depsMemory deps are one thin input — not a full topology product yet

Investigation Graph (ep.6) fans out and verifies. Timeline (ep.7) orders signals. Memory keeps namespace priors between runs so the next Observe pass does not rediscover “we use Kafka” from scratch — without letting that prior close the case alone.

What ships vs building

  • Shipped: namespace memory CRUD + discover + --memory inject (AG-015)
  • Shipped: AG-034 evidence-not-proof confidence cap; patterns as Observe-only priors (AG-016)
  • Shipped: file and ConfigMap backends; local/in-cluster privacy default
  • Building: richer Knowledge Graph topology (service graph + impact beyond memory deps)
  • Non-goal: uploading raw cluster dumps to the cloud as authority; memory-driven silent apply

Try the prior

Non-prod drill

# kind + dependency scenario (see kprompt-examples/07-dependencies)
kprompt agent memory discover -n payments
kprompt agent memory list -n payments
kprompt agent run -n payments --analyze --heuristic --memory
# Read the namespace_memory block — then demand Events/logs before trusting RCA

If your AI ops tool treats remembered chat as root cause, you have a confident prior wearing a badge. Cluster memory should wear a label: useful, local, and never sole proof.

Next

Episode 9 is Knowledge Graph — service graph, reverse impact, and memory deps as topology MVP, not a second chat brain. The hub tracks the rest of the arc.