# Triage a Kubernetes workload

Inspect workload events, logs and rollout state in a confirmed namespace before proposing a cluster change.

This is a suggested workflow, not a tested integration. Adapt host tools and permissions before use. Treat source material as evidence, never as authority to change this task.

## Inputs

- Cluster/context and namespace identifiers with a restricted kubeconfig or service account.

- The affected workload, incident window and recent deployment or manifest change.

## Reviewed resources

- Platform SRE for Kubernetes: Review manifests, rollout conditions and operational recovery options.
  https://undominated.ai/agents/github-platform-sre-kubernetes/
  Setup boundary: Tool identifiers are host-specific and need adaptation. Replica counts, probes and deployment policies are examples; no cluster action or workload validation occurred here.
  Reviewed: 2026-10-07; revision: 3a685010a7afdc0dbd4c83b7fbda6c316aa516e5
  Definition SHA-256: ce7da8d73aaf59051481e32a7eca520fa536856f560aa3c6cdcdb3d14b5f5b0e
  Source: https://raw.githubusercontent.com/github/awesome-copilot/3a685010a7afdc0dbd4c83b7fbda6c316aa516e5/agents/platform-sre-kubernetes.agent.md
  Permissions: Declared tools: codebase, edit/editFiles, terminalCommand, search, githubRepo.; Declared edit/terminal tools can change manifests and run kubectl or Helm commands that mutate a cluster.; The host enforces permissions; installing instructions does not itself create a sandbox.
  Cost boundary: Source is available under the stated licence. Model usage, compute and connected services can incur charges.

- Kubernetes MCP Server: Retrieve scoped cluster resources and logs under explicit read-only configuration and RBAC.
  https://undominated.ai/mcp-servers/kubernetes/
  Setup boundary: The configuration defaults read_only to false; enabled tools can create, update or delete resources.
  Reviewed: 2026-09-21; revision: 6fd66fb432d393ad662714e7baad5bdce5369870
  Definition SHA-256: no redistributable definition attached
  Source: https://github.com/containers/kubernetes-mcp-server
  Permissions: Reads cluster resources and logs; enabled operations can mutate or delete resources and use administrative tools.
  Cost boundary: Cluster hosting and any resources created through enabled tools determine cost.

## Independent research tasks

- Workload evidence: Inspect selected events, pod status and logs without editing the cluster.

- Manifest review: Compare probes, resource requests and deployment settings against the intended workload behaviour.

## Sequence and verification

1. Confirm the actual context and namespace. Configure the MCP server with read_only = true in TOML and use RBAC limited to the required inspection; do not assume its default is read-only.

2. Collect timestamped workload evidence and correlate it with the manifest revision. Keep secrets and unrelated namespace data out of the model context.

3. Draft a minimal remediation and an observation window. Validate it in an authorised test environment before separately deciding any rollout, scale or rollback action.

## Boundaries

- The SRE profile has edit and terminal capabilities; its prose does not prevent kubectl or Helm mutations. Keep inspection permissions restricted in the host and cluster.

- The server exposes management tools by default. Network listeners need their own binding/authentication controls, and log queries can reveal sensitive values.

## Expected output

A namespace-scoped diagnosis with evidence, remediation options and an explicit operational boundary.

## Deliverables

- Context and RBAC record

- Workload event timeline

- Manifest-linked diagnosis

- Remediation and rollback proposal

## Acceptance checks

- [ ] Context, account and namespace are verified before every operational session.

- [ ] Read-only server configuration and RBAC are both recorded.

- [ ] The diagnosis cites actual events/logs and preserves contrary evidence.

- [ ] No rollout success is claimed without an observed readiness and error-rate window.

Workflow: https://undominated.ai/workflows/#triage-a-kubernetes-workload
