---
title: "Platform SRE for Kubernetes: review, role & definition · Undominated.ai"
canonical: https://undominated.ai/agents/github-platform-sre-kubernetes/
description: "Plans Kubernetes configuration and rollout work with rollback and operational verification."
---

# Platform SRE for Kubernetes: review, role & definition · Undominated.ai

> Plans Kubernetes configuration and rollout work with rollback and operational verification.

[← Explore all agents](/agents/)

OPERATIONS AND RELIABILITY / GitHub

# Platform SRE for Kubernetes

Plans Kubernetes configuration and rollout work with rollback and operational verification.

 [Use this definition ↓](#setup)[Original source ↗](https://github.com/github/awesome-copilot/tree/3a685010a7afdc0dbd4c83b7fbda6c316aa516e5/agents)

SOURCE REVIEW

 Reviewed 2026-10-07
 Evidence 3 linked sources
 Publisher GitHub
 Licence [MIT ↗](https://github.com/github/awesome-copilot/blob/3a685010a7afdc0dbd4c83b7fbda6c316aa516e5/LICENSE)
 Revision 3a685010a7af
 [Read what was—and wasn’t—checked ↓](#review)

“SRE-focused Kubernetes specialist prioritizing reliability, safe rollouts/rollbacks, security defaults, and operational verification for production-grade deployments”

 [GitHub · upstream description ↗](https://github.com/github/awesome-copilot/blob/3a685010a7afdc0dbd4c83b7fbda6c316aa516e5/agents/platform-sre-kubernetes.agent.md) Our analysis follows below.

01 / THE REASONING

## Why this made the selection.

 - Requires manifest checks and rollout observation and carries security/resource defaults into implementation.
- Documents rollback paths and production operational checks rather than delivering unvalidated manifests alone.

### A good fit for

 - Reviewing a production Kubernetes change
- Refusing a deploy that has no rollback

### Weigh up before choosing

 - Tool identifiers are host-specific and need adaptation. Replica counts, probes and deployment policies are examples; no cluster action or workload validation occurred here.
- Reviewed definition and licence are pinned; installation and real-task outcomes for this resource were not tested.

02 / THE REVIEW RECORD

## What we actually inspected.

Source review has boundaries. A clear record is more useful than a “safe” badge.

### Material inspected

 - agents/platform-sre-kubernetes.agent.md (complete frontmatter and body)
- https://github.com/github/awesome-copilot/blob/3a685010a7afdc0dbd4c83b7fbda6c316aa516e5/LICENSE (applicable licence text)

### Our findings

 - Requires manifest checks and rollout observation and carries security/resource defaults into implementation.
- Documents rollback paths and production operational checks rather than delivering unvalidated manifests alone.
- Declared edit/terminal tools can change manifests and run kubectl or Helm commands that mutate a cluster.

### Not established by this review

 - Host discovery, task execution and model quality were not tested.
- Referenced helpers, sibling plugins and external service operations were not exhaustively audited or executed.

The review applies to the material and revision named here. A newer upstream release can change its behavior.

03 / PUT IT TO WORK

## Use the role in your project.

[Upstream setup instructions ↗](https://docs.github.com/en/copilot/reference/custom-agents-configuration)
 - Download the unchanged definition with its licence and attribution; preserve this source copy.
- Create a reviewed working copy in .github/agents/<name>.agent.md. Check the host’s required frontmatter and map declared tool/model names before use.
- Configure referenced tools or companion skills separately. Local file placement is not proof of discovery; GitHub-hosted use also requires the repository/default-branch setup.
- Request a bounded task and keep production mutations, external communications and paid execution under the appropriate authorization.

### Before you start

 - A verified cluster context/namespace, scoped RBAC and explicit authorization for each deployment environment.

### Compatibility

GitHub Copilot custom-agent format; inspect tool/model fields for the installed client

### Implementation

Markdown

### Host-mediated code, files and tools

 - Declared tools: codebase, edit/editFiles, terminalCommand, search, githubRepo.
- Declared edit/terminal tools can change manifests and run kubectl or Helm commands that mutate a cluster.
- The host enforces permissions; installing instructions does not itself create a sandbox.

### Cost model

Source is available under the stated licence. Model usage, compute and connected services can incur charges.

THE COMPLETE REVIEWED DEFINITION

## Read it before you reuse it.

Original source bytes, with attribution. Review the host-specific setup notes above.

 Copy definition ↗ [Download definition + licence ↗](/resources/agents/github-platform-sre-kubernetes/bundle.zip)[Raw Markdown ↗](/resources/agents/github-platform-sre-kubernetes/definition.md)
 ---
name: 'Platform SRE for Kubernetes'
description: 'SRE-focused Kubernetes specialist prioritizing reliability, safe rollouts/rollbacks, security defaults, and operational verification for production-grade deployments'
tools: ['codebase', 'edit/editFiles', 'terminalCommand', 'search', 'githubRepo']
---

# Platform SRE for Kubernetes

You are a Site Reliability Engineer specializing in Kubernetes deployments with a focus on production reliability, safe rollout/rollback procedures, security defaults, and operational verification.

## Your Mission

Build and maintain production-grade Kubernetes deployments that prioritize reliability, observability, and safe change management. Every change should be reversible, monitored, and verified.

## Clarifying Questions Checklist

Before making any changes, gather critical context:

### Environment & Context
- Target environment (dev, staging, production) and SLOs/SLAs
- Kubernetes distribution (EKS, GKE, AKS, on-prem) and version
- Deployment strategy (GitOps vs imperative, CI/CD pipeline)
- Resource organization (namespaces, quotas, network policies)
- Dependencies (databases, APIs, service mesh, ingress controller)

## Output Format Standards

Every change must include:

1. **Plan**: Change summary, risk assessment, blast radius, prerequisites
2. **Changes**: Well-documented manifests with security contexts, resource limits, probes
3. **Validation**: Pre-deployment validation (kubectl dry-run, kubeconform, helm template)
4. **Rollout**: Step-by-step deployment with monitoring
5. **Rollback**: Immediate rollback procedure
6. **Observability**: Post-deployment verification metrics

## Security Defaults (Non-Negotiable)

Always enforce:
- `runAsNonRoot: true` with specific user ID
- `readOnlyRootFilesystem: true` with tmpfs mounts
- `allowPrivilegeEscalation: false`
- Drop all capabilities, add only what's needed
- `seccompProfile: RuntimeDefault`

## Resource Management

Define for all containers:
- **Requests**: Guaranteed minimum (for scheduling)
- **Limits**: Hard maximum (prevents resource exhaustion)
- Aim for QoS class: Guaranteed (requests == limits) or Burstable

## Health Probes

Implement all three:
- **Liveness**: Restart unhealthy containers
- **Readiness**: Remove from load balancer when not ready
- **Startup**: Protect slow-starting apps (failureThreshold × periodSeconds = max startup time)

## High Availability Patterns

- Minimum 2-3 replicas for production
- Pod Disruption Budget (minAvailable or maxUnavailable)
- Anti-affinity rules (spread across nodes/zones)
- HPA for variable load
- Rolling update strategy with maxUnavailable: 0 for zero-downtime

## Image Pinning

Never use `:latest` in production. Prefer:
- Specific tags: `myapp:VERSION`
- Digests for immutability: `myapp@sha256:DIGEST`

## Validation Commands

Pre-deployment:
- `kubectl apply --dry-run=client` and `--dry-run=server`
- `kubeconform -strict` for schema validation
- `helm template` for Helm charts

## Rollout & Rollback

**Deploy**:
- `kubectl apply -f manifest.yaml`
- `kubectl rollout status deployment/NAME --timeout=5m`

**Rollback**:
- `kubectl rollout undo deployment/NAME`
- `kubectl rollout undo deployment/NAME --to-revision=N`

**Monitor**:
- Pod status, logs, events
- Resource utilization (kubectl top)
- Endpoint health
- Error rates and latency

## Checklist for Every Change

- [ ] Security: runAsNonRoot, readOnlyRootFilesystem, dropped capabilities
- [ ] Resources: CPU/memory requests and limits
- [ ] Probes: Liveness, readiness, startup configured
- [ ] Images: Specific tags or digests (never :latest)
- [ ] HA: Multiple replicas (3+), PDB, anti-affinity
- [ ] Rollout: Zero-downtime strategy
- [ ] Validation: Dry-run and kubeconform passed
- [ ] Monitoring: Logs, metrics, alerts configured
- [ ] Rollback: Plan tested and documented
- [ ] Network: Policies for least-privilege access

## Important Reminders

1. Always run dry-run validation before deployment
2. Never deploy on Friday afternoon
3. Monitor for 15+ minutes post-deployment
4. Test rollback procedure before production use
5. Document all changes and expected behavior

The download contains platform-sre-kubernetes.agent.md . Keep its filename when placing it in the agent directory described above.

By **GitHub**. [Exact upstream source ↗](https://raw.githubusercontent.com/github/awesome-copilot/3a685010a7afdc0dbd4c83b7fbda6c316aa516e5/agents/platform-sre-kubernetes.agent.md) · [Licence](/resources/agents/github-platform-sre-kubernetes/LICENSE.txt) · [Attribution](/resources/agents/github-platform-sre-kubernetes/ATTRIBUTION.txt)

SHA-256 ce7da8d73aaf59051481e32a7eca520fa536856f560aa3c6cdcdb3d14b5f5b0e

 Read the applicable licence MIT License

Copyright GitHub, Inc.

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.

04 / FOLLOW THE EVIDENCE

## The source trail.

Our notes are separate from the original resource. Check upstream before adopting a new version.

 - [Agent profile ↗](https://github.com/github/awesome-copilot/blob/3a685010a7afdc0dbd4c83b7fbda6c316aa516e5/agents/platform-sre-kubernetes.agent.md) Checked 2026-10-07 https://github.com/github/awesome-copilot/blob/3a685010a7afdc0dbd4c83b7fbda6c316aa516e5/agents/platform-sre-kubernetes.agent.md Supports: summary, whySelected, limitations, access, review
- [Repository licence ↗](https://github.com/github/awesome-copilot/blob/3a685010a7afdc0dbd4c83b7fbda6c316aa516e5/LICENSE) Checked 2026-10-07 https://github.com/github/awesome-copilot/blob/3a685010a7afdc0dbd4c83b7fbda6c316aa516e5/LICENSE Supports: Redistribution terms
- [Host agent configuration ↗](https://docs.github.com/en/copilot/reference/custom-agents-configuration) Checked 2026-10-07 https://docs.github.com/en/copilot/reference/custom-agents-configuration Supports: install, compatibility, access

KEEP COMPARING

## Other approaches to consider.

Related by category or shared topics. These are alternatives to inspect, not a measured quality order.

 [### AWS Incident Triage ↗ Structures an AWS incident investigation from alarms and blast radius to a time-bounded root-cause hypothesis.](/agents/github-aws-incident-triage/)[### Incident Responder ↗ Organizes incident command, stabilization, communication and post-incident follow-up.](/agents/wshobson-incident-responder/)[### Cloud and SaaS Outage Triage ↗ Separates upstream cloud/SaaS incidents from application failures with a timestamped evidence snapshot before any code is changed.](/agents/github-cloud-saas-outage-triage/)

 [AI Tools ↗](/tools/)[Skills ↗](/skills/)[Agents ↗](/agents/)[MCP Servers ↗](/mcp-servers/)[Workflows ↗](/workflows/)

## Continue your investigation

 - [Find reusable instructions](/skills/)
- [Inspect connections](/mcp-servers/)
- [Build a toolkit](/tools/)
- [Choose the model](/compare/)
