Overview
Most Kubernetes clusters accumulate risk quietly. Namespaces get created without RBAC boundaries, pods run without resource limits, and network policies are "on the roadmap" for eighteen months straight. Nothing breaks — until it does, usually at the worst possible time.
This audit is a structured, hands-on review of your cluster against the same checklist I use when I onboard any new client's infrastructure. You get a clear picture of where you actually stand, not a generic vendor scorecard.
What's Included
- Security posture assessment — RBAC bindings, service account scoping, Pod Security Standards, and secrets handling
- Resource optimization review — requests/limits sanity check, over- and under-provisioned workloads, HPA/VPA configuration
- Network policy audit — default-deny posture, cross-namespace exposure, ingress/egress rules
- Ingress & TLS review — certificate management, rate limiting, and exposed endpoints
- Availability, backups & recovery review — disruption budgets, replica configuration, backup coverage, and recovery readiness
- Observability & cluster configuration review — critical signals, alerting coverage, control-plane and add-on configuration where accessible
- A professional written report ranking every finding by severity, with concrete remediation steps — not just "this is bad"
Our Process
- Access & scoping call (30 min) — I get read access to the cluster and confirm which namespaces and workloads are in scope.
- Automated + manual review (2-3 days) — I run the audit checklist against your cluster, cross-checking automated tool output by hand.
- Findings report — a written document ranking issues by severity (critical / high / medium / low), each with a specific fix.
- Walkthrough call — we go through the findings together, and I answer questions about prioritization and effort.
What You Receive
A professional written assessment with findings categorized as critical, high, medium, or low. Every finding explains the risk, affected area, and a practical remediation recommendation so your team can prioritize work confidently. See the sample Kubernetes audit report for an anonymized example of the format.
Access & Delivery
Read-only access is preferred wherever possible. The audit is delivered in 3–5 business days and includes a live walkthrough of the findings and recommended next steps.
Kubernetes Production Readiness Framework
The audit uses a practical Kubernetes Production Readiness Framework: twelve areas that help teams evaluate whether a cluster is secure, operable, and ready to support production workloads. The depth of each area depends on the agreed scope and available access.
- SecurityCluster hardening and workload exposure
- Identity & RBACAccess boundaries and service accounts
- NetworkingConnectivity, segmentation, and NetworkPolicy
- WorkloadsConfiguration and deployment safety
- Resource ManagementRequests, limits, and capacity signals
- AvailabilityReplicas, disruption handling, and failure modes
- Ingress & TLSPublic entry points and certificate handling
- SecretsStorage, access, and workload consumption
- ObservabilityMetrics, logs, alerts, and operational visibility
- Backup & RecoveryCoverage and recovery readiness
- CI/CD & GitOpsDeployment controls, drift, and rollback paths
- Operational ReadinessRunbooks, ownership, and day-two operations
For focused follow-on work, see Kubernetes Consulting, Kubernetes Security Consulting, and Kubernetes Observability & Monitoring.
What Happens After the Audit?
Who This Is For
Teams running production workloads on EKS, GKE, AKS, or self-managed Kubernetes who have never had a formal security or resource review — or haven't had one in over a year. Also a good fit before a compliance audit, after a security incident, or when onboarding a new platform team.
FAQ
Do you need production access? Read-only access is sufficient for almost everything in the audit. I'll flag the few checks that need broader access up front.
What if you find something critical? Critical findings are flagged and communicated immediately — I don't wait for the final report to tell you your cluster is wide open to the internet.
Can you also fix what you find? Yes. The audit stands alone, but most clients follow up with a scoped remediation engagement based on the report. There's no obligation either way.