Two weeks. One monitoring dashboard. A €5,000/month reduction.
That's why I'm not ditching Kubernetes — and why the recent wave of "we're happier without K8s" articles miss the point entirely.
I recently read "I Stopped Using Kubernetes. Our DevOps Team Is Happier Than Ever." The author abandoned K8s to simplify operations and cut costs. I respect that decision for their context. But after years running Kubernetes in production across multiple organizations, I've learned something different: the problem isn't Kubernetes — it's how we adopt it.
The €16,500 Problem (And How I Fixed It)
Let me start with the case that changed my perspective on Kubernetes costs.
I inherited a K8s cluster on AWS EKS burning €16,500/month. The team was convinced Kubernetes was "just expensive" and that simpler alternatives would save money. Before recommending any platform change, I spent two weeks analyzing what we actually had.
What I found:
The cluster wasn't expensive because of Kubernetes. It was expensive because of poor operational practices:
- Over-provisioned nodes: Running m5.4xlarge instances (16 vCPU, 64GB RAM) when workloads needed m5.xlarge (4 vCPU, 16GB RAM). Nobody had profiled actual usage.
- No resource limits: Pods could request unlimited CPU/memory, causing massive over-allocation.
- Zombie workloads: 12 deployments from old experiments running 24/7, consuming resources nobody used.
- No autoscaling: The cluster ran 20 nodes at 3 AM when traffic was 5% of peak.
The fix I implemented:
First, I deployed monitoring that mattered — not vanity dashboards, but real resource consumption data:
# Prometheus with kube-state-metrics
helm install prometheus prometheus-community/kube-prometheus-stack \
--set prometheus.prometheusSpec.retention=30d \
--set grafana.enabled=true
Two weeks of metrics revealed the truth:
- Average CPU utilization: 23%
- Average memory utilization: 31%
- Peak traffic hours: 9 AM - 6 PM weekdays
Then I implemented changes over a two-week period, starting with the least risky:
Resource quotas per namespace:
apiVersion: v1
kind: ResourceQuota
metadata:
name: compute-quota
spec:
hard:
requests.cpu: "50"
requests.memory: 100Gi
limits.cpu: "100"
limits.memory: 200Gi
Horizontal Pod Autoscaling:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: app-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: main-app
minReplicas: 3
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
High Availability with PodDisruptionBudgets:
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: main-app-pdb
spec:
minAvailable: 2
selector:
matchLabels:
app: main-app
Cluster Autoscaler configuration:
apiVersion: v1
kind: ConfigMap
metadata:
name: cluster-autoscaler-config
data:
scale-down-delay-after-add: "10m"
scale-down-unneeded-time: "10m"
skip-nodes-with-local-storage: "false"
balance-similar-node-groups: "true"
With these changes:
- Switched to m5.xlarge instances
- Set min nodes to 5, max to 25
- Enabled scale-down during off-peak hours
The results after 30 days:
- Monthly cost: €16,500 → €11,500 (30% reduction)
- Average cluster utilization: 23% → 68%
- Zero performance impact (actually improved p95 latency by 15ms)
- Bonus: Autoscaling handled a Black Friday traffic spike without manual intervention
ROI including labor: Even counting two weeks of engineering time, the changes paid for themselves within the first billing cycle.
The team didn't have a "Kubernetes cost problem." They had an operational maturity problem.
Tools that made this possible:
- Kubecost - Real-time cost allocation per namespace/deployment
- Prometheus + Grafana - Historical metrics and capacity planning
- AWS Cost Explorer - Validating spend against Kubernetes metrics
Why Teams Fail at Kubernetes (It's Not the Platform)
The pattern I've seen across multiple organizations is consistent: teams abandon Kubernetes not because the platform fails, but because they adopt it wrong.
Common failure modes:
- Jumping in without expertise - Treating K8s like a simple PaaS instead of distributed systems infrastructure
- Skipping monitoring - Flying blind without observability leads to resource waste
- No training investment - Expecting engineers to figure out pod networking, service meshes, and CRDs on their own
- Over-engineering from day one - Starting with multi-cluster, multi-region before mastering basics
The complexity exists because Kubernetes solves genuinely hard problems. When you're orchestrating hundreds or thousands of containers across multiple environments with auto-scaling, self-healing, and zero-downtime deployments, you need sophisticated tooling.
Simpler alternatives work — until they don't. And when you hit that ceiling, migration becomes exponentially more painful than starting with the right foundation.
When Kubernetes Makes Sense (And When It Doesn't)
Kubernetes is the right choice when you need:
- Multi-cloud portability - Same manifests across AWS, GCP, Azure, or on-premises
- Advanced networking - Service meshes, network policies, ingress controllers
- Declarative infrastructure - GitOps workflows with full audit trails
- High availability - Self-healing, rolling updates, pod disruption budgets
- Team isolation - Namespace-based multi-tenancy with resource quotas and RBAC
Skip Kubernetes if you have:
- Simple, stateless web apps - Managed container services (ECS, Cloud Run) are simpler
- Serverless-friendly workloads - Event-driven functions don't need orchestration
- Small teams without ops expertise - Training investment might outweigh benefits
- Low traffic volume - Over-engineering for problems you don't have yet
I'm not advocating Kubernetes everywhere. I'm advocating informed decisions based on actual requirements, not reactions to complexity.
How to Win With Kubernetes: The Operational Playbook
Success with Kubernetes comes down to four pillars:
1. Observability From Day One
Monitoring isn't optional — it's the foundation of cost control and reliability.
The stack I deploy first:
- Prometheus - Metrics collection and alerting
- Grafana - Visualization and dashboards
- Loki - Log aggregation (EFK alternative that's lighter weight)
- Tempo - Distributed tracing for debugging microservices
Together, these turn incidents into insights. When something breaks, you have data to understand why.
2. GitOps as the Source of Truth
Every change should be declarative, reviewable, and reversible. This is where Kubernetes becomes infrastructure-as-code at scale.
Our GitOps workflow:
# ArgoCD Application manifest
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: main-app
spec:
project: production
source:
repoURL: https://github.com/company/k8s-manifests
targetRevision: main
path: apps/main-app
destination:
server: https://kubernetes.default.svc
namespace: production
syncPolicy:
automated:
prune: true
selfHeal: true
Benefits:
- Every change has a Git commit (audit trail)
- Rollbacks are
git revertcommands - Disaster recovery is
kubectl apply -f repo/ - Compliance teams love the paper trail
3. Compliance & Security by Design
Beyond scalability, Kubernetes enforces compliance discipline. Pod Security Standards, NetworkPolicies, and RBAC allow fine-grained separation that maps directly to ISO 27001 and GDPR Article 32 requirements.
Example NetworkPolicy for zero-trust networking:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: api-to-db-only
spec:
podSelector:
matchLabels:
app: database
policyTypes:
- Ingress
ingress:
- from:
- podSelector:
matchLabels:
app: api-server
ports:
- protocol: TCP
port: 5432
This level of control is difficult or impossible in simpler platforms. For regulated industries, it's non-negotiable.
4. Continuous Training Investment
The learning curve is real. Engineers need time to understand pod networking, storage classes, operators, and service meshes.
Our approach:
- Monthly "Kubernetes office hours" for Q&A
- Dedicated Slack channel for sharing solutions
- Budget for KubeCon attendance and CKAD/CKA certifications
- Internal documentation wiki with runbooks
The upfront investment pays exponential returns. Trained teams don't just "use" Kubernetes — they leverage it.
The Advantages Nobody Talks About
Multi-Cloud Isn't Just About Avoiding Lock-In
It's about disaster recovery and negotiating leverage. When your infrastructure is portable, you have options:
- DR across regions: Same manifests, different clusters
- Cost negotiation: "We can move to GCP" isn't a bluff
- Hybrid cloud: On-prem for compliance, cloud for burst capacity
I've seen teams save 20% on cloud contracts just by demonstrating multi-cloud capability.
Disaster Recovery Becomes Trivial
Kubernetes' declarative model means environments are reproducible:
- GitOps repo holds all manifests
- Velero backs up cluster state and volumes
- New cluster +
kubectl apply= recovered environment
Compare this to imperative alternatives where rebuilding from scratch requires tribal knowledge and manual steps.
Multi-Tenancy Without Duplication
Namespace isolation + resource quotas + RBAC = shared infrastructure for multiple teams.
Alternatives force full environment duplication (separate ECS clusters, separate Lambda accounts), multiplying costs and management overhead.
What I'd Do Differently (Lessons Learned)
Start with managed Kubernetes, not self-hosted: EKS, AKS, and GKE handle control plane complexity. Don't build what you can buy.
Automate observability from day one: Don't wait until you have cost or performance problems. Deploy Prometheus on day zero.
Invest in training immediately: The two-week "figure it out" approach wastes more time than structured learning.
Document everything as code: Not just manifests, but runbooks, architecture decisions, and troubleshooting guides in your Git repo.
Start small and scale up: Run non-critical workloads first. Learn the platform before migrating production.
Key Takeaways
- Kubernetes costs aren't inherent — they're operational. Monitoring and rightsizing save 30%+ immediately.
- The complexity exists to solve hard problems. Simpler tools hit ceilings that Kubernetes doesn't.
- Success requires investment in training, automation, and observability — not optional "nice-to-haves."
- GitOps + Kubernetes creates reproducible, auditable, compliant infrastructure that alternatives can't match.
- Multi-cloud capability provides resilience and negotiating leverage, not just "lock-in avoidance."
I haven't stopped using Kubernetes because, with the right strategy, it's an investment that pays off. It's transformed how I approach infrastructure and enabled my teams to build systems that meet modern demands for scale, reliability, and compliance.
Different tools work for different problems. But for distributed systems at scale, alternatives haven't replicated what Kubernetes provides.
About the author: Henrik Jess is a DevOps & MLOps engineer with experience running Kubernetes in production across multiple clouds and industries. This reflects lessons learned from both successful implementations and expensive mistakes.