47 clusters. €92,000/month in cloud spend. A team on the brink of burnout.
I recently read a compelling article about a DevOps team that ditched Kubernetes after years of struggling with its complexity. Their deployment success jumped 89%. Infrastructure costs dropped 62%. Team morale transformed overnight.
Their story resonated because I've lived both sides of it. I've seen teams abandon Kubernetes in frustration. I've also optimized poorly-managed clusters to save 30% monthly while improving reliability. The difference? Understanding why Kubernetes fails — and addressing root causes instead of symptoms.
The Pattern Behind Kubernetes Failures
When teams abandon Kubernetes, the narrative is always similar: too complex, too expensive, team exhausted. But when you dig into specifics, a different pattern emerges.
The team that ran 47 clusters wasn't drowning because of Kubernetes. They were drowning because of architectural decisions that would have created problems on any platform:
- Cluster proliferation instead of namespace isolation - Running separate clusters for every service instead of proper multi-tenancy
- Multi-cloud without justification - Distributing across AWS, GCP, and Azure without clear business requirements
- No consolidation strategy - Allowing clusters to multiply as services grew instead of rightsizing
- Over-engineering from day one - Implementing enterprise patterns before understanding basic operations
The same team spending €23,000/month on control planes alone could have run the same workloads on 3-5 well-architected clusters with proper namespace isolation.
This isn't a Kubernetes problem. It's an adoption problem.
The Hidden Cost of "Simplifying"
Moving away from Kubernetes feels like relief — until you hit the ceiling of simpler alternatives.
What you gain immediately:
- Faster deployments with less configuration
- Lower cognitive overhead for junior engineers
- Reduced monitoring complexity
- Smaller operational surface area
What you lose over time:
- Portability - Lock-in to cloud-specific services (ECS, Fargate, Lambda)
- Advanced orchestration - Self-healing, rolling updates, sophisticated scheduling
- Declarative infrastructure - GitOps workflows with full audit trails
- Multi-tenancy - Namespace-based isolation without environment duplication
- Unified networking - Service meshes, network policies, ingress control
The migration savings are real. But they come with technical debt that compounds as your architecture grows.
When Kubernetes Is the Wrong Choice
Let's be clear: Kubernetes isn't for everyone. The team that moved to ECS made a rational decision for their context.
You probably don't need Kubernetes if:
- You're running fewer than 10-15 services
- Your traffic patterns are predictable and don't require complex autoscaling
- Your team is small (< 5 DevOps engineers) and lacks orchestration expertise
- You're primarily using managed services (databases, queues, caching)
- Your deployment patterns are simple (blue-green or basic rolling updates)
- You don't need multi-cloud or hybrid cloud capabilities
For these scenarios, simpler alternatives win:
- AWS ECS/Fargate - Managed container orchestration without control plane overhead
- Google Cloud Run - Serverless containers with auto-scaling
- Scalingo/Platform.sh - European PaaS providers with GDPR compliance built-in
- AWS Lambda/Cloud Functions - Event-driven, serverless workloads
- Hetzner Cloud - Cost-effective European infrastructure for simple container workloads
There's no shame in choosing simplicity when it matches your requirements.
When Kubernetes Is the Right Choice
But if your architecture looks like this, walking away from Kubernetes creates different problems:
You need Kubernetes if:
- GDPR and data sovereignty matter - Keep data within EU boundaries with multi-region deployments
- Multi-cloud is a business requirement - Regulatory compliance, disaster recovery, or negotiating leverage
- Advanced networking is essential - Zero-trust architecture, service mesh, sophisticated traffic routing
- You're running 20+ microservices - Namespace isolation beats duplicating environments
- Compliance demands audit trails - GitOps with declarative manifests provides paper trail for ISO 27001, NIS2
- Your team has orchestration expertise - The learning curve is already paid
- You need sophisticated scheduling - Resource quotas, pod affinity, priority classes
- Self-healing is critical - Automatic restart, rescheduling, health checks
For distributed systems at scale, alternatives require building custom tooling that replicates what Kubernetes provides.
How to Succeed With Kubernetes
The teams that win with Kubernetes follow patterns that prevent the problems others encounter.
1. Start with Managed Control Planes
Don't run your own control plane. Use EKS, AKS, or GKE.
The team spending €23,000/month on 47 control planes could have run managed Kubernetes for €450-€1,400/month total with proper consolidation.
Benefits of managed Kubernetes:
- Control plane upgrades handled by cloud provider
- Built-in high availability and disaster recovery
- Enterprise support included (critical for NIS2 compliance)
- Reduced operational burden on your team
- European data centres available (Frankfurt, Paris, Amsterdam, Stockholm)
2. Design for Multi-Tenancy, Not Cluster Proliferation
One cluster per environment, not per service.
Use namespace isolation with resource quotas and RBAC:
apiVersion: v1
kind: Namespace
metadata:
name: team-alpha
---
apiVersion: v1
kind: ResourceQuota
metadata:
name: compute-quota
namespace: team-alpha
spec:
hard:
requests.cpu: "20"
requests.memory: 40Gi
limits.cpu: "40"
limits.memory: 80Gi
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: team-alpha-admin
namespace: team-alpha
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: admin
subjects:
- kind: Group
name: team-alpha
apiGroup: rbac.authorization.k8s.io
Result: Isolated teams sharing infrastructure without duplication.
3. Implement Observability From Day One
The team that abandoned Kubernetes ran five monitoring tools with 147 false positive alerts. That's not observability — that's alert fatigue.
Start with the essentials:
# Deploy Prometheus + Grafana on day one
helm install prometheus prometheus-community/kube-prometheus-stack \
--set prometheus.prometheusSpec.retention=30d \
--set alertmanager.enabled=true \
--set grafana.enabled=true
Focus on metrics that matter:
- Resource utilization (CPU, memory, disk)
- Pod restart frequency
- Deployment success rate
- Request latency and error rates
- Cost per namespace
Consolidate tooling: One logging solution (Loki or CloudWatch), one metrics solution (Prometheus), one tracing solution (Tempo or Jaeger).
4. Automate Cost Optimization
The €16,500 cluster I optimized to €11,500 came down to three changes:
Rightsized instance types:
# Before: m5.4xlarge (16 vCPU, 64GB RAM)
# After: m5.xlarge (4 vCPU, 16GB RAM)
# Based on actual utilization data
Implemented autoscaling:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: api-server
minReplicas: 3
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
Set resource limits:
resources:
requests:
cpu: 200m
memory: 256Mi
limits:
cpu: 500m
memory: 512Mi
Deploy cost monitoring:
# Kubecost for real-time cost allocation (supports EUR pricing)
helm install kubecost kubecost/cost-analyzer \
--set kubecostToken="your-token-here" \
--set prometheus.server.global.external_labels.currency="EUR"
5. Invest in Team Training
The learning curve is real. Engineers need structured training, not "figure it out."
Our approach:
- CKAD certification for all platform engineers (3-month timeline)
- Monthly Kubernetes office hours for Q&A and troubleshooting
- Internal runbook wiki documenting common operations
- Buddy system pairing new hires with experienced engineers
- KubeCon Europe attendance budget (Paris, Amsterdam, Valencia rotations)
- Local meetups - Kubernetes community groups in Berlin, London, Stockholm, Zurich
ROI: Trained engineers build reliable systems. Untrained engineers create the complexity that drives teams away.
6. Adopt GitOps for Infrastructure as Code
Every change should be declarative, reviewable, and reversible.
# ArgoCD Application manifest
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: production-api
spec:
project: production
source:
repoURL: https://github.com/company/k8s-manifests
targetRevision: main
path: apps/api
destination:
server: https://kubernetes.default.svc
namespace: production
syncPolicy:
automated:
prune: true
selfHeal: true
Benefits:
- Every change has a Git commit (compliance audit trail)
- Rollbacks are
git revertcommands - Disaster recovery is
kubectl apply -f repo/ - No configuration drift between environments
The Real Choice: Right-Sizing vs. Walking Away
The team that moved from 47 Kubernetes clusters to ECS made a smart decision for their situation. But their problems were solvable without abandoning orchestration.
If they had:
- Consolidated to 5 clusters with namespace isolation
- Used managed Kubernetes (EKS/GKE)
- Implemented proper monitoring and cost controls
- Invested in team training
- Adopted GitOps for declarative management
They would have achieved:
- Similar cost savings (50-60% reduction)
- Maintained portability and flexibility
- Avoided cloud lock-in
- Preserved orchestration capabilities
- Built team expertise instead of abandoning it
What I'd Tell Teams Considering Kubernetes
Start with these questions:
- Do we need orchestration? If you have < 10 services, probably not.
- Do we have the expertise? If not, can we invest in training?
- What's our scale trajectory? Growing fast or relatively stable?
- What are our compliance requirements? Audit trails, multi-cloud, zero-trust?
- What's our team size? Small teams benefit from simplicity; larger teams need structure.
If you're already using Kubernetes and struggling:
- Audit your architecture - Are you over-engineered? (47 clusters is a red flag)
- Measure actual costs - Break down control plane, compute, networking, and human time
- Assess team capability - Is the problem Kubernetes or lack of training?
- Look for quick wins - Consolidate clusters, implement autoscaling, set resource limits
- Consider alternatives - But understand what you're trading away
Key Takeaways
- Kubernetes failures are usually adoption failures, not platform failures. Poor architecture, lack of training, and over-engineering create the pain points.
- The complexity exists to solve real problems. Simpler alternatives work until they don't — and migration back is expensive.
- For European companies, Kubernetes provides critical GDPR and NIS2 compliance capabilities. Multi-region deployment within EU boundaries and declarative audit trails matter.
- Success requires investment in monitoring, training, and operational discipline. These aren't optional.
- There's no shame in choosing simpler tools when they match your needs. Right-sizing beats over-engineering.
- The teams that win with Kubernetes consolidate infrastructure, automate cost controls, and invest in expertise. It's an operational playbook, not magic.
Different problems require different tools. For small teams with simple workloads, European PaaS providers like Scalingo or Hetzner Cloud make sense. For distributed systems at scale with multi-cloud and data sovereignty requirements, Kubernetes provides capabilities that alternatives can't match.
The question isn't "Kubernetes vs. no Kubernetes." It's "What problems am I solving, what are my compliance requirements, and what's the right tool for my context?"
About the author: Henrik Jess is a DevOps & MLOps engineer who has seen both successful Kubernetes implementations and expensive mistakes. This article reflects lessons from teams on both sides of the orchestration debate.