12 conflicts. 2 teams doing their jobs correctly. 1 missing actor: Management.
During a Kubernetes cluster rebuild, 12 distinct technical conflicts emerged between a platform team and a DevOps team. ArgoCD debugging limitations. Kubectl access removed. Namespace collisions. Ingress conflicts. Support boundary disputes.
The pattern was clear: two competent teams, locked in friction, with no one above them aligning expectations.
I've watched this pattern repeat across technology shifts—VMware migrations in the 2000s, cloud migrations in the 2010s, and now container orchestration. The technology changes. The management failure doesn't.
Here's the uncomfortable truth: Platform engineering conflicts aren't technical problems. They're organizational gaps dressed in technical clothing.
The 12 Conflicts (What Happened)
Let's start with the facts. Here's what actually went wrong during the cluster rebuild.
During the migration, these conflicts emerged:
Access & Debugging (The Tool Wars)
1. Kubectl access removed
Access to cluster API (kubectl/k9s) was removed during the new cluster build. No alignment discussion happened beforehand. DevOps team's debugging workflow broke overnight.
2. Debugging limitations
ArgoCD became the mandated debugging tool with no kubectl fallback. ArgoCD doesn't show complete error messages in many cases. Debugging time increased significantly without terminal access.
3. ACR push access restricted
Push access to Azure Container Registry was removed. This blocked testing and deployment of images outside pipelines. Access was restored only after explicit escalation.
Configuration & Architecture (The Standards Fight)
4. Namespace and Helm conflicts
Helm charts contained namespace definitions, which conflicted with Argo's namespace management approach. ServiceAccounts were created multiple times, causing sync failures.
5. Argo out-of-sync loops
Applications terminated and recreated in loops due to ApplicationSet naming collisions. Argo showed generic "out of sync" without details. Root causes were difficult to trace without cluster access.
6. Ingress host/path duplicates
Multiple namespaces contained ingresses with identical host/path combinations. NGINX admission controller rejected these silently. Resources couldn't sync until conflicts were manually identified.
7. Azure App Configuration Provider issues
Multiple issues emerged: missing Helm installation and incorrect values.yaml, incompatible manifest fields causing automatic removals, immutable selector errors when changing Deployment via nameOverride, and missing or incorrect federated credentials for workload identity. No clear documentation existed for ArgoCD integration.
8. Configuration management debate
Disagreement emerged over Azure App Configuration vs. External Secrets/Key Vault. DevOps team needed GUI for developer self-service. Platform team considered App Configuration "too Azure-specific."
Ownership & Support (The Accountability Gap)
9. Unclear responsibility boundaries
When ApplicationSets diverged from platform standards: "Not our problem." DevOps team was still expected to deliver services, but without full access. No escalation path existed when conflicts arose.
10. Support model gaps
Support was principle-bound ("use Argo only"). Complex errors required manual, reactive intervention. No defined SLAs or response time expectations existed.
11. Expectation mismatch on rebuild
Test cluster was torn down and rebuilt. DevOps team expected same access model as before. Platform team implemented new restrictions without discussion.
12. Governance vs. operational needs
Platform team optimized for control, standardization, and security. DevOps team needed flexibility, access, and self-service debugging. These goals were never aligned or prioritized by management.
Summary: Twelve conflicts, all rooted in misaligned expectations. But the real question is: why did these conflicts emerge in the first place?
You're Not Alone (Industry Validation)
Before diving into root causes, let's validate something important.
These aren't unique problems. This exact pattern plays out across the industry.
ArgoCD Debugging Is a Known Pain Point
ArgoCD excels at deployment but struggles with operational debugging. The platform team's "use ArgoCD only" mandate ran into a well-documented tool limitation.
This isn't just one organization's experience. GitHub issues about ArgoCD debugging visibility and dependency tracking have hundreds of reactions and years of community discussion. The pattern is clear across the industry.
"One Tool to Rule Them All" Doesn't Work
Tool standardization hits practical limits. Teams need multiple tools for different contexts—kubectl for debugging, Terraform for infrastructure, Helm for packaging. Mandating a single tool often drives workarounds rather than compliance.
The pattern shows up repeatedly in developer discussions: platform teams over-index on perfect automation for infrastructure, under-index on practical developer experience.
The Industry Is Having the Wrong Conversation
Platform engineering discussions focus overwhelmingly on tools, architectures, and technical patterns. Search industry forums for "platform team conflict resolution" or "organizational design" and you'll find remarkably little.
Even YC-backed startups building internal developer platforms position the problem as technical: "Kubernetes for both devs and ops." The market recognizes the friction. But the conversation stays technical.
Almost no one is talking about management's role in platform conflicts.
The Missing Actor (What Should Have Happened)
Now let's identify the real problem.
When you analyze these 12 conflicts, something becomes obvious: two technical teams doing their jobs correctly, locked in conflict because no one above them aligned expectations.
Here's what actually happened:
- Platform team: "Implement secure, standardized Kubernetes with GitOps best practices"
- DevOps team: "Get 50+ applications running with minimal disruption"
- Management: [cricket sounds]
The Pattern Repeats
This isn't unique to Kubernetes. The same dynamic plays out across technology shifts:
- VMware migrations in the 2000s
- Cloud migrations in the 2010s
- Container orchestration in the 2020s
The technology changes. The management failure doesn't.
What Management Didn't Do (But Should Have)
Three critical conversations never happened. Here's what was missing.
Before the Cluster Rebuild
Missing Conversation #1: Access Model Alignment
What happened:
- Platform team decided: "ArgoCD only, no kubectl"
- DevOps team expected: "Same access as before"
- Management never asked: "What are we trading for this security?"
The framework that was missing:
Decision: Remove kubectl access
Cost: Significantly longer debugging time, 3-day incidents become 1-week+ incidents
Benefit: Reduced security surface, cleaner audit logs
Trade-off: Acceptable? To whom?
Exception process: What happens when ArgoCD can't solve it?
Management should have facilitated this discussion. They didn't.
Missing Conversation #2: Ownership Boundaries
When DevOps team's ApplicationSets diverged from platform standards, a critical question emerged: who owns what?
The platform team had built templates for "standard" deployments. But real applications rarely fit perfectly into standard templates.
When the DevOps team needed custom configurations—different namespace structures, specific Helm values, unique ServiceAccount setups—the platform team drew a line.
What happened:
- Platform team: "Not our problem, doesn't match our template"
- DevOps team: "Still need to get it working"
- Management: [absent]
No one had defined where platform responsibility ended and application team responsibility began.
No one had created a process for handling edge cases.
No escalation path existed when teams disagreed about what was "reasonable variance" versus "unsupported customization."
The framework that was missing:
Standard Config: What platform team supports fully
Supported Variance: What platform team helps with
Unsupported (Self-Service): What teams own completely
Escalation Path: When conflicts arise, who decides?
These boundaries needed management definition. They were never set.
Missing Conversation #3: Support Model & SLAs
The most painful conflicts emerged when things broke.
The platform team had mandated ArgoCD as the primary debugging interface, but ArgoCD's visibility has known limitations. When applications failed to sync, the UI often showed generic "out of sync" messages without exposing the underlying Kubernetes errors.
DevOps engineers needed deeper access to diagnose issues—kubectl access to inspect resources, check logs, describe pod states, and understand what the admission controller was rejecting.
Real scenario from the conflict list:
- Argo shows "out of sync"
- ArgoCD UI doesn't reveal root cause
- Platform team: "Use Argo to debug"
- DevOps team: Blocked for days
No one had defined what level of support the platform team would provide.
No response time expectations existed. No process for temporary exception access during incidents.
No escalation path when the platform team's recommended approach wasn't working.
The DevOps team was left in limbo—responsible for fixing production issues but without the tools to diagnose them.
The framework that was missing:
Severity 1 (Production Down):
- Access granted: kubectl read-only within 15 minutes
- Platform team response: 1 hour
- Escalation path: CTO after 4 hours
Severity 2 (Feature Blocked):
- Platform team response: 4 hours
- Joint debugging session: 24 hours
- Management mediation: If unresolved in 3 days
Management never defined support expectations. Both teams operated in a vacuum.
During the Conflicts
What Actually Happened:
- DevOps hits blocker (can't see Ingress conflict in ArgoCD)
- DevOps asks Platform for kubectl access
- Platform says "use ArgoCD" (following their mandate)
- DevOps works around it (manual inspection via cloud console)
- Multiple days lost
- Management never sees the cost
What Should Have Happened:
- DevOps escalates: "Blocked, ArgoCD insufficient, need resolution path"
- Management asks both teams: "What's needed to unblock?"
- Temporary exception granted with timeline
- Platform team: "We'll improve ArgoCD visibility for this case"
- Documented as known gap in support model
- Added to platform roadmap
The difference? Management participation.
The Costs Management Didn't See
Here's why this matters beyond team frustration.
Let me quantify what this conflict actually cost:
Visible Costs (Management Saw These)
- 2-week delay in cluster migration (team capacity)
- ApplicationSet rebuilds (platform team time)
- Direct infrastructure costs
Hidden Costs (Management Didn't See)
- Significantly extended debugging time across 12+ incidents (engineer time)
- Workarounds and shadow processes (ongoing tax)
- Knowledge gaps and communication overhead (repeated explanations)
- Team morale and trust damage (attrition risk)
- Reduced velocity on feature work (opportunity cost)
The Hidden Cost Multiplier
For every hour of visible platform friction, there are multiple hours of hidden productivity loss.
Management can't fix what they can't see.
Why Management Stayed Out
So why does this pattern repeat? Management stays out because of four common misconceptions.
1. "Let the Technical Experts Figure It Out"
Management thought: "We hired smart platform engineers. They'll work it out."
The flaw: Technical teams optimize for different goals. Platform team prioritizes security, standardization, and maintainability. DevOps team needs velocity, flexibility, and operational reality.
Without management aligning these goals, conflict is guaranteed.
2. "Platform Team Knows Best"
Management assumed platform team's technical authority = decision authority.
The flaw: Platform team can define "how" but not "what trade-offs are acceptable." For example: "ArgoCD only is more secure" is a technical fact. But "Is significantly slower debugging acceptable?" is a business decision.
Management delegated a business decision to a technical team.
3. "It's Just Temporary Friction"
Management thought: "Growing pains, they'll adapt."
The flaw: Patterns established during migration become permanent. Workarounds become standard process. Teams learn to avoid each other. Trust erodes incrementally. "Us vs. them" culture solidifies.
By the time management notices, the damage is structural.
4. "We Don't Understand Kubernetes Enough"
Management avoided mediating because "it's too technical."
The flaw: The conflicts aren't about Kubernetes. They're about decision authority (who owns what), resource trade-offs (security vs. speed), support expectations (who helps whom), and communication patterns (how do teams align).
These are management problems dressed in technical clothing.
The Framework Management Needed
Now for the practical part. Here's what actually works.
The solution isn't complicated—it's about creating the right conversations at the right time:
1. Pre-Migration: Alignment Workshop
Attendees: Platform team lead, DevOps team lead, Engineering Manager, CTO
Topics:
- Access model: What changes, what's the cost, what's the exception process?
- Ownership boundaries: Where does platform support end?
- Support model: Response times, escalation paths
- Success metrics: Shared goals both teams optimize for
Output: Written agreement, signed by all parties, visible to teams
Time Required: 4 hours Cost of NOT Doing This: Months of friction and reduced productivity
2. During Migration: Weekly Sync
Format: 30-minute standing meeting
Agenda:
- Blockers escalation (what needs management decision)
- Workarounds review (temporary solutions becoming permanent)
- Cost visibility (time lost to friction)
- Process adjustments (what's not working)
Output: Decision log, action items with owners
Management Role: Decision maker, not observer
3. Post-Migration: Retrospective
Questions:
- What conflicts emerged that we didn't anticipate?
- What cost us the most time?
- What would we do differently?
- What needs to change in our operating model?
Output: Updated platform team charter, revised support model
Purpose: Learn, don't blame
4. Ongoing: Shared Success Metrics
Wrong Metrics (Creates Conflict):
- Platform team: "100% ArgoCD adoption" vs. DevOps team: "Zero deployment blockers"
- Platform team: "Zero security exceptions" vs. DevOps team: "Feature velocity"
Right Metrics (Creates Alignment):
- Shared: "Mean time to production from code commit"
- Shared: "Deployment success rate"
- Shared: "Developer satisfaction score"
- Shared: "Platform adoption rate"
When both teams have the same scorecard, conflict becomes collaboration.
Summary: Four conversations at four stages—alignment workshop, weekly sync, retrospective, and shared metrics. Not complicated, but requires management to show up.
What Good Looks Like
Let's look at an organization that got this right.
Real Example: How Spotify Does Platform Engineering
Spotify's platform team (Golden Path) operates with this principle:
"The platform is a product. DevOps teams are our customers. If customers aren't using it, we failed."
This one sentence changes everything.
Their Model:
- Default path is easy: 80% of use cases "just work"
- Exceptions are supported: 15% of use cases get help
- Custom is allowed: 5% of use cases opt out entirely
- Platform team measures: Adoption rate, satisfaction, time-to-production
Management's Role:
- Platform team doesn't get to mandate adoption
- DevOps teams don't get to skip security reviews
- Both teams measured on shared outcomes
- Regular alignment meetings with decision authority
Result: 95%+ platform adoption, high developer satisfaction, minimal friction
The Key Difference
Spotify treats platform as a product to adopt, not a policy to enforce.
When teams opt out, platform asks: "What did we miss? How do we improve?"
Not: "They're non-compliant, that's their problem."
The Management Checklist
Ready to take action? Start here.
If you're a manager with platform and DevOps teams, use this checklist:
Before Platform Changes:
- Access model changes discussed with all affected teams?
- Trade-offs explicitly acknowledged (security vs. speed)?
- Exception process defined and communicated?
- Support model and SLAs documented?
- Ownership boundaries crystal clear?
- Shared success metrics defined?
During Conflicts:
- Escalation path used within 24 hours of blocker?
- Both sides heard before decision?
- Decision documented with rationale?
- Temporary exceptions time-boxed?
- Hidden costs tracked (debugging time, workarounds)?
After Migration:
- Retrospective conducted within 2 weeks?
- Lessons learned captured and shared?
- Operating model updated based on reality?
- Team relationships assessed (trust, communication)?
- Follow-up on unresolved friction points?
Ongoing:
- Weekly sync for first 3 months?
- Monthly review of shared metrics?
- Quarterly platform roadmap alignment?
- Annual team interaction model review?
Key Takeaways
-
Platform conflicts aren't technical failures - they're management gaps dressed in technical clothing
-
Two teams can both be right - when management doesn't align goals and boundaries
-
The cost of conflict is mostly hidden - visible friction is just the tip of the iceberg
-
Management can't delegate business decisions - "how secure vs. how fast" is a management call, not a platform team call
-
Silence is a decision - by not mediating, management chose platform team authority by default
-
Platform teams need product management - they're building for internal customers who can opt out (via shadow IT)
-
Shared metrics drive shared behavior - if teams have different scorecards, expect conflict
-
Exception processes are mandatory - "no exceptions" means "shadow workarounds"
The real lesson: Platform engineering is 20% technology selection, 80% organizational design.
Management can't outsource the 80%.
From the Ground Level
After maintaining self-hosted infrastructure with 99.9% uptime since 1994 and leading Kubernetes migrations for 50+ applications across multiple countries, the pattern is clear: these conflicts aren't failures of ArgoCD, Kubernetes, or GitOps.
They're failures of organizational design.
The good news? This is fixable. Not with better tools, but with better conversations.
The platform engineering movement is three years old. It's time to grow up and admit: we've been optimizing the wrong layer.
Stop debating ArgoCD vs. Helm. Start defining ownership boundaries.
Stop mandating tools. Start measuring shared outcomes.
Stop letting technical teams fight. Start leading.
The technology is the easy part. The organizational design is the work.
About the author: Henrik Jess is a DevOps engineer with 30+ years of infrastructure experience, specializing in Kubernetes migrations and cloud-native architecture. He's maintained self-hosted infrastructure with 99.9% uptime since 1994 and currently helps organizations bridge the gap between DevOps practices and enterprise governance requirements.