November 18, 2025. 6-hour global outage. Dashboard down. Workers offline. A single database permission change brought down one of the internet's most reliable platforms.
I've been running infrastructure since 1994—back when "redundancy" meant two physical servers in different racks, not multiple cloud vendors. These days, after migrating 50+ applications across various cloud platforms, I've learned something critical: the problem isn't that Cloudflare had an outage. The problem is that most teams architect like it never will.
The Cloudflare incident on November 18th wasn't a wake-up call about their reliability—it was a reminder about how we think about dependencies. A bot management file doubled in size, exceeded hardcoded limits, and cascaded into a complete platform failure. Workers KV relied on the core proxy. Access depended on KV. The dashboard required Turnstile. When one component failed, everything failed.
Sound familiar? It's the same pattern I see when DevOps teams build on a single cloud provider, or when ITSM teams centralize everything through one vendor's stack. Different domains, same architectural mistake: we optimize for convenience until an outage forces us to optimize for resilience.
What Running Infrastructure Taught Me
Running production infrastructure since 1994 taught me to think about failure differently.
Early lessons:
- Physical diversity (different racks, different power circuits)
- Provider diversity (multiple ISPs with BGP failover)
- Geographic diversity (services spread across locations)
- Component diversity (avoiding single vendor lock-in)
The pattern hasn't changed—just the scale. What used to be "don't put all servers in one rack" is now "don't put all services with one cloud vendor." The principle is identical: every single point of failure is a future outage waiting to happen.
Here's the uncomfortable truth: Cloudflare makes it easy to consolidate. DNS, CDN, Workers, WAF, Zero Trust, object storage—all in one dashboard, with volume discounts. It's operationally simple and financially attractive. Until the dashboard goes down and you realize you've built a house of cards.
The Real Cost of Convenience
When I talk to teams about vendor diversity, the pushback is always the same: "But it's more complex! More expensive! More to manage!"
Let me translate that: "We're optimizing for today's operational simplicity at the expense of tomorrow's availability."
Here's what that actually costs when a 6-hour outage hits:
Direct costs:
- Lost revenue (e-commerce, SaaS, subscriptions)
- SLA penalties (enterprise contracts)
- Support overhead (incident response, customer communication)
Hidden costs:
- Customer trust erosion
- Reputation damage (social media amplification)
- Team burnout (6-hour all-hands incident response)
- Executive scrutiny (explaining single points of failure)
Organizations often find that a single major outage costs 3-4x their annual resilience investment. One six-hour incident can pay for years of multi-vendor infrastructure—if you've architected for it in advance.
Beyond Cloudflare: Practical Alternatives
The goal isn't abandoning Cloudflare—it's building operational insurance through diversity. You're not replacing everything, just creating escape hatches for critical paths.
CDN & DDoS Protection
What Cloudflare does well:
- Global PoP distribution
- Free tier for small sites
- DDoS protection at scale
Hybrid alternatives:
Primary: Cloudflare (70% of traffic) Failover: Fastly or Bunny CDN (30% of traffic) Origin: Self-hosted or colocation (maintains independence)
Configuration:
DNS: GeoDNS with health checks every 30 seconds
CDN Layer: Multiple providers with origin pull
Failover: Automatic DNS switch on CDN failure
I run this setup for a production environment serving 2M+ requests daily. Cloudflare handles the bulk because they're good at it. Fastly takes the overflow and acts as hot standby. If Cloudflare goes down, DNS shifts traffic automatically. The origin stays independent—it's not locked into any CDN's proprietary features.
DNS: Your Most Critical Dependency
DNS failure means everything fails. No exceptions.
What actually works:
Primary: Azure DNS or Route53 (compliance documentation, enterprise SLAs) Secondary: Self-hosted BIND or PowerDNS (operational independence)
Why both?
Cloud DNS provides geographic distribution and compliance posture needed for regulated industries (GDPR, ISO 27001, SOC 2). Self-hosted provides the escape hatch if your cloud provider has issues.
Cost reality: Running two authoritative nameservers on cheap VPS instances costs €15/month. That's insurance you can afford.
WAF: Defense in Depth
The Cloudflare outage started with their bot management system. If your WAF is down, you're exposed.
Layered approach:
Edge: Cloudflare or Azure Front Door + WAF (broad DDoS protection) Application: Self-hosted ModSecurity or similar (critical services) Internal: Network segmentation and zero-trust principles
Each layer uses different vendors/technologies. An edge WAF failure doesn't expose your application because the app-level rules still apply.
Edge Compute: The Portability Problem
Workers and similar edge compute platforms have high switching costs because they're proprietary runtimes.
Design for portability:
- Keep edge logic thin
- Use containerized workloads where possible
- Avoid vendor-specific APIs (or abstract them)
- Document assumptions about runtime behavior
Alternatives with lower lock-in:
Fastly Compute@Edge - WebAssembly-based, language-agnostic, more portable Fly.io - Simple edge deployment, standard containers Self-hosted containers - Ultimate portability, higher operational overhead
The goal isn't avoiding edge compute—it's avoiding architectures that can't be moved if your provider goes down for 6 hours.
Object Storage: The Data Diversity Challenge
R2 is cheap and fast. It's also another Cloudflare dependency.
Hybrid storage strategy:
Primary: Azure Blob Storage or AWS S3 (enterprise SLAs, compliance docs) Secondary: Self-hosted MinIO (disaster recovery, data sovereignty) Sync: rclone cronjob every 15 minutes
This isn't real-time replication—it's disaster recovery with acceptable lag. If your primary object storage fails, you have a recent copy you control. The cost? Storage is cheap. The peace of mind? Priceless when an outage hits.
Zero Trust / Access: The Identity Challenge
Access control is sticky. Users hate authentication friction. That makes it hard to fail over.
Pragmatic approach:
Primary: Azure AD + Conditional Access (enterprise SSO, compliance) Backup: Tailscale or self-hosted WireGuard (critical internal access)
Most users authenticate through Azure AD because it's integrated with their corporate identity. But for critical internal services—monitoring dashboards, deployment tools, emergency access—there's a parallel Tailscale mesh network that doesn't depend on any single vendor.
When the Cloudflare outage took down Access, teams with this setup could still reach critical infrastructure. Teams without it were locked out of their own monitoring systems.
The Architecture Pattern That Works
After countless migrations, here's the pattern that delivers actual resilience:
Three-Tier Dependency Model
Tier 1 - Primary (Cloud Provider)
- Azure/AWS/Cloudflare for 80-90% of operations
- Optimized for cost, features, operational simplicity
- Accept this will fail occasionally
Tier 2 - Failover (Different Cloud Provider)
- Hot standby at smaller scale (10-20% capacity)
- Different vendor to avoid correlated failures
- DNS/load balancer health checks for automatic failover
Tier 3 - Critical Path (Self-Hosted)
- DNS (authoritative nameservers)
- Object storage (DR copy)
- Core business logic (the service you absolutely cannot lose)
- Emergency access (VPN, bastion hosts)
This isn't "duplicate everything." It's strategic redundancy on critical paths.
What to Self-Host vs. Buy
Self-host for:
- Services you cannot afford to lose (core business logic)
- Data sovereignty requirements (GDPR-sensitive data)
- Control requirements (heavily regulated industries)
- Emergency access (when cloud dashboards fail)
Buy from cloud for:
- DDoS mitigation (you need their scale)
- Global CDN (you need their PoP distribution)
- Primary object storage (you want their availability SLAs)
- Edge compute (you want geographic reach)
The pattern I've seen work: cloud for reach, self-hosted for control.
Auditing Your Dependency Graph
The Cloudflare outage revealed cascading failures through hidden dependencies. Workers KV → Core Proxy. Access → KV. Dashboard → Turnstile. One failure, complete collapse.
Your stack probably has similar hidden dependencies. Here's how to find them:
1. Map Your Dependency Graph
Draw it out. Literally. What services depend on what?
Example from a real migration I did:
Application → Authentication → Database
↓
CDN (Cloudflare)
↓
DNS (Cloudflare)
↓
Monitoring Dashboard (Cloudflare)
Problem: Cloudflare outage means no CDN, no DNS, no monitoring. You're blind and unreachable.
2. Identify Single Points of Failure
Where do multiple services converge?
In the Cloudflare case: their core proxy. Everything depended on it.
In your infrastructure: probably authentication, DNS, or a central API gateway.
3. Test Failover Scenarios
Don't assume it works. Verify it works.
I test failovers quarterly: - DNS failover (switch primary to secondary, verify resolution) - CDN failover (disable primary, verify traffic shifts) - Object storage recovery (restore from secondary, verify integrity) - Access failover (primary auth down, verify backup path works)
Every test reveals assumptions that were wrong. Better to find them in testing than during a production incident.
4. Calculate Blast Radius
If service X goes down, what breaks?
Example blast radius analysis:
Cloudflare CDN failure:
- Static assets unreachable (site broken)
- API calls timeout (no caching)
- Users see origin directly (DDoS exposure)
Cloudflare DNS failure:
- Domain doesn't resolve (complete outage)
- Email stops working (MX records unreachable)
- Monitoring alerts fail (can't reach monitors)
Knowing the blast radius helps prioritize what gets redundancy first.
The Translation Layer Nobody Builds
This connects directly to what I wrote about in "DevOps Meets ITIL"—different teams speak different languages about the same problem.
Infrastructure teams say: "We need vendor diversity." Finance teams hear: "We want to spend 20% more on redundant services."
DevOps teams say: "Single points of failure are unacceptable." Business teams hear: "Operational complexity is more important than cost efficiency."
The translation layer is showing cost-benefit in business language:
Don't say: "We should implement multi-vendor resilience architecture."
Say: "A 6-hour outage typically costs 3-4x our annual resilience investment. That's clear ROI on insurance against downtime."
Finance understands insurance. Executives understand risk mitigation. Frame vendor diversity as operational insurance, not infrastructure complexity.
The Compliance Angle Nobody Mentions
If you operate in regulated industries (finance, healthcare, government), vendor diversity isn't just good practice—it's often required.
Compliance frameworks that care about single points of failure:
ISO 27001: Requires documented risk management and business continuity SOC 2: Availability criteria demand redundancy and failover GDPR: Data sovereignty means understanding where vendor outages leave data exposed NIS2 (EU): Critical infrastructure must have incident response and resilience plans
When I helped a financial services client migrate to multi-vendor architecture, their auditor explicitly noted it as a positive finding. Vendor diversity became a compliance strength, not just operational best practice.
Key Takeaways
- Vendor consolidation optimizes for today's convenience at tomorrow's availability expense. The operational simplicity of single-vendor platforms creates architectural fragility.
- Every single point of failure is a future outage. The Cloudflare incident proved this again. Your stack has similar dependencies—find them before they find you.
- Multi-vendor infrastructure costs 20-30% more but typically pays for itself after one major incident. Calculate downtime cost vs. infrastructure investment. Insurance is worth it.
- Self-host strategically, not religiously. Cloud for scale, self-hosted for control. You don't need to run everything—just the critical paths you can't afford to lose.
- Test failovers regularly. Assumptions about backup systems working are just assumptions until you verify them under real conditions.
- Translate infrastructure concerns into business language. Frame vendor diversity as operational insurance with clear ROI, not technical complexity.
- Compliance frameworks reward resilience. Multi-vendor architectures often satisfy regulatory requirements around availability and risk management.
The goal isn't replacing Cloudflare or any single vendor entirely. It's building escape hatches for critical paths so when—not if—your primary provider fails, you have options.
I've learned this: the best time to architect for failure is before the outage. The second-best time is right now.
About the author: Henrik Jess is a DevOps engineer with 30+ years of infrastructure experience, specializing in cloud migrations and resilience architecture. He's run production infrastructure since 1994, evolving from bare-metal Unix servers through VMs to modern container orchestration. Three decades of operations taught him that infrastructure fundamentals remain constant—and that high uptime is earned through understanding failure, not claimed through statistics.