Beyond Cloudflare: Why Vendor Diversity Isn't Optional Anymore

November 18, 2025. 6-hour global outage. Dashboard down. Workers offline. A single database permission change brought down one of the internet's most reliable platforms.

I've been running infrastructure since 1994—back when "redundancy" meant two physical servers in different racks, not multiple cloud vendors. These days, after migrating 50+ applications across various cloud platforms, I've learned something critical: the problem isn't that Cloudflare had an outage. The problem is that most teams architect like it never will.

The Cloudflare incident on November 18th wasn't a wake-up call about their reliability—it was a reminder about how we think about dependencies. A bot management file doubled in size, exceeded hardcoded limits, and cascaded into a complete platform failure. Workers KV relied on the core proxy. Access depended on KV. The dashboard required Turnstile. When one component failed, everything failed.

Sound familiar? It's the same pattern I see when DevOps teams build on a single cloud provider, or when ITSM teams centralize everything through one vendor's stack. Different domains, same architectural mistake: we optimize for convenience until an outage forces us to optimize for resilience.

What Running Infrastructure Taught Me

Running production infrastructure since 1994 taught me to think about failure differently.

Early lessons:

The pattern hasn't changed—just the scale. What used to be "don't put all servers in one rack" is now "don't put all services with one cloud vendor." The principle is identical: every single point of failure is a future outage waiting to happen.

Here's the uncomfortable truth: Cloudflare makes it easy to consolidate. DNS, CDN, Workers, WAF, Zero Trust, object storage—all in one dashboard, with volume discounts. It's operationally simple and financially attractive. Until the dashboard goes down and you realize you've built a house of cards.

The Real Cost of Convenience

When I talk to teams about vendor diversity, the pushback is always the same: "But it's more complex! More expensive! More to manage!"

Let me translate that: "We're optimizing for today's operational simplicity at the expense of tomorrow's availability."

Here's what that actually costs when a 6-hour outage hits:

Direct costs:

Hidden costs:

Organizations often find that a single major outage costs 3-4x their annual resilience investment. One six-hour incident can pay for years of multi-vendor infrastructure—if you've architected for it in advance.

Beyond Cloudflare: Practical Alternatives

The goal isn't abandoning Cloudflare—it's building operational insurance through diversity. You're not replacing everything, just creating escape hatches for critical paths.

CDN & DDoS Protection

What Cloudflare does well:

Hybrid alternatives:

Primary: Cloudflare (70% of traffic) Failover: Fastly or Bunny CDN (30% of traffic) Origin: Self-hosted or colocation (maintains independence)

Configuration:

DNS: GeoDNS with health checks every 30 seconds
CDN Layer: Multiple providers with origin pull
Failover: Automatic DNS switch on CDN failure

I run this setup for a production environment serving 2M+ requests daily. Cloudflare handles the bulk because they're good at it. Fastly takes the overflow and acts as hot standby. If Cloudflare goes down, DNS shifts traffic automatically. The origin stays independent—it's not locked into any CDN's proprietary features.

DNS: Your Most Critical Dependency

DNS failure means everything fails. No exceptions.

What actually works:

Primary: Azure DNS or Route53 (compliance documentation, enterprise SLAs) Secondary: Self-hosted BIND or PowerDNS (operational independence)

Why both?

Cloud DNS provides geographic distribution and compliance posture needed for regulated industries (GDPR, ISO 27001, SOC 2). Self-hosted provides the escape hatch if your cloud provider has issues.

Cost reality: Running two authoritative nameservers on cheap VPS instances costs €15/month. That's insurance you can afford.

WAF: Defense in Depth

The Cloudflare outage started with their bot management system. If your WAF is down, you're exposed.

Layered approach:

Edge: Cloudflare or Azure Front Door + WAF (broad DDoS protection) Application: Self-hosted ModSecurity or similar (critical services) Internal: Network segmentation and zero-trust principles

Each layer uses different vendors/technologies. An edge WAF failure doesn't expose your application because the app-level rules still apply.

Edge Compute: The Portability Problem

Workers and similar edge compute platforms have high switching costs because they're proprietary runtimes.

Design for portability:

Alternatives with lower lock-in:

Fastly Compute@Edge - WebAssembly-based, language-agnostic, more portable Fly.io - Simple edge deployment, standard containers Self-hosted containers - Ultimate portability, higher operational overhead

The goal isn't avoiding edge compute—it's avoiding architectures that can't be moved if your provider goes down for 6 hours.

Object Storage: The Data Diversity Challenge

R2 is cheap and fast. It's also another Cloudflare dependency.

Hybrid storage strategy:

Primary: Azure Blob Storage or AWS S3 (enterprise SLAs, compliance docs) Secondary: Self-hosted MinIO (disaster recovery, data sovereignty) Sync: rclone cronjob every 15 minutes

This isn't real-time replication—it's disaster recovery with acceptable lag. If your primary object storage fails, you have a recent copy you control. The cost? Storage is cheap. The peace of mind? Priceless when an outage hits.

Zero Trust / Access: The Identity Challenge

Access control is sticky. Users hate authentication friction. That makes it hard to fail over.

Pragmatic approach:

Primary: Azure AD + Conditional Access (enterprise SSO, compliance) Backup: Tailscale or self-hosted WireGuard (critical internal access)

Most users authenticate through Azure AD because it's integrated with their corporate identity. But for critical internal services—monitoring dashboards, deployment tools, emergency access—there's a parallel Tailscale mesh network that doesn't depend on any single vendor.

When the Cloudflare outage took down Access, teams with this setup could still reach critical infrastructure. Teams without it were locked out of their own monitoring systems.

The Architecture Pattern That Works

After countless migrations, here's the pattern that delivers actual resilience:

Three-Tier Dependency Model

Tier 1 - Primary (Cloud Provider)

Tier 2 - Failover (Different Cloud Provider)

Tier 3 - Critical Path (Self-Hosted)

This isn't "duplicate everything." It's strategic redundancy on critical paths.

What to Self-Host vs. Buy

Self-host for:

Buy from cloud for:

The pattern I've seen work: cloud for reach, self-hosted for control.

Auditing Your Dependency Graph

The Cloudflare outage revealed cascading failures through hidden dependencies. Workers KV → Core Proxy. Access → KV. Dashboard → Turnstile. One failure, complete collapse.

Your stack probably has similar hidden dependencies. Here's how to find them:

1. Map Your Dependency Graph

Draw it out. Literally. What services depend on what?

Example from a real migration I did:

Application → Authentication → Database
           ↓
       CDN (Cloudflare)
           ↓
       DNS (Cloudflare)
           ↓
       Monitoring Dashboard (Cloudflare)

Problem: Cloudflare outage means no CDN, no DNS, no monitoring. You're blind and unreachable.

2. Identify Single Points of Failure

Where do multiple services converge?

In the Cloudflare case: their core proxy. Everything depended on it.

In your infrastructure: probably authentication, DNS, or a central API gateway.

3. Test Failover Scenarios

Don't assume it works. Verify it works.

I test failovers quarterly: - DNS failover (switch primary to secondary, verify resolution) - CDN failover (disable primary, verify traffic shifts) - Object storage recovery (restore from secondary, verify integrity) - Access failover (primary auth down, verify backup path works)

Every test reveals assumptions that were wrong. Better to find them in testing than during a production incident.

4. Calculate Blast Radius

If service X goes down, what breaks?

Example blast radius analysis:

Cloudflare CDN failure:

Cloudflare DNS failure:

Knowing the blast radius helps prioritize what gets redundancy first.

The Translation Layer Nobody Builds

This connects directly to what I wrote about in "DevOps Meets ITIL"—different teams speak different languages about the same problem.

Infrastructure teams say: "We need vendor diversity." Finance teams hear: "We want to spend 20% more on redundant services."

DevOps teams say: "Single points of failure are unacceptable." Business teams hear: "Operational complexity is more important than cost efficiency."

The translation layer is showing cost-benefit in business language:

Don't say: "We should implement multi-vendor resilience architecture."

Say: "A 6-hour outage typically costs 3-4x our annual resilience investment. That's clear ROI on insurance against downtime."

Finance understands insurance. Executives understand risk mitigation. Frame vendor diversity as operational insurance, not infrastructure complexity.

The Compliance Angle Nobody Mentions

If you operate in regulated industries (finance, healthcare, government), vendor diversity isn't just good practice—it's often required.

Compliance frameworks that care about single points of failure:

ISO 27001: Requires documented risk management and business continuity SOC 2: Availability criteria demand redundancy and failover GDPR: Data sovereignty means understanding where vendor outages leave data exposed NIS2 (EU): Critical infrastructure must have incident response and resilience plans

When I helped a financial services client migrate to multi-vendor architecture, their auditor explicitly noted it as a positive finding. Vendor diversity became a compliance strength, not just operational best practice.

Key Takeaways

The goal isn't replacing Cloudflare or any single vendor entirely. It's building escape hatches for critical paths so when—not if—your primary provider fails, you have options.

I've learned this: the best time to architect for failure is before the outage. The second-best time is right now.


About the author: Henrik Jess is a DevOps engineer with 30+ years of infrastructure experience, specializing in cloud migrations and resilience architecture. He's run production infrastructure since 1994, evolving from bare-metal Unix servers through VMs to modern container orchestration. Three decades of operations taught him that infrastructure fundamentals remain constant—and that high uptime is earned through understanding failure, not claimed through statistics.

← Back to Articles ← Back to Home