The Myth of the Perfect Deployment
After eight years of running production Kubernetes clusters, I’ve seen every deployment strategy fail in spectacular ways. The industry loves to talk about blue-green deployments like they’re some kind of silver bullet. They’re not. Rolling deployments get treated as the default safe choice. They’re not that either. The truth is messier: each strategy solves specific problems while creating others, and the best engineers I know mix and match based on what they’re actually trying to achieve.

Most teams pick a deployment strategy based on what they read in a blog post or what their platform team mandated. This is backwards. You should choose based on your application’s failure modes, your team’s operational maturity, and the real-world constraints of your business. A payment processor has different needs than a content management system. A startup with three engineers operates differently than a bank with compliance requirements.
The question isn’t which strategy is best. It’s which strategy fails in ways you can tolerate and recover from quickly.

Rolling Deployments: The Double-Edged Default
Rolling deployments ship with Kubernetes because they handle the common case reasonably well. Pods get replaced gradually, traffic shifts automatically, and if something breaks you can rollback without downtime. In theory. In practice, I’ve seen rolling deployments cause more production incidents than any other strategy, and it’s usually because teams don’t understand what “gradually” actually means for their workload.
The killer feature of rolling deployments is also their biggest weakness: they run old and new code simultaneously. This works fine if you’re serving static content or your API is perfectly backward compatible. It becomes a nightmare when you have database migrations, message queue format changes, or any shared state that both versions need to access differently. I’ve debugged incidents where a rolling deployment created a two-hour window of data corruption because the new code expected a database column that didn’t exist yet.
The readiness probe is your best friend here, but most teams configure it wrong. Don’t just check if your web server responds to HTTP requests. Verify that your database connections work, your dependencies are healthy, and your application can actually serve traffic successfully. A proper readiness probe should fail if deploying right now would cause problems. This sounds obvious, but I’ve seen production clusters where the readiness probe was literally just “return 200 OK” regardless of application state.
Rolling deployments work best for stateless applications with rock-solid backward compatibility and comprehensive health checks. If you can’t check all those boxes, consider other options.
Blue-Green: High Stakes, High Rewards
Blue-green deployments get all the press because they look elegant in diagrams. Run two identical environments, deploy to the inactive one, test it thoroughly, then switch traffic over in one atomic operation. Zero downtime, instant rollback capability, and complete isolation between old and new versions. What’s not to love?
The cost. Blue-green requires doubling your infrastructure for every deployment. Not just compute resources, but databases, caches, external integrations, and anything else your application touches. I’ve worked with teams that spent months building blue-green pipelines only to discover they couldn’t afford to run them in production. The economics work for applications that generate significant revenue per instance, but they fall apart for internal tools or high-volume, low-margin services.
Even if you can afford the infrastructure, blue-green creates operational complexity that teams underestimate. Managing two complete environments means twice the monitoring, twice the configuration drift potential, and twice the places where things can go wrong. I’ve seen blue-green deployments fail because the team forgot to update DNS records, because load balancer health checks weren’t configured correctly, or because the “identical” environments had subtle differences that only appeared under production load.
Blue-green shines for applications where you absolutely cannot tolerate mixed versions, where rollback speed matters more than resource efficiency, and where you have the operational maturity to manage the complexity. Financial trading systems, critical APIs with strict SLAs, and applications with complex state transitions are good candidates. Internal dashboards and development tools probably aren’t worth the overhead.
Canary Deployments: Controlled Risk
Canary deployments split the difference between rolling and blue-green strategies by routing a small percentage of traffic to the new version while keeping most users on the stable version. This gives you real production data about the new version’s behavior without exposing your entire user base to potential issues. When done right, canary deployments catch problems that no amount of staging environment testing would reveal.
The challenge is defining “done right.” Most teams start with a simple traffic percentage split, discover that 5% of users hitting a broken feature is still a lot of angry customers, then spend months building sophisticated routing rules based on user segments, geographic regions, or request characteristics. The tooling ecosystem has improved dramatically with service meshes like Istio and specialized tools like Flagger, but you’re still building a complex system that can fail in subtle ways.
Effective canary deployments require three things that many teams don’t have: comprehensive metrics that update quickly, automated rollback triggers that you trust, and a traffic routing layer that doesn’t become a single point of failure. I’ve seen teams build beautiful canary pipelines that detected issues immediately but took twenty minutes to rollback because their automation wasn’t designed for speed. I’ve also seen canary deployments succeed technically while failing operationally because the team couldn’t respond to alerts fast enough.
Start simple with canary deployments. Route 10% of traffic to the new version, watch your key metrics for ten minutes, then proceed or rollback based on error rates and response times. Add sophistication only after you’ve proven the basic workflow works reliably in production.
Feature Flags: The Deployment Strategy Nobody Talks About
The most successful deployment strategy I’ve seen in production isn’t technically a deployment strategy at all. It’s deploying dark code behind feature flags, then enabling features gradually through configuration changes. This separates code deployment from feature release, giving you the benefits of blue-green’s instant switching with the resource efficiency of rolling deployments.
Feature flags solve the problem that all deployment strategies struggle with: determining whether a change is working correctly requires real production traffic, but exposing production traffic to changes carries risk. With feature flags, you deploy code that won’t execute, verify the deployment succeeded, then enable features for small user segments while monitoring the impact. Problems get fixed by toggling flags, not by deploying code.
The downside is complexity in a different dimension. Your application code becomes more complex because it needs to handle multiple feature states gracefully. Your configuration management becomes more complex because feature state is now part of your production configuration. Your testing becomes more complex because you need to verify that all feature flag combinations work correctly. Teams that embrace feature flags often find themselves building internal tooling for flag management, user targeting, and gradual rollouts.
Feature flags work best for teams that already have strong configuration management practices and applications that can tolerate the additional code complexity. They’re particularly powerful for consumer-facing products where you want to measure user behavior changes, or for enterprise software where different customers might want different feature sets.
Choosing Your Own Adventure
The right deployment strategy depends on your specific constraints, not on what worked for someone else’s team. Consider your application’s architecture, your team’s operational capabilities, your infrastructure budget, and your tolerance for different types of failures. Most successful teams end up using different strategies for different services based on their requirements.
What deployment strategies have worked or failed for your team? I’m curious about the edge cases and unexpected challenges that don’t make it into the documentation.
