The CI/CD Strategy War — Which One Actually Wins?

#CI/CD Pipeline #Deployment Strategies #Automated Testing #Staging Environments #Feature Flags #Canary Deployments #DevOps
💬 Chat with this Video
Ask anything about this video…
Let's say you're working on an e-commerce platform and your team just finished building a new checkout feature. With this new feature, customers can now save multiple payment methods and there is a buy now button that skips the card entirely. You've tested it on your laptop. Everything works. So now you need to deploy it to production where real customers will actually use it. But here's the question. Do you deploy directly from your development environment straight to production or do you first deploy to a test environment then staging environment then maybe pre-production and then finally to production? Different teams have completely different answers to this question and a lot of them think the other approaches are wrong and what they're doing is the correct way. Some teams deploy straight to production with just a code review. Other teams have five environments between development and production and it takes 2 weeks to deploy a simple change. So in this video I want to show you the three main CI/CD pipeline strategies that I see in real companies and real projects. I'm going to explain why each one exists when each one makes sense and help you choose the right approach for your team. Because here's the important point. There's no universally correct number of environments. If there was, then we would just have one CI/CD strategy and that's it. But it really depends on what you're building, how fast you need to deliver, and most importantly, what the cost of a bug in production is for your project or company. So, let's start with the fastest, most aggressive approach. So, strategy one is the minimalist development to production straight away. So let's go back to our e-commerce platform where your team decides we're going to deploy directly from development straight to the end users. No intermediate stages. So here's how it works. You write the checkout feature code on your laptop. You write automated tests, unit tests that verify that the logic for saving payment method actually works. Then you write integration tests that verify that the buy now button correctly creates an order. And once you're done, you push your code to Git repository. The CSD pipeline gets automatically triggered. It runs code quality checks, security scans, and if everything passes, your code is automatically deployed to production. So from your laptop to live customers in minutes. Now, some of you may be thinking, yeah, that's the way to do it. But a lot of you may be thinking, why would anyone do this? This seems risky, completely irresponsible, deploying straight to production. But here's the thinking behind it. If your automated tests are really good and reliable, you don't need manual testing environments. So, let me explain with our checkout feature. You've written tests that verify payment methods are saved correctly to the database. The buy now button creates orders with the right data. Edge cases are handled. What if someone clicks buy now twice or does some other weird stuff that users usually do with your application? Integration with the payment processor works. If these tests all pass, the code should work in production. So why wait and introduce manual bottlenecks in between? So the core philosophy here is that if you invest heavily in proper automated testing when you're going to trust your tests to test everything before it gets deployed to the production. So instead of spending time maintaining staging environments and doing manual testing, you spend the time writing better, more complex automated tests. And I came up with this. So think of it like when cooking. Some people taste their food multiple times while cooking. After each ingredient or every time they put spices in it, they taste and adjust after each set. The minimalist approach is like following a precise recipe with exact measurements. And if you follow the recipe correctly, then you don't need to taste every step. You know that the dish is going to be fine at the end. But as I said, every strategy fits specific type of projects. So this approach works especially well for startups that are moving fast because you need to ship features quickly. You probably don't have tens of millions of users. So your main priority is to basically iterate really fast based on user feedback. So waiting 2 weeks for a deployment may kill your momentum in your entire startup. It also works really well if you're building tools for your own company like internal tools because the customers are your co-workers. So if something breaks, they can tell you immediately and you can fix it. there's not going to be a really dramatic event if you deploy something that doesn't work properly. And obviously works for teams of any projects that have excellent test coverage. So if you have 80% plus code coverage, but also very important, the quality of tests are reliable. So they test very complex scenarios end to end, you can trust the tests. And finally, you can deploy code to production with feature flex where the features that are in development or in progress are kept hidden for the end users, but it's turned on for just 1% of users first to see if it actually works, nothing breaks, and then gradually increase if everything looks good. However, there are a few issues that you may run into when using this approach. First of all, you may realize your tests are actually not as good as you thought they were. So you wrote tests for many scenarios. You're happy with it, but you forgot to test some edge cases or maybe you forgot to test what happens when payment processor is down. So all your happy path tests actually pass. You deploy and suddenly production breaks when Stripe has an outage and then you jump into a panic mode fixing the bug in the production. And when that happens, your customers basically become your QA team because when bugs sleep through to the production, the actual customers keep them. They get frustrated. They complain on Twitter. they complain on Reddit or X and your support team gets flooded with tickets. And there is one non-obvious thing that sometimes happens in teams that use this approach which is the pressure not to break things slows you down. So because everyone knows that each commit will lend directly into the production. It slows down or makes the team actually extra careful about every single commit. So ironically deploying straight to production can make teams way more conservative. So people become afraid to commit when the deployment is risky which kind of defeats the purpose of the strategy. So the minimalist approach is fast but it requires discipline and really excellent automated testing that is really good in reality not just in your head. And many teams think they have good tests and they're ready to go until they try this approach and realize they don't. But let's look at the exact opposite approach which is the paranoid with five plus environments. So imagine a completely different team working on the same e-commerce checkout feature. And the team says, "We cannot afford bugs in production. It's going to kill our credibility. We have too many users on our platform. We cannot afford that." So every bug costs real money, lost sales, lost customer trust, and sometimes potential legal or regulatory issues. So they set up five environments between development and production. They have dev and test and staging and pre-production and prod. and maybe they squeeze another one in just to be sure and just to be on the safe side. So if we follow our checkout feature through this pipeline, what happens is you finish coding on your laptop and deploy to the dev environment. This is the first shared environment where your code that you wrote locally now runs on actual servers, not your laptop. The dev environment catches basic integration issues, right? Maybe your code works on your Mac but fails on Linux servers. Maybe you forgot to add an environment variable. So dev environment will catch this. Next your code goes to the test environment. So once it passes the dev stage, it basically moves or advances to the test environment. And this is where QA engineers manually test your feature. And they test the checkout feature thoroughly, right? They try to break it. They try to do all types of weird things that they think users may also do. Try all the edge cases. What happens if they click it twice? What happens with invalid credit cards? Does it work on mobile browsers? And so on. and they may spend hours just testing every single scenario. Usually they find bugs and report them to you and you fix the bugs and redeploy to test. So it goes through the same cycle. Once your feature passes the manual testing on the test environment, your code moves to staging. This environment is a clone of production. Same database size, same number of servers, same configuration. So and staging is where you test performance and integration with production data. You run load tests. So what happens if 10,000 customers use the checkout at the same time? Maybe it works for 100 customers but it completely breaks under load. So you test that. You also test migrations here. If your checkout feature requires database schema changes, for example, you test the migration process in staging first. Once everything is fine there, then your code goes to pre-production. And this environment is even closer to production. It might even use the same actual database as production servers, maybe just in readonly mode. So pre-production is basically a final sanity check. You test the deployment process itself. You verify that the monitoring and alerts work properly. You make sure rollback procedures work. So basically any single imaginable scenario is tested here. And finally after passing through all these environments, your code reaches production. It's been 2 weeks since you wrote code or maybe even more. But you're very confident at this point that it works correctly because it's been tested in four different environments. Now again, a lot of people here may think that's exactly how you should release anything. But a lot of people would ask why would anyone do this? It seems like a waste of time and waste of engineer resources. It seems excessive, right? Four environments before production. But here's a logic behind it. Each environment catches different types of problems. An analogy here would be an airport security. You might think going through multiple checkpoints is an overkill, but each checkpoint checks different things. One checks your ticket, then the next one checks your ID or passport, another one scans your bags, then there is another one that does body scanning. So multiple layers that each check for different things for the same person. And as I said, there are completely valid use cases for each approach. For this one specifically, think of banking and financial services. The main difference here is that a bug in production may cost a lot of money for the institution itself but also for the customers. So a bug that loses customer money or exposes financial data can be catastrophic. So the cost of slow deployments is actually worth the safety. The same way in healthcare systems for example a bug in a medical record system could actually harm people's lives. So you need multiple verification steps or think of large enterprises with compliance requirements. You need audit trails, approval gates, verification at each step for regulatory compliance. So it may not be even code logic and functionality that gets checked as much as is the code or the application compliant. Does it go through security checks and so on. So generally as you see the pattern is that systems where bugs cost millions or cause regulatory issues actually justifies having multiple testing environments even though it may be painfully slow to deploy any new change. But of course it comes with lots of downsides which is it takes forever because imagine you are deploying a button color change and it's been sitting there in test environment for a week waiting for QA approval. And another problem may be that environments drift from production because when you have to maintain five deployment environments and keep them in sync, that's a huge overhead. So very often in practice, what happens is your staging environment maybe was a clone of production 6 months ago, but a lot of things changed on production. Different database sizes, different traffic patterns, you did some security patches, you added some servers, and staging does not represent production anymore. And it's also more expensive because you have to actually create and pay for those extra environments, servers, database, maintenance, monitoring. So for a large application, this could be tens of thousands of dollars per month just for maintaining these extra environments. So the paranoid approach is safe, but it's slow and expensive. So we saw these two approaches for deployment, which are kind of two extremes of each other. However, in practice, most teams actually end up somewhere in the middle. So let's call this strategy the practical and this is what most pretty successful software teams actually do. They use three environments devstaging prod. So if we follow our checkout feature through this pipeline what happens is you finish coding and deploy to dev which we saw it's a shared environment where the team can see and test new features before they're ready for customers. And again the basic integration testing is done here to see if the feature actually works when connected to the database to the actual APIs and so on. And you also use dep for collaboration. So your product manager may see the new feature and give you feedback about the usability user friendliness and stuff like this not just the functionality. And your designer can verify that it matches the actual designs. Once the feature looks good in dev it moves to staging. And here's a key difference from the paranoid approach. Staging is specifically a production clone for testing the scary changes. So what are scary changes? Database migrations that modify millions of rows. That's a scary change. Or major architectural changes. Switching payment processor for example or code refactoring where multiple parts of the code actually changes to critical flows like checkout or login or registration. So for our checkout feature, you test in staging whether the database migration that adds a new table for new payment method is actually working or you test integration with payment processor under realistic load and you also test the deployment process itself any underlying dependencies that your code changes actually need. But here is what you don't do in staging. You don't manually test every single feature. Your automated tests already verify that the feature works. So staging is for testing things that automated tests cannot verify like deployment process, database migrations, load testing, production grade load for the application and so on. And after staging looks good, you deploy to production. And here is where the practical approach adds safety in this deployment phase. You deploy gradually. You can use techniques like canary deployments where you deploy to 5% of your servers first. You monitor for errors. You'll just let it run for hours or maybe days and see if everything looks good. Then you deploy to 20% of the servers and then you increase that to 100 as long as nothing breaks. Another technique is feature flex. So the checkout feature gets deployed with feature flex which means by default it's hidden. You can think of it like a switch. So it's kind of switched off by default and you can switch it on or turn it on for maybe just 1% of customers first and see what the feedback is. everyone using the feature without issues. Then you increase it to 10% of the customers and you turn it up to 100% of the customers. And the final safety lever that teams often use with this practical approach is monitoring and roll back. So you have monitoring setup that automatically checks within the first hours, first days, first week of the new deployment to see if there are any issues and if error rates spike for example, it automatically rolls back to the previous version. So even though you are deploying with way less testing compared to the paranoid approach, you still have these guard rails that protect you from completely making a mess in case you deploy something that's broken. So the practical approach kind of balances speed and safety. So it's fast enough to deploy daily. You're not waiting weeks for changes to go through multiple environments. So a typical feature might move from deaf to staging to prod in two to three days, but it's safe enough to sleep well at night once you hit that deploy button because staging catches major issues and gradual rollout catches production specific issues quickly. And if you have good monitoring, it means you know immediately if something breaks and the biggest upside here is that it's affordable. You just have two extra environments that don't cost as much as four. And staging does not need to be as large as production if you're just testing deployments and not performance at scale. And this approach actually works for majority of use cases. Most software companies like SAS products, web applications, mobile backends, this can be a standard approach for these type of companies, but also teams that want to ship frequently but still have low-risk deployments. You can deploy with this multiple times per week while still feeling safe about each deployment. And for any customerf facing application where bugs are not catastrophic, they are annoying and inconvenient, but they're not like dramatically bad like e-commerce, social media, productivity tools, this type of stuff. It's also a perfect approach for this. You still have some issues and risks with it obviously because the two additional stages still need to be maintained and they still may drift from production. And of course, when the staging does not exactly match production, there are less things that you can test. that is really realistic. So you end up still doing a lot of the load testing or integration testing on production environment. That's why we have the gradual roll out and you still need good automated tests for this because staging does not replace testing. It's actually an extra step. So if your automated tests are weak then you're going to have a lot of bugs that reach production even with these extra two stages. So now that you understand all three approaches to CI/CD deployment and why each one makes sense in which type of projects, I think you can answer this specific question yourself, which is how do you decide which approach is right for your team and your project. The simple formula is if a bug in production will cost the company a lot, then you need to be a little bit more paranoid. But if you're a fast-moving startup with relatively low number of users, then you can be risky, especially at the beginning. And for any mature project that has a lot of users, but the cost of a bug in production is not dramatic, which is the majority of the projects can use the practical approach. That's why it's called practical because it works for most use cases in practice. My personal recommendation if you're unsure where to start is always start with practical and then you can shift up or down depending on whether you decide okay we need to shipping is actually much faster then you go to the minimalist approach. But if you see okay there are too many bugs that are ending up in production that is costing our company reputation that is frustrating our customers we cannot afford that then you can introduce additional stages. Now before we wrap up, I want to show you a few common mistakes that I've seen in practice and you may recognize yourself in one of them which can be very practical for a lot of teams. Mistake number one is too many environments that nobody uses. I've seen companies with six environments and nobody can explain what makes prepro different from staging for example and why they need both. They added environments some years ago for reasons nobody remembers anymore and now they maintain all six and deployments take a month. And here's a formula to avoid this. Audit each environment and for each one answer this one question. Is this environment catching a bug or increasing our confidence of deployment? If not, remove it. Mistake number two is where staging does not match production or is not even close to production. And this is probably the most common one that I've seen where staging environment was set up at the same time as production. So there was a match at one point, but over time it completely drifted. It has different versions of libraries, different configurations, different environment variables, different data volumes. So it kind of makes it useless because testing and staging gives you false confidence. It doesn't tell you anything about how the application will perform on production. So you're actually testing without purpose because the purpose of testing is to make sure that the application does not break in production which means naturally you should test in a production-like environment otherwise how do you know that it's going to behave the same way. So the solution here is either keep staging synchronized with production or use staging just for basic integration testing but also become aware and realize testing on staging does not give you any data about how the application will perform on a production environment. This one is very clear where teams do not have automated tests but they jump into minimalist approach and it happens because of practical reasons where a lot of teams think we have great tests I think we're good to go and besides Netflix deploys straight to production so should we and here obvious solution is just to do a reality check whether you really trust your automated tests and whether your tests validate complex end toend processes in your application not just simple happy scenarios. So you start with the practical approach, you add the automated tests until you feel fully confident. You add monitoring and some guardrails and then move to the minimalist approach. And finally, this is a very common and interesting one which is adding environments instead of fixing tests. Very often when teams have flaky, unreliable tests and bugs keep reaching production. So basically instead of writing good automated tests, they add environments and manual testing which slow down the process. So the better fix is always invest in test quality. Improve coverage. Good tests are always better than more environments because they cost less. You don't have to maintain them like you have to maintain environments and they don't take hours like manual testing. They're much faster. And that's a wrap- up for this video. I hope you learned a lot here and I hope that you could take at least one practical action step for your project and share it with the team. or if you have some other approach that I haven't covered here that is super custom and interesting, please share it below in the comments. I think it's going to be fun to see what everybody is using in their real projects. And with that, thank you for watching and I'll see you in the next video.

Generated algorithmically for Search Engine Indexing.

Summarize Another Video