Continuous Delivery: Scaling Continuous Delivery

Its been a while since I posted. Main reason is that we have been very focused on our main deliveries and feature development for the last six month. Whenever the feature train hits central station its always work such as build, release, test automation that gets hit first.

Though there are upsides to not touching your Continuous Delivery process for a few months. If you just keep working on your backlog you don't get time to analyze the impact of the changes you just made. Several times we have realized that the number two/three items in the backlog have dropped significantly in priority as we have fixed the most important issue and others rising fast in priority.

Now we have had time to analyse a lot of new issues and its time for us to pick up the pace again.

Scaling the Organization

The good thing, the awesome thing (!) is that during these six or so months our organization has changed and we have actually be able to create a line organization that owns and takes responsibility for the continuous delivery process.

One of the major bottlenecks we found in our process was our platform/tools team. The team was small and resources in that team where always first to go when feature pressure increased. The team became just another "IT function" that didn't have time to be proactive due to all the reactive support work it had to do.

There was a few reasons behind this first it was the way the team worked in the past. It actually built the pipes and processes for all the teams by hand and tailored to the custom needs of each team. On some teams there were individuals who picked up the work and kept on configuring the jenkins jobs to tailor them even more but on some teams there was no interest whatsoever and their jobs degraded.

The result of this was that no one really knew how the pipes looked and how they should look. Introducing process change was a horribly slow process as it was all manual and dependent on the platform/tools team.

One of the first changes we made was to increase the bandwidth of the team and reducing the dependency on that team. @TobiasPalmborg provided a great solution for this over a chat this summer. Instead of the platform/tools team supporting the development teams the development teams put resources into the platform/tools team. Each team was invited to add a 50% resource on a volunteered basis. This way the real life issues got much better attention in the platform/tools team and the competence about the Continuous Delivery process got spread in a much more organic way.

This did not eliminate the bottleneck organization but it gave us bandwidth to change the way we work and long term gave us the ability to scale with the number of teams that use the process.

Scaling the Process

The main issue with why we were a bottleneck was the way we worked. We preached Automate Everything, Test Everything, If its hard do it more often, ect but when it came to the Continuous Delivery process we didn't do what we where teaching.

We had ONE Jenkins Environment so all the changes happened directly in production. Testing plugins and new configurations on a production environment isn't really the way to delivery stability, reliability and performance.

Manually created Jenkins Pipes isnt really a way to create sustainable pace and continuous improvements.

Developing Deploy scripts without explicit unit tests isnt really a good way of creating a stable process. We have been priding ourselves with our deployment being tested hundreds of times pre production deploy which was true but very dumb. Implicit testing means that someone else takes the pain for my mistakes. Deployment scripts are applications and need to be treated as first class citizens.

This had to change.

First thing we did was to use the extra bandwidth we had obtained to build a totally new way of delivering continuous delivery. Automate everything, obvious, hu?

We also decided to deliver a continuous delivery environment per development team and not have them all in one environment. So we started with automating provisioning of Jenkins & Test environments. We dont have a cloud solution in our company at this time so we have a fake cloud that we work with which is a huge pool of virtual servers. This pool we provision and maintain using chef.

Second thing was to automate the build pipe setup. We built us a little simple pipe generator which has defined pipe templates of 5-8 different layouts to support the different needs. We actually managed to get the development teams to adjust to a stricter maven project naming convention to use the generated pipes as everyone saw the benefits of this.

The pipes we have are basically typed by what they build if its libs or deployable components and how they are tested as we still need to initiate our Fitnesse tests a bit differently from our other tests.

We made it the responsibility of the platform/tools team to develop the pipe templates and the responsibility of the development teams to configure their generator to generate the pipes they needed for their components.

Getting to this stage was a lot of work and a lot of migration work for all the teams but the results have been terrific. The support load has gone down alot on the platform/tools team and each bug fix is rolled out within minutes to all the pipes.

We have also be able to take on new development teams very easily. Not all teams in our company are ready to do Continuous Delivery but they are all heading in this direction and we can now provide environments and pipelines that match their maturity.

Summary

We have gone from a process developed as skunkworkz to Continuous Delivery as a Service within our organization. We always run into new bottlenecks and challenges this time the bottleneck was much more us than anything else. I assume that the next big bottleneck is going to be hardware and our inability to deliver on a cloud solution, since we now can roll out to more and more teams. But who knows I can be wrong only time will tell.

4 comments:

J.February 21, 2014 at 5:44 PM
If your issue is scaling and bandwidth, you should be moving to tools designed for enterprise, not teams. http://www.electric-cloud.com/products/
Tomas RihaFebruary 27, 2014 at 1:53 PM
J,

Tooling is actually a small part of our scaling problem. All our main issues have been People and Organization. The maturity of us as individuals and the our understanding of what needs to be done.

As an example our template based pipline generator is super simplistic. It is the ability to mass produce and mass change pipes that allows us to scale not what tool we used to implement it. In order to get to where we are with generated pipes we had to mature our architecture, stream line our testing, improve our adherence to name standards and grow to realization that we actually need generated pipes.

Tools are imho only about 10% of the time and effort we invest in Continuous Delivery.
AnonymousMarch 4, 2014 at 3:07 PM
Good post. A lot of your issues sounds all too familiar. We have managed to keep complexity and maintainability under control in our 300+ jobs Jenkins installation, by naming conventions, templates and job generation, etc., but still need to improve a little on the testing and releasing of new features in the pipe framework.

But until we have the capability to promote VMs instead of artifacts, it's still going to be hard to test the prod deploy job to 100%. A recent idea we implemented was to have a 'hello world' service/component that is built, tested and deployed using the framework, exactly like the other components. This way we can actually test the prod deploy. :)
Tomas RihaMarch 17, 2014 at 2:46 PM
Marcus,

That is exactly what we do. We have a simple hello world application that we use to integration test our Install and Deploy process. We have integration tests written in JUnit (Groovy) that provision a virtual node using Vagrant+Virtual Box that we have installed on all our Jenkins Slaves. Then we install and deploy the test applicaiton into that virtual node and assert the everything we want to assert using a mix ssh and http calls.

Tomas

Continuous Delivery

Pages

Friday, February 21, 2014

Scaling Continuous Delivery

4 comments: