A lot of teams say they want better CICD. What they usually mean is they want fewer weird failures, less shared state nonsense, less tribal knowledge, and way less time wasted fighting the delivery system itself. Because the problem is not just that Jenkins is old.
The problem is when your pipelines are noisy, Fragile, hard to reuse, hard to observe, and tightly coupled to a bunch of infrastructure decisions developers should not have to think about. And when you finally fix the problem, you get a different problem. Success. Because once delivery gets fast and easy, people start using it for everything.
Bots, maintenance changes, dependency bumps, mass rollouts, and now the real question becomes, how do you keep the system smooth when everybody finally trusts it? Like and subscribe! Hey, I'm Brian. I work in DevOps and SRE, and I run Tellers Tech. Ship It Weekly is where I filter the noise and focus on what actually matters when you are the one running infrastructure and owning reliability.
Most weeks, it's a quick news recap. In between those, I do interview episodes with people who have actually built things, migrated real systems, and learned what works the hard way. Today is one of those conversations. I'm joined by Stefan Moser from Pipedrive.
He helped lead a big move from Jenkins to GitHub Actions, built a self -hosted runner platform on Kubernetes, moved delivery towards GitOps with Argo CD, and helped roll that model out across a large internal estate with hundreds of services. And what I like about this one is it's not just tool talk. We get into why Jenkins had become painful, from groovy fiction to noisy neighbor problems on shared VMs.
Why GitHub Actions ended up fitting better. How reusable workflows and custom actions helped. Why they chose Argo CD over other deployment options. And how they had to build better internal observability because GitHub alone was not enough at their scale. We also talk about the migration strategy. Which honestly is one of the best parts.
Dogfooding first, migrating in batches, using internal teams as the first proving ground, letting the process get polished before pushing it wider, and building something self -service enough that teams eventually started migrating on their own. And then there is the mobile story, which is its own thing.
Mac minis, messy runner drift, different toolchains and the surprisingly practical path they landed on after testing a few different options for stabilizing mobile CI. If you care about CI CD architecture, platform engineering, GitHub actions at scale, or how to do a migration like this without setting your org on fire, this one is worth your time. All right, let's jump in.
Today, I'm joined by Stefan Moser from Pipedrive. He helped lead a big move from Jenkins to GitHub Actions and built a self -hosted runner platform on Kubernetes, plus a GitOps CD flow with Argo CD. And we're going to talk about what worked, what broke, and what's worth copying. Stefan, thanks for joining me. Thank you. Thank you for the opportunity to talk about this adventure.
Basically, it's not the first time I'm talking about this. I already had... Meetup session that was recorded on YouTube, two blog posts, and basically just sharing again these adventures. But now with a little sparking because something already after one year already changed. So I have more stuff to add to these blog posts and meetup that we had. Awesome. Well, I'm excited to learn more.
So starting off, can you give me a thesis? Why did Jenkins stop working for you, and what were you trying to optimize with this new system? So basically, the first problem that we had with Jenkins is Jenkins decides being the pipelines being written in Groovy. That was the first problem. In Pipedrive, we are mainly working with TypeScript and Go. Groovy was really not very needed for us, so it was a bearer.
For the DevOps teams that we had at that time, plus the engineers trying to do writing something. That was the first issue. The second issue was basically the setup. So we had VMs, and that VMs were not isolated. So that means if a big pipeline lands in one VM, and we had other pipelines like to build the Docker containers, then we had the issue with the nice...
Noisy neighborhood so basically we get a starvation of the resource so it's really not predictable we tried to to improve that so we had a crazy idea first to build a internal ci engine so that was really the most crazy thing they did so basically copy the idea of yaml from github's gitlab ci we build the engine and then start to work on that but then we had another issues People need to learn and use syntax.
And after GitHub Land launched GitHub Actions, people started using it, some developers started using it in their app stores. And then we were thinking, why not experimenting? And the first idea was basically, besides Jenkins, we had another product, CodeChip, that was to do the pull request validation. So every time the pull request is open, we had a typical lint and test work.
So basically linting to building the... Code, leading the code, and doing the unit tests. They were running in the code chip. And I can tell it was worse than Jenkins because every project was individual. So imagine if I had to change or to share something, it's very difficult. So we had an idea. Well, let's try to win the GitHub Actions. And basically, I team up with a colleague, Gregor.
And basically, it was the idea that... Let's replace CodeChip with GitHub. Actually, let's see what happens. And that was really the kickoff. So we get a green card to export idea. So the first thing was thinking that is, okay, I need to run this somewhere. Probably GitHub, even we had a free package, will not be enough. So let's go to find solutions for self -hosting.
We found, after some ideas, we found a community project. It is ARC, Actions Run Controller. At this moment, it's already belonging to or maintained by Gita, but at that time was really purely community -based. One thing that it was very... Important for us, or very convenient for us, it was a Kubernetes controller. So we spent the last years working with Kubernetes.
We are understanding the Kubernetes API, these CRDs, all talking in the Kubernetes way. So it was more easy to work with that. So we define resource, get runners running, fine. One thing that is also awesome, we can grab a pre -built image and then... And put that customization that you want. So basically, developer needs HGK, IKH, or different tooling. Just put it inside of that custom image and then we ship it.
It was even for easy. Then we already know how to restrict and monitoring Kubernetes in production. Just apply the same idea in the...
CI environment so basically I did a thing that is normally not normal standard so basically I don't I didn't set the requests to be the smallest possible most necessary no I just was brute force so I say for example set the request and the limits to be almost the same so that I could have one thing that is pretty bit in the So that every time a developer gets a runner, it's always the same CPU and memory.
So that avoids the problem of noisy neighborhoods. Then we had already ideas how to, we already know how to see and control the resource in Kubernetes. So apply the same metrics. So basically we bring back the, of the really tools that we had in production to.
The ci cluster but then we also did want to be more straightforward in in the process to maintain the clusters basically instead of building the cluster from scratch we use ets to be more easy and then we want to to have scalability in terms of the nodes we decide then to go to the new project at that time was carpenter so basically the the magic way of AWS to spawn nodes faster than a cluster of auto -scaling.
So basically, then we get this solution. So a controller that was listing the GitHub app books to spawn jobs, spawn when the instance creates a new pod that represents a runner. If that runner doesn't have a space, the carpenter then creates a new node.
And that was really bring the flexibility okay I just need to set the max number of thoughts or the runners and the ecosystem for the cluster scale up and scale down when when necessary and this was really basically the why we we moved to uh two changes and first steps we then basically did the migration of code chip to github actions at that time also was released the reusable workflow so that means we are able to
Reusable to create a workflow and spread that workflow to the old repositories so reducing the the repetition that we had and the manual configuration we had with Jenkins with the code chip and Basically, after we dropped out of CodeChip, it was the time that we had the opportunity to revamp CI -CD.
So we bring a group of engineers, I think four engineers plus a developer experience product manager. So basically, we had a product manager dedicated for the developer experience. And then we decided to revamp.
Uh necessity and they basically said to put some kind of competition in in terms of tooling so I remember that the first contender was basically github actions in terms of ci but then also we bring argo workflows and tecton because were two projects that I was curious about using And in terms of deploying, I'm thinking Argo CD plus Flux.
We even tried with Spinnaker, but it was so messy to spin up that system that we just dropped it. And how we chose, basically, first was...
Readability or feasibility or the easiness to create the workflows and the specific customizations so and that is really shine for us the kit of actions basically people complain about the action but really the fact that the actions are written in javascript and then we are using typescript so it's natural to create custom logic with typescript even then we can go to creating composite actions with some more for some batch scripting.
Scripts, we just use that language to create the customization. And then it was really easy to pack up. So basically, we create action, we have a package, like a package, and then can put it a block that we can put in different place. And then the other point was basically about how we can expose or show the workflows to the developers. And that is... The less clicks the developer does, find lots better.
And that, of course, makes the GitHub action to be a big content. So the fact that the workflows was next to the repository, the execution was next to the repository, the developer didn't need to switch to the platforms to see was really a big plus. Yeah. And that's really the reason why we went to... To GitHub Actions. Even at some point after this migration, what we had, it's basically a monorep of GitHub Actions.
And I think we have 50 or 60 GitHub Actions. So basically, we just build our actions and even it's our monorep of all the code base that you want to reuse in the organization. In terms of development, of deployment, one thing that we want to make. Clear is that we want to avoid push code.
So one of the issues that we had always is that with Jenkins and this push idea is that fact that I need to have a runner in the clustered push code. That means that I need to give the Kubernetes credentials to a runner to be able to write stuff. And that means that in some way I can try to capture that runner and do malicious stuff.
Let's do the reverse so that is that the cluster is really in isolation and we get what we have it's someone inside that check something that it's really it's already trusted and use it to um to apply change in the cluster and that is really why you want to go to the to do githubs and of course in githubs we have two two tools to choose argo cd or flags yeah And in that case, it was basically one of the reasons why Argos CDF had a good UI comparing with Flux at that time.
It was like two, three years ago. So that was really the big reason that we picked up with Argos CDF Flux for the UI. And then we had already a way to show to developers what has happened in a more easy way.
Because again, that is that it's easy for developers to visualize doesn't mean that I allow them to change manifest with rcd yeah I know what I allow is to basically scope in inside of the application saying okay this is your application see your status you can go see the status of your resource you can see the logs if you need it you can see if the content the pod is it's restarting or not really um we added that idea.
Yeah, you can show. Yeah, after that idea, I think that picking up the GitHub Actions took me like two days to create MVP of the deployment flow, basically because I was reusing everything that I already did in the CodeChip migration, and we decided to have that MVP, and then... See what are the gaps.
So one thing that I forgot to mention is that initially in the first step or first iteration that we had with migrating to CodeChip, we had a problem with a lack of observability in terms of actions. GitHub at that time didn't provide any statistics about the actions. So if you want to know how long was the runner, we have any idea. You can just... See at the repository and was not good enough.
But especially in our organization that we already have, at that time we had like 700 service. We have architecture of microservice plus libraries. It was impossible to track at the repository level. So what we did is basically we create a service that lists these events and then starting putting everything inside of a database so that could build.
A source that we can then do queries and then find uh what are the the performance of of the the workflow because at some point a manager a director or even cto can comes to us it says what is the failure rate of these uh workflows where what are the workflows that are taking or these guys what are the deployment flows that are taking more time what are the steps that we need to to optimize and then basically this is really our gut feeling that we need that information and we are storing that information.
And until now, we're still using that service to track, even now to tracking more billing information, but it's really the source of truth. Okay, that was the phase one and phase two, phase three. It's basically making the product ready for the consumption. And one thing that we discovered that we lack, it's basically, again, organization -wide of the visibility.
GitHub actually is great to see your workflows, to see the organization. It's bad because we don't have anything. So we create a service. That is basically a registration service plus a UI that we have in our back office. So basically we have a UI in our back office that shows all the deployments, but they need to have some source.
And basically we create a service that consume events that we are producing in our workflows. And after consuming these events can do the mapping of what it's app. And basically we have basically our own interpretation of the events of. Because they have a different interpretation of what is deployment, what is a deploy, what is a region.
And then basically we add that internal logic or that internal understanding what is build, test, unit test, functional test, end -to -end test, deploy to a region, deploy to a... Deployment at a whole scale. And we have that understanding inside of a service.
And basically this was initial, the service that was consuming everything to create observability of what we need to do or what we have in terms of deployments to production. And after that, we have that, after having that UI, we are more confident to go. Now goes the question, how you do this kind of release? 700 service. I think we have five teams or 10 teams at that time. It was really a big migration.
How you start migrations? I think like any product should migrate or should be released, you first try it yourself. So basically... And we decided to migrate our own service. My team doesn't only produce CI, CD. We have a couple of service that supports the deployment flows and support the developer experience. So for example, we had one service that creates and do all the lifecycle of remote dev environment.
So basically we have remote dev environments that developers can use that it's a replica of production. And then we have a system that creates that. For example, this is one of these servers that we need to migrate. And that is really basically like this. So I have scripts, I have workflow, I have documentation to do it. So I deliver to my teammates so that my teammates can migrate. These bring two good things.
First, we test our process to someone that doesn't know about the process. And second, we allow... To share knowledge. So basically, knowledge that was restricted to that five people that create systems starting to be spread to the team because I'm sharing that knowledge by forcing the people to do the migration. And basically, it was basically this alpha grouping. Then we had a different team.
So basically, it was a platform team that it's more experienced than the normal developers around stuff with Kubernetes. And the deployment process and they have more out of the box or out of the standard service, then we go to them and say, okay, we have this Polish workflow. Let's now teach you, do a session and then allow them to migrate.
And then it was basically like the open beta of the system, of the migration. Again, the idea was to... Make everything more polished, more clear for the users. After they give us our feedback and then we improve the workflows, define, we start doing to think how we do all the role. In this case, we pick up the prioritization that SRE team define for the service. So basically SRE define tiers.
So tier one is a service that's very critical. Tier two, it's more or less. And then tier three is less critical. So we decided, okay, let's split this in batch. Let's go team by team, starting in tier three. Then they can move to tier two and then tier one. And that was the idea. And then the plan was basically split it in batch and assign a DevOps engineer to each batch, not to do the...
The migration but to be the assistant so basically the person who does the introduction of the new way ui the new workflow and then this is the the the interaction goal that we have to to do your mind and also to make them so basically if I want everyone one person that's responsible to achieve a goal of migrate x number of of service until the end of the month it will keep the momentum in the team and then we
Started doing that so team by team we started moving we have a schedule we have a limit number of of engineers so basically we have a queue until some moment we had guys from the middle of the queue at the end of saying okay I already saw the other guys doing migration.
I think I can do it alone. Okay, just check the recording that we had. Try yourself. Yeah, and they decide to try alone and starting doing migration alone. So basically, the system was already so well oiled and everything was moving smoothly. They were able to migrate alone. And that's really, I think, the real sex story of doing these migrations by batch.
Doing your dog food and all these things in a way that makes everything automatic or with less human intervention to allow people to use it. And then you also can think this is really the idea of maybe platform engineering, making everything self -service. When you have a product that's already usable for a normal product engineer that doesn't need to have the context of the CICD and how to deploy to Kubernetes.
It's a service that he can use and can work and can do his work autonomously. And yeah, that's then after five months, I think we migrate everything. Yes, we found what we get as issues, basically our success starting break things. So the process was so smoothly. The deployment flow was so good. That then we're starting an introduction of bots to look at deployment for the maintenance tasks.
So basically, the SRA team developed a service to create pull requests to adjust the resource usage in production and then automatically deploy that to production. So basically, it creates a pull request, we introduce a set of resources, and this pull request is already pre -approved, so it moves by all the normal pipeline flow and goes to production. Then we have the PandaBot.
Creating pull requests for the pendants and in some case the developers were already confident enough in the unit test they have in the functional test they have allowed that obligations or updates of of dependence to go without any approval from a human so we had then a situation that we have multiple um departments happen and then was struggling or doing impact in our system. Even we can grow, grow, grow.
Then we add some bottlenecks. At that point, we decide to improve that service that was the deployment registrar to have a queue. So the idea is that you then have a way that you, when you start the deployment, you send event, I have a deployment. And then what happens is that the deployment registrar register that deployment and put in the queue.
Evaluate the queue size so basically we define that we have the number of 50 deployments per available in parallel and also we do a tricky thing that is we only allow 10 of that queue to be used by bots so basically we always want to have some kind of free space for the moments to develop that develop features to ship and basically that's a and so we have a queue with all the idea to put a margin for humans.
And then if everything is okay, then we rerun the workflow that allows to execute everything. Yeah, it was basically that thing that we had to improve. Also, one of the bottlenecks that we had was we had issues with how we commit to an environment state. So basically, it was one of our... Really bottlenecks is when you have multiple deployments, you need to be, only can commit once, one at a time, the deployment.
So basically, we had a strategy that we create one commit per region. So we have a lot, a pile of commits to push to that, this environment site repository. And that was really a bottleneck. Fortunately, we were not very clever at that time to create the queue for that specific step. And then we had to re -engineering all the process because we had the issue that we didn't do like a FIFO.
So basically, we did implement a hot hobby. And that means that some guys was very unlucky that it was the first one to arrive and was not the first one to get their stuff deployed. Yeah. This is really the story that we have. What was missing at that time was in the presentation and blog post was really the migration to the mobile team. I don't know if you want to go to already that story.
It's a new story that I have. Or you want any other questions? Let's actually get to the mobile story in a second. So going back over your CICD process, I'm curious. First off, are you worried about The GitHub change to private self -hosted runners, they're changing the pricing model. So I guess they're going to charge for self -hosted as well now. Is that going to change your implementation at all?
That's the problem. I can think that if they're starting to charge that, then I need to start to thinking, if they start charging like that, then I need to really be very picky in their SLAs.
So I they how they can charge me a fee for the control plane when their control plane it's not really 99 or doesn't yeah meet the slas it certainly hasn't been lately that's for sure that that is really the thing even even we don't even we get that uh we more or less exclude that from our matrix of of um I already starting to think already seeing some page that shows the SLA in the last couple of months.
And then I'm starting to thinking if you charge me for this, I need to have a better SLA. I cannot stop working. And then for the size of what we have, it's basically or I need to find a different CI tooling and going to that path or even can go more crazy. Don't forget that, okay, we are a special case, or probably not a normal case. That is, we are already in enterprise level.
And we have GitHub hosted enterprise server. So basically, you can bring yourself. And if this is starting to be expensive, probably that's the thing I need to do. I bring down and then try to be less dependent of... The ability of the cloud version and then be my concern. And then I can then point to myself. Yeah, your GitHub is down because of myself and not because of some change in the cloud version.
But yeah, it will be really a concern. And if they force us, it's basically now I will go to my legal team and then check the SLA and let's see if they don't like the SLA. They're starting to get some notice from our legal team to get a recharge back or something. So also you had mentioned, that's a fair statement for sure. You had mentioned looking at Flux originally and then going with Argo.
Are you using anything like Crossplane with Argo yet or no? No. So basically the fact that we don't use Crossplane is basically because it's not in our domain of work. Okay. From what I understand, really, the idea is that cross -plane is a good way to provision a resource using Kubernetes as a native language.
Even I try to understand from Victor, Victor Werczek, that's working with cross -plane, the difference between Terraform and cross -plane, and that's really the reason that... We didn't go deep in crossfire because the main player of Terraform and provision infrastructure is the infrastructure team, not the engineering excellence department. Basically, we are the middle layer. Imagine it's like a lasagna or a burger.
Basically, the team is the lowest band and I am the lettuce. And then we have even up on me, we have the burger and then we have the tomato and they have the other. So basically we have.
These layers and I'm in that layer that I consume service from infrastructure team and then I deliver service to to to the to the platform team and to the to the developers so that's the reason that we don't look for crossplane that's thing that I would like to to experiment but then I need to really have a good case of why did a terraform to use crossplan yeah it's more I guess if you want the infrastructure
Definitions closer to the actual service right so if you want that all defined together and there's there's pros and cons for both ways I would say honestly even in our organization we kind of or I've worked in organizations where we've done both yeah even even listen uh what victor said about the idea of cross plan and how to use cross plan and that year for example is one of the examples that I can tell or I can get a model in terraform that's getting a postgres to be from Mother bless.
But probably if I want to talk with the developer, developer wants to have the minimal settings to change. So basically, I think the WinCross plan can abstract that with their internal resources and say, okay, define this YAML manifest or that YAML resource or, sorry, that CRD in YAML and then the controls and everything will set up everything for you. So at this moment, we don't add this kind of...
Need in terms of organization to have exposed so much the infrastructure to developers. And that's real. Without need, I don't have a way to force a tech to be used. Oh, that's fair. It's always a balance. Yes, it's cool, but I need to really, it makes sense to the case. So tell me about this, the new mobile deployments. How is that going? And how did you set that up?
The mobile team was using Jenkins, but even with more strange setup. So basically, they had a farm of Mac minis, and they basically reconnected with the Jenkins controller master. And they were using that. So basically, they were using groovy, pretty groovy with...
In their checking pipelines plus with fast line for one for the ios team the android team will use a different tooling and that was really that massive because part of the team needs to be operations to understand to upgrade the nodes to fix issues and also again the same issue of isolation when we had cases that the nodes were not very identical and a job was if lands in one machine was passing if lands in the other
Machine was failing or less force between artifacts between was screwing up with node versions hubi versions and was really a mess so lack of productivity and then that was really the idea We need to move to GitHub actions, but plus also find a way to make their compute power or compute resource be very stable, having the same identity that we had in the CI -CD for the microservice.
So that was a multiple deployment and isolate. And at that point, it was really a surprise for me, the end solution, but basically it was this case. So I went to research. That basically I put 4Ks on top of the table. First, using VMs inside of Mac. So basically a product from CircleCI. So basically a company that's already providing GitHub Action Runners for Macs and have the system to work.
And basically it's on top of the one tooling at the start. And basically it's a... A nice tooling to spawn VMs inside of Mac. Then the idea, okay, let's try to use Nix for some independence in isolation. And the other idea was basically also to use AWS. So why not spawn Macs in AWS and use it as a GitHub address? And at last I was thinking, I need to at least think in the way of outsourcing this to a company.
This guy is GitHub Actions, but... I could also think to other GitHub action providers or first to GitHub itself and then to other providers because you can see that Blacksmith and I think Depot are providers of GitHub action trends that you can offload and don't depend of GitHub to have the best performance in your machines. Basically, it was really a poor scientific research. I have three, four hypotheses.
I have one month to test it. And I set like one week for each hypothesis and then try to go. And this is really the first. It was in summer of last year. And also that culminate in the appearance of the... AI native mindset. So in this experimental research, after I collided with teachers, with tooling, I found that I had to use start to create VMs. So first hypothesis.
And then I would decide, well, if I try to create a controller, like I have the idea or already the use case that to have a control for Kubernetes, to run runners. So why I don't have an action runner control for that? And that was really the idea.
Starting to think I did a poc with shell script because it was a very easy command but then I decided okay I have a nice shell script let me do a really nice mvp and then I did my first specification development project so basically I used this idea I then signed to define my specifications in a markdown file what I want, what were the toolings, what were the constraints.
And then use, in this case, was already still using GitHub Copilot and say, I have this idea in this file, let's make a plan. And starting elaborating the plan, creating the plan. And then the process was really that way that after I have the plan, we need to have to -do lists or to -dos for each point. And then basically I force the Copilot to use, go for each to -do or each step. Do the implementation, I review it.
Okay, it's fine. Let's move to the next one. Fine. And then also ask to do a summary of each implementation. So basically to have a history of what I did. Just important for me, but also important to share with the team all this process. And then after one day and a half, I got a control. So basically I had a Mac mini in my desk. I put the control there and it was spawning. And doing the lifecycle of the VM.
So basically, a workflow, it's my runner, execute a niche, drop, tear down the VM, start a new one, register against GitHub, like normal flow. And I will say, yeah, nice. I just need to make some improvements. Then I moved to the Nix. AI, in this case, also it was Copilot, helped me a lot. How to build the recipes with Nix, but it fails tremendously just because of the way Xcode works. It was very annoying to work.
And then I had the issue of how to distribute Xcode. At that time, I was running out of time. AWS was not really an option to investigate. And then I started doing some calculations about cost if I use GitHub as a provider of headers. And surprise, surprise. It was cheap for our use case. So it was basically a matter of after spending three or four months collecting metrics in Jenkins, I say, yeah, we can use it.
And basically it was really the idea. I said, okay, this is the amount of money and comparing the working hours that it's necessary for an engineer to fix this, it's a good balance. And then after I convinced my director that it really makes sense in terms of financial terms.
You approve and then we move and then we move to migration and that this migration was I did with the junior the first thing that we did when this migration was really sit down with the developers and I asked anoint us to to to my junior in this case it was to internities sorry but it was really important to that yeah I asked him go for each um jenkins pipeline they have and start doing a flowchart so basically we
Had a flowchart for each Jenkins pipeline with steps and the steps in the way that what is supposed to do and what are the commands that are executed and then I sit down with the with each uh team from android and and ios and then let's go forward for each step and then try to understand as this makes sense this flow I don't care about what is the command that is executed.
Does this step make sense? Does this test make sense? Does this fork in the logic make sense? And then we also understand some good things that is some workflows or some checking job that we have was already redundant. We could refactor the input and combine in one single pipeline. That was really the idea. And then that was the second time we used the AI to speed up.
And this time I already added the session of agenting coding with my engineering department. So basically the engineering excellence create agenting sessions to teachers all to use AI tools in a more agentic way. So not to auto -complete features, but to give context, to give a goal. Have this guy, this AI, as really a partner to execute. And then at this time, already we was using cloud code.
And then, okay, let's bring these flowcharts, convert to something that is more digestible by AI. So I was able to export as a CSV file. And then, okay, these are CSV files, contains flowcharts of our workflows. Let's build GitHub Action workflows. And documentation. And he's starting doing the old workflow. Was not really exactly what we want, but was close enough.
Imagine that this was really a good best draft, a good first draft. This was really, now just adjust this step, this step, this step, and then we just starting building on that. And then, of course, this really make the work very easy. So we add like 17 workflows to migrate. What is the difference from this migration to the other migration? The other migration, we add like one or two workflows for all the process.
So basically, the issue was to replicate that to use workflows to 700 service. In this case, we have only two service, two repositories, the Android and iOS, but we have multiple workflows to migrate. And then we had to rewrite a lot of stuff. And that was the way that we use AI to basically at the pace of each day we migrate a workflow and then cloud code was able to digest some part of the code base.
So I had two issues. First, parts of the customization that they have in the iOS team was using FastLine that is written in Ruby. I don't know Ruby. So I used it to understand what was that.
What was the logic behind that ruby scripts and had extra features then I'm not very well versatile in the ios test and compiler so I never I'm not irs developer I don't know all the quirks about the the process to compile language compile ios app And I use protocol for that. So basically, it was already in a way that I had an issue in my pipeline. And I tell them, OK, I have problems like this.
I have an issue in my pipeline. This is the idea of the workflow. This is the idea of job. This is a step that's failing. Fetch.
Using github cli in that time even we using mcp probably it's best if you have a cli to tell the ai model to use that cli to fetch the information instead of uh bloat the contacts with mcps use the cli fetch the logs in that section and that let's go investigate what it's what it's failing and it was really able at some point I get get surprised because when I go in deep mode of troubleshooting the model in this case
Even the the agent that is called cloud cloud was able to go to the internet and find github issues about the problem pointing out I have found this issue probably it's about this let's double check and then I double check yeah probably makes sense let's try this change and it was really basically it was a more family language or I cannot what I can say it was like while we were in the past doing Googling.
So you put the problem, you try to find the issues. Here, I have the problem. Also find the issues in the internet and get me back the information and then validate with me and then explore. And basically, after four weeks, we migrate everything. We even add extra features they want. And they were very happy. And still, they are very happy. Of course, the initial costs failed because it was more than I expected.
But for one reason, developers were delivering more. So it was really a situation that the workflows and run are so stable, they can focus more in future and then increase the cost. But it's because they are shipping more features than they did in the past. That's a good cost problem to have, you know? Yeah. Yeah. I think that's the thing I have. Do you have any questions about that topic?
I did want to ask, wrapping up, if someone that's listening, they wanted to pull off a big CICD switch like you did, are there some lessons learned that you could give that they could follow so they don't set their org on fire? I mean, because this is a complex... Lift and shift, right?
Going from Jenkins to GitOps and introducing Argo and all the complexities around, you know, the Mac mini pipelines, like be interested. There's some like core lessons learned that you could impart. So the first thing is, I think you need to, even you have a big pipeline, I bet that you have a small part. So a niche. So try to find. A small part that you can replace and do it in isolation.
In my case, it was basically the pull request validation. It's detached in some way of the big flow. Try on that. If you cannot do that, try to find segmentations that you have in your organization in terms of teams. That's a good way to approach so that you can move parts of your team. Basically, if you have 10 teams or five teams, pick one team and try to go in that way.
So this is a way that you can try to reduce the buster radius. And of course, I think that was a thing that was very important for us, basically doing dogfood. I think it's very unfair for someone that is developing tools for developers not using that in their work. So I think this is really the most important thing. It's doing dogfooding.
And then if you want, try to find in each place if you don't have anything think do you have internal tooling that doesn't uh provide for your final customers that can be bad but for trying to apply this to your internal tooling yeah that makes sense cool where can people find your uh your posts and and where can they reach out to you okay so my posts are in the medium probably you can find them have the links in the description also I have that both was based in the talk that I did.
So also in YouTube, I will have that talk. Also, sometimes I do some publishing in LinkedIn. So you can go there, try to reach me in LinkedIn. Awesome. I'll leave the links for your Medium posts and your LinkedIn and anything else in the show notes. Stefan, thanks for coming on. Really appreciate it. Okay. Thank you. All right. That's my conversation with Stefan Moser.
My biggest takeaway from this one is that good CICD is not just about picking a newer tool. It is about building a delivery system that is predictable. Observable, isolated, and usable enough that engineers can trust it without needing constant help from platform teams. That is really the thread running through this whole episode. They did not just swap Jenkins for GitHub Actions.
They reduced noisy neighbor problems. They standardized runners. They leaned into reusable workflows. They moved deployment towards GitOps. They built their own visibility layer when the platform was not giving them enough, and they rolled it out in a way that let teams build confidence instead of forcing a giant overnight cutover. I also liked that he was honest about what happens when the new system works.
Once deploys get easier, people use them more. Bots start shipping changes, automation starts piling up, and then you discover the next bottleneck, whether that is queuing, fairness, or protecting enough room for humans to still get work out. That is the real platform lesson. Success creates new load.
And the better your self -service story gets, the more you have to think about throughput, guardrails, and the system behavior under trust. The other part I liked was his migration advice at the end. Start with a niche. Reduce blast radius. Dog food your own system first. And if you are building tools for developers, use them yourself first before asking everyone else to bet on them.
That is probably the cleanest takeaway from the whole episode. If you enjoyed this episode, follow Ship It Weekly wherever you listen to podcasts. If you want the show notes, links to Stefan, his write -ups, and the resources we talked about, head over to shipitweekly .fm. Thanks for listening, and I'll see you later this week. Thank you.