"Even a small change could mean waiting close to an hour for CI, and then the release could still get stuck because a migration failed. We can't operate like this any longer." - Client representative
A large European streaming service for web, mobile, and connected devices.. Infrastructure Scale: Dozens of Java services on Kubernetes, many CI pipelines every day.
2 Alpacked DevOps/Platform engineers.
~8 weeks.
Our client is a European streaming platform with an established market position and a large audience. It delivers video on the web, in mobile apps, and on connected devices. The backend consists of dozens of Java services: they're built with Maven, released through GitLab CI/CD, and run on Kubernetes. Several development teams ship updates frequently, so pipelines run many times a day.
The client's tools themselves are modern. The problem was how they were wired together. The delivery process had gradually accumulated new services and modules, and eventually it became the main constraint on the speed of the entire engineering team. The client wanted to release more often, but without taking on more operational risk.
Two of our engineers worked alongside the client's development and platform teams. Over 8 weeks, we optimized the Maven build, redesigned GitLab CI, put artifact management in order, introduced GitOps with ArgoCD, rebuilt how migrations run, and took all of it to production.
30–60 → 3–6 min CI build
0 migrations in init containers
1 image for all environments
A developer pushes a small change and waits anywhere from half an hour to an hour for CI feedback. If something's wrong, they fix it and wait just as long again. At the deployment stage, another trap could be waiting: a failed migration in an init container kept new pods from starting and blocked the entire rollout. After a failure, getting back to the next attempt often dragged on because of the long build cycle. The client couldn't stop development for a large rewrite. So the new CI and CD had to work with the existing Java/Maven codebase, Kubernetes, and GitLab, and they had to be introduced incrementally without halting releases. When we took the process apart, we found three technical causes of the problem.
1
Challenge #1: The Build Did Too Much for Every Change
In practice, the Maven build behaved like a monolith. The pipeline was static, so modules the change didn't touch still went through expensive build stages. Dependency caching was poorly organized: libraries were downloaded over and over, and the same work was repeated across modules and stages.
2
Challenge #2: The Image Wasn't an Immutable Artifact
The container image could be rebuilt between environments. That meant production could end up running a different image from the one tested in earlier stages. And every rebuild added yet another long build.
3
Challenge #3: Deployment Was Tightly Coupled to the Database Schema
Migrations ran at pod startup, so schema changes, application startup, and the rollout were merged into a single process. A release could break either because of the migration itself or because of the wrong deployment sequence. Controlling the order of steps and handling failures in that setup was difficult.
"Even a small change could mean waiting close to an hour for CI, and then the release could still get stuck because a migration failed. We can't operate like this any longer." - Client representative
First, we tried the obvious: we added runners and tuned individual jobs. The gain was marginal because the process architecture stayed the same. A large codebase refactoring wasn't an option either, since development had to continue. That left one path: change the delivery mechanics themselves on top of the existing code. So we took on building the CI/CD process from scratch, from commit to production, and focused the work on three areas.
Build Only What's Needed We split the Maven project into smaller parts that can be built independently. The pipeline is generated for each specific change. Dependencies come from the cache, and build outputs are passed between stages without being built again.
Build Once, Deploy Everywhere CI builds the image once, and that same image is promoted across environments. Production receives the same container image that passed checks in the earlier environments.
GitOps, and Migrations Separate from the Application We moved the deployment logic out of CI scripts and into ArgoCD. We moved database migrations into dedicated Kubernetes Jobs and set the "migration first, then service" order with Sync Waves. We make changes like these as part of our Kubernetes consulting work, when a team already has a cluster but releases on it are painful.
CI, attempt #1: optimize Maven jobs and add runners. Almost everyone starts here, and so did we. Things got a bit better, but the problem didn't go away: the pipeline kept rebuilding overlapping components and doing work it could do without. You can't fix an architectural problem with more powerful runners.
CI, attempt #2: split the build and generate the pipeline per change. It worked. The gain came from three things together: splitting the build into smaller units, clear artifact boundaries between them, and a pipeline built according to which parts of the code a change touched. That's what delivered the bulk of the speedup. This kind of process breakdown, rather than tuning individual jobs, is where we start our CI/CD consulting.
CD, attempt #1: retry failed deployments. When a deployment failed because of a migration, it was simply restarted. Sometimes that helped, but the root cause remained: the migration was baked into pod startup. As long as that's the case, every release with a schema change is a gamble.
CD, attempt #2: move migrations out of init containers entirely. It worked. Now a migration is a dedicated Kubernetes Job with its own status and logs. Through Sync Waves, ArgoCD runs it before deploying the service. If the migration fails, the failure is visible immediately and in one place, and the new service version isn't deployed until the problem is fixed. New application pods no longer get stuck on repeated attempts to run the migration in an init container.
The hardest part was reworking a delivery process that had formed around a large Java/Maven codebase without interrupting the releases that were going out at the time.
First, we had to map the dependencies between modules correctly. If you parallelize the build without understanding those relationships, it's easy to end up with inconsistent artifacts: a module gets built against an old version of a dependency, and that only surfaces once it's in an environment.
Second, moving the migrations. The switch from init containers to dedicated Jobs had to preserve the same guarantee: the schema is updated before the new application version starts working with it. A mistake in the order of steps here would mean a broken release, so during the move, our first priority was keeping that sequence safe.
Dynamic child pipeline generator. At the start of the pipeline, a dedicated job determines which components a change touched and generates the child pipeline configuration. Only the components and checks needed for that change make it into the child pipeline.
Image promotion automation. After a successful CI run, a script updates the image reference in the environment configuration and points it to the exact immutable image CI built. ArgoCD picks up the change and brings the cluster to the desired state, so a separate build for each environment is no longer needed.
Before
After
Just fill the form below and we will contaсt you via email to arrange a free call to discuss your project and estimates.
Language & build: Java, Maven CI/CD: GitLab CI/CD, GitLab dynamic child pipelines, GitLab artifacts Containers & orchestration: Docker, Kubernetes, Helm, container registry GitOps: ArgoCD
The client's codebase, Kubernetes, and GitLab stayed the same. What changed were the delivery mechanics: the pipeline no longer does unnecessary work, and releases follow a clear order. Here's what that delivered.
1
Delivery Speed
The typical CI build went from 30–60 minutes down to 3–6, roughly 10x faster. Developers see build results within minutes instead of waiting up to an hour. The release pipeline no longer repeats a full build for each environment.
2
CI Resources and Team Time
Runners no longer repeat Maven work or rebuild images for each environment, so there's less unnecessary compute in CI. Thanks to GitOps, engineers spend less time manually troubleshooting failures and intervening during releases.
3
Release Reliability
The same image goes all the way from dev to production, and ArgoCD continuously reconciles the cluster with the state defined in Git. A migration failure no longer hides inside pod init containers: it's visible in a dedicated Job, so the cause is easier to find, and recovery or a rerun is easier to plan.