How do you migrate dozens of services to a new infrastructure without taking down production?
In the Wayo* case study, we had one month to do it.
The task wasn't just about moving infrastructure from an old Kubernetes cluster to a new one. The services also had to be adapted to an updated internal engineering platform.
The most interesting part was the dependencies. Each service had its own set of configurations, databases, Redis and other KV stores, brokers, internal services, and external APIs. Moreover, some connections weren't obvious from the original configuration and were only discovered during analysis.
So for each service, we defined a unified preparation process and a set of readiness criteria: startup, health checks, accessibility within the cluster, connections to storage systems and APIs, logs, and key user scenarios.
In production, the old and new versions ran in parallel for some time. First we validated the new version, then switched traffic over to it, and only after a monitoring period did we decommission the old one. If something went wrong, we had the option to roll back.
By the end of the fourth week, all services scheduled for migration had been moved to the new platform and launched in the new cloud environment. The user-facing application continued running without interruption.
We covered the dependency audit, the migration process, and the traffic switchover in detail in our case study: https://evrone.com/cases/wayo
And what stage of migration usually turns out to be the trickiest for you?
*The company name and certain identifying details have been changed because the project was completed under an NDA.
Post #763
145

- ❤ 5
- 🔥 3
- 👍 1