Flaky Tests Overhaul
If you've got a substantial set of tests, chances are very high that you've encountered situations where some test outcomes fluctuate between runs, even when there are no changes in the code. This inconsistency is called "flakiness".
Teams often may find themselves repeatedly retrying pipelines containing flaky tests (timeouts are also considered a form of flakiness), trying to make the pipeline green. However, this process often results in wasted engineering hours and CI resources. As the code base and number of teams grow, managing flaky tests becomes increasingly challenging, leading to more potential issues and accumulating technical debt.
The Uber engineering team recently published an article detailing their approach to improve CI stability and address the issue of test flakiness.
Key points:
📍Separate service Testopedia is introduced to visualize history of test execution and test performance characteristics
📍Testopedia is language/repo-agnostic. It operates with the term ‘test entity’. Each test entity is uniquely identified by a “fully qualified name” (FQN) that usually includes a full test address in the repo.
📍Tests can be grouped into realms, each realm is owned by some responsible team.
📍Testopedia analyzes test execution stats (including flakiness, reliability, staleness, execution time) , groups problem tests and triggers a JIRA ticket with the deadline to fix.
📍GenAI integration is a future step to auto-generate fixes for flaky tests. It’s under research now.
As a result, the authors noted that implementing the Testopedia approach significantly improved the reliability of CI and reduced the number of retries. If this tool were available as an open-source project, I would certainly give it a try, but unfortunately, it's not.
However, in the absence of such a tool, what steps can we take on our own to address this issue? Here are some suggestions:
📍Visualize pipeline health by implementing simple monitoring of CI statistics, including the number of retries, execution time, and other relevant metrics. To improve something it must be measurable.
📍Treat problematic tests as work items with clear deadlines for resolution.
📍Prioritize CI issues, recognizing them as critical technical debt that will require attention anyway
📍Implement measures to make retries more difficult or even impossible (validations, webhooks, etc.)
📍Clearly define roles and responsibilities for maintaining CI stability; otherwise there is a risk of collective irresponsibility.
#engineering #ci
Post #23
361
- 👍 3
- 🔥 3