TGViewer
TechLead Bits TechLead Bits @techleadbits · 517 subscribers
Post #23 361
Flaky Tests Overhaul

If you've got a substantial set of tests, chances are very high that you've encountered situations where some test outcomes fluctuate between runs, even when there are no changes in the code. This inconsistency is called "flakiness".

Teams often may find themselves repeatedly retrying pipelines containing flaky tests (timeouts are also considered a form of flakiness), trying to make the pipeline green. However, this process often results in wasted engineering hours and CI resources. As the code base and number of teams grow, managing flaky tests becomes increasingly challenging, leading to more potential issues and accumulating technical debt.

The Uber engineering team recently published an article detailing their approach to improve CI stability and address the issue of test flakiness.

Key points:
📍Separate service Testopedia is introduced to visualize history of test execution and test performance characteristics
📍Testopedia is language/repo-agnostic. It operates with the term ‘test entity’. Each test entity is uniquely identified by a “fully qualified name” (FQN) that usually includes a full test address in the repo.
📍Tests can be grouped into realms, each realm is owned by some responsible team.
📍Testopedia analyzes test execution stats (including flakiness, reliability, staleness, execution time) , groups problem tests and triggers a JIRA ticket with the deadline to fix.
📍GenAI integration is a future step to auto-generate fixes for flaky tests. It’s under research now.

As a result, the authors noted that implementing the Testopedia approach significantly improved the reliability of CI and reduced the number of retries. If this tool were available as an open-source project, I would certainly give it a try, but unfortunately, it's not.

However, in the absence of such a tool, what steps can we take on our own to address this issue? Here are some suggestions:
📍Visualize pipeline health by implementing simple monitoring of CI statistics, including the number of retries, execution time, and other relevant metrics. To improve something it must be measurable.
📍Treat problematic tests as work items with clear deadlines for resolution.
📍Prioritize CI issues, recognizing them as critical technical debt that will require attention anyway
📍Implement measures to make retries more difficult or even impossible (validations, webhooks, etc.)
📍Clearly define roles and responsibilities for maintaining CI stability; otherwise there is a risk of collective irresponsibility.

#engineering #ci
  • 👍 3
  • 🔥 3
More from @techleadbits
  1. Oct 7, 2026AI & Repository Strategy For many years, there has been an ongoing debate between monorepo…
  2. Oct 1, 2026Tracer Bullets Continuing the topic from the previous post, let's talk in more detail abou…
  3. Sep 28, 2026Why Software Factories Fail "Read the Code!" is one of the key ideas from Dex Horthy's tal…
  4. Sep 21, 2026Illustrations from The Culture Map showing how different cultures compare on the scales. #…
  5. Sep 21, 2026The Culture Map Have you ever worked in international distributed teams? Or collaborated w…
  6. Sep 10, 2026Loop Engineering from First Principles Continuing the topic of Loop Engineering, I'd like…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →