Skip to content
Testavi
Blog

Why your end-to-end tests break after every redesign

Shaig Khaligli

Shaig Khaligli, founder of Testavi
· 4 min read

Most end-to-end test failures aren't bugs. They're the test noticing that a button moved, a class name changed or a page loaded half a second slower than usual. The product still works. The test doesn't.

Before Testavi I built test automation at SAP, and a lot of that work was rewriting tests after UI changes rather than catching real bugs. This post is about why that happens and what you can do about it, whether or not you ever use our product.

Key takeaways

  • Most E2E breakage comes from tests describing how the page is built instead of what the user is trying to do.
  • Google found about 1.5% of all its test runs were flaky, and almost 16% of its tests had some flakiness.
  • Role-based locators, dedicated test attributes and fewer fixed waits cut a lot of noise if you write tests in code.
  • Intent-based tests, written in plain English and run by an agent that reads the page visually, avoid the selector problem altogether.

Tests break because they describe the page, not the user

An end-to-end test is a test that drives your real app through a full user journey, the way a customer would, from the first page to the result. A typical Playwright or Cypress test says something like "click the element with the class .btn-primary inside #checkout-form". That's a description of your HTML. Your customer doesn't know or care what .btn-primary is. They're looking for the button that says "Pay now".

So when a designer swaps the button component, or a developer renames a class during a refactor, the test fails even though a human could still check out without thinking about it.

The framework authors are upfront about this. Playwright's documentation recommends role-based locators and warns that CSS and XPath selectors "can break when the DOM structure changes" (Playwright docs). Cypress's best-practices guide lists brittle selectors as a common mistake and suggests adding data-* attributes purely for tests (Cypress docs).

That advice works. It also means every component in your app now carries attributes that exist only to keep the test suite alive, and somebody has to remember to add them.

Flakiness is the other half of the problem

Even when the selectors are fine, end-to-end tests fail randomly. An animation hasn't finished, an API responds slowly, a cookie banner shows up on one run and not the next.

Google measured this across its own codebase. About 1.5% of all test runs reported a flaky result, almost 16% of tests had some level of flakiness, and about 84% of the pass-to-fail transitions they saw involved a flaky test (Google Testing Blog, 2016).

That last number is the one that hurts. If most of the times a test goes from green to red it's noise, people stop trusting red. Then a real checkout bug sits in a failing build that everyone assumed was "just flaky again".

What you can do if you write tests in code

If your team is committed to Playwright or Cypress, these changes remove a lot of the pain:

  1. Locate elements the way a user would. Prefer getByRole('button', { name: 'Pay now' }) over CSS selectors. If the label changes, that's a real change worth reviewing.
  2. Add stable test attributes to the few elements that need them. Payment buttons, cart totals, the things your most important tests depend on. Not every div.
  3. Stop using fixed waits. cy.wait(2000) is a guess. Wait for the thing you need to be visible or for the request you depend on to finish.
  4. Quarantine flaky tests instead of retrying forever. Retries hide the problem. A short list of known-flaky tests that someone owns is better than a suite where everything passes on the third try.
  5. Keep end-to-end tests for the journeys that matter. Signup, checkout, subscription changes. Everything else is usually cheaper to cover with unit and integration tests.

Or describe intent and let an agent handle the clicks

The other route is to stop describing the page at all.

A test like "add the first pair of running shoes in size 42 to the cart and check out as a guest" says nothing about your markup. An agent that reads the page visually can find the size picker and the checkout button the same way a person would, so a redesign that keeps the flow intact doesn't break the test. Interruptions such as cookie banners and promo pop-ups get closed instead of failing the run.

This is how Testavi works, and it's the reason we built it. The trade-off is honest: you give up some of the fine-grained control code gives you, like mocking a specific network response. For the long click-through journeys that break most often, that's usually a good trade. We've written more about it in Testavi vs Playwright and Cypress.

Where to start

Pick the one flow that would cost you the most money if it broke tomorrow. For most online businesses that's checkout. Write it as a single test, in whatever tool you choose, and make sure a failure tells you what the user couldn't do, not which selector went missing.

If you'd rather we set that up for you, our free 30-day pilot starts exactly there: we write and maintain tests for your core flows so you can see whether the approach fits before paying anything. And if checkout is the flow you're worried about, our guide to checkout flow testing covers what to check.

Protect your most important user flows

We write and maintain tests for your core flows during a free 30-day pilot. No credit card. Plans start at €299/month after that.