JTjason.teixeira() Docs
services Book a call →
Home / Docs / Test Automation & QA / How AI Is Changing Test Automation: Real vs Hype
Test Automation & QA

How AI Is Changing Test Automation: Real vs Hype

AI test-automation demos are dazzling and half of it is smoke. Here is the part that actually holds up.

Watch enough AI testing demos and the pattern shows up fast. An agent writes a passing test from a plain-English sentence, the crowd claps, and nobody asks what happens on the second run. AI does move the needle on a few boring parts of test automation. It also makes other parts worse if you trust it too far. Here is where the line falls.

#The part that actually holds up: maintenance

The durable win is not writing new tests. It is keeping the ones you have alive. Selector-healing is the clearest case. When a button's id changes and a normal locator breaks, an AI layer can re-find the element by its surrounding text, role, and position, and keep the test green instead of failing the suite on a cosmetic change. That is real leverage, because flaky-locator churn is where teams burn the most hours.

Triage is the other one. When forty tests fail, AI is good at clustering them: these thirty share one root cause, these ten are separate. A morning of log-reading drops to a few minutes. Both wins have the same shape. A human already decided what correct behavior is, and the AI does pattern-matching on top of that decision.

#The hype: generate a full suite from a sentence

The demo that gets funded is "describe your app, get a test suite." It breaks on the oracle, the part of a test that knows what the right answer is. AI will happily generate a hundred tests that assert the app does whatever it currently does. That is a screenshot with extra steps, and it will lock in a bug as expected behavior.

A concrete tell: ask the generated test what it would catch. A good checkout test asserts the total equals the sum of line items plus tax. An AI-generated one often asserts the total field is visible and non-empty. The first catches a pricing bug. The second passes while the company loses money. When you audit AI-written tests, grep the assertions. If most of them check existence and rendering instead of specific values and state transitions, you have coverage theater.

#The honest middle: a drafting partner with a supervisor

Between the win and the hype is the actual daily use. AI is a fast first-draft writer for tests you were going to write anyway. Hand it a function and two example cases and it scaffolds the edge cases you would have typed by hand: empty input, huge input, a unicode name. That is a real speedup on the tedious 80 percent.

The catch is that it drafts confidently wrong assertions right next to the good ones, so it needs a supervisor who knows the domain. The workflow that holds up is simple. AI drafts, the human owns the oracle. You keep judgment over what correct means and let the model handle the typing around it. Teams that invert this, letting the model decide correctness and using humans only to skim, ship suites that are green and meaningless.

▸
AI is strong at the mechanics: healing locators, clustering failures, drafting edge cases. It is weak at deciding what correct actually means. Keep that decision human and the rest is fair game.

#The bottom line

Point AI at the maintenance drag and the first-draft tedium and it earns its keep this quarter. Let it own the assertions and you get a suite that stays green while your product breaks, which is worse than no suite because you trust it. Better models have not moved this line. The machine can do the work around the test. A human still has to know what the test is for.

Want this on your product, not just in theory?
Get a free mini-eval on your live AI feature, or book a call to talk it through.
Build your plan → 2 minor book a call →
© 2026 Jason Teixeira · Sage Ideas LLC · Documentation home · privacy · terms