Run Playwright Tests Against an AI-Generated UI
Test a UI that keeps changing under you, with Playwright tests that lean on roles and text instead of fragile selectors.
You will write Playwright tests that survive an AI tool regenerating your UI. The trick is to find elements by their role and visible text. Avoid CSS classes and DOM structure, since that is the stuff a model changes when it rewrites the markup. This guide sets up Playwright and writes a couple of resilient tests. Budget about thirty minutes.
#Before you start
- Node.js 18 or newer installed.
- A web app you can run locally with a known URL (the examples use http://localhost:3000).
- Basic comfort with running commands in a terminal.
#Install Playwright
Run the init command in your project. It installs the test runner, drops in a config and an example test, and offers to download browsers. Say yes to the browser download. You need it to run anything.
npm init playwright@latest#Point the config at your app
Set a baseURL so tests can navigate with short paths like /. Turn on webServer so Playwright starts your app before the run and shuts it down after. This keeps the suite self-contained, so you do not have to remember to start the server.
import { defineConfig } from '@playwright/test';
export default defineConfig({
testDir: './tests',
use: {
baseURL: 'http://localhost:3000',
},
webServer: {
command: 'npm run dev',
url: 'http://localhost:3000',
reuseExistingServer: !process.env.CI,
},
});#Find elements by role and visible text
This is the whole point. Use getByRole, getByLabel, and getByText to find elements the way a user would. A button is still a button after the AI rewrites your markup, even if the class name and wrapper divs all change. Avoid page.locator('.some-class'). That is the first thing to break.
import { test, expect } from '@playwright/test';
test('user can sign in', async ({ page }) => {
await page.goto('/login');
await page.getByLabel('Email').fill('user@example.com');
await page.getByLabel('Password').fill('hunter2');
await page.getByRole('button', { name: 'Sign in' }).click();
await expect(
page.getByRole('heading', { name: 'Dashboard' })
).toBeVisible();
});#Run it and use UI mode to debug
Run the suite from the terminal. When something fails, open UI mode to step through the run and see the page state at each action. The codegen command is also handy. It clicks around your app and writes role-based locators for you, which gives you a starting point to clean up.
npx playwright test
npx playwright test --ui
npx playwright codegen http://localhost:3000#Run it in CI
Add a workflow so the tests run on every pull request. The Playwright setup installs browsers and runs the suite. A failing test fails the job, which is your gate against a UI regeneration quietly breaking a flow. Keep any test credentials in repo secrets, never in the workflow file.
name: e2e
on: [pull_request]
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 20
- run: npm ci
- run: npx playwright install --with-deps
- run: npx playwright test
env:
TEST_USER_PASSWORD: ${{ secrets.TEST_USER_PASSWORD }}#Watch out for
- Role and text locators only work if the AI-generated UI is reasonably accessible. If buttons are divs with no labels and inputs have no associated text,
getByRolefinds nothing. Treat that as a signal that the markup is bad and fix the UI. Do not reach back for brittle CSS selectors. - Text locators are sensitive to copy changes. If the AI renames 'Sign in' to 'Log in', the test breaks. That is usually the test doing its job. Expect to update the expected strings when wording legitimately changes.
- Do not chase flakiness with fixed sleeps. Playwright auto-waits for elements, so use web-first assertions like
expect(...).toBeVisible()instead ofpage.waitForTimeout. Hardcoded waits are the main source of flaky suites.
#What you built
You have a Playwright suite that finds elements by role and text, so it keeps passing when an AI tool rewrites the markup underneath it. It runs on every pull request. Tests break only when behavior or copy actually changes, which is the kind of failure you want. Next, add data-testid attributes to the handful of elements with no good accessible name, and grow the suite to cover your other critical flows.