Agentic E2E Testing: What Works in 2026
How AI agents plan, run and repair browser tests, where Playwright's test agents fit, when self-healing hides bugs, and what evidence a run should leave.
Agentic E2E testing works in 2026 when the AI agent explores the app, drives a real browser and drafts tests, and the pass or fail decision rests on checks the agent cannot quietly change, with evidence a person can review. It goes wrong when the same agent writes the test, runs it and "repairs" it until it passes. AI-written tests are now mainstream: Linear wrote on 21 September 2026 that "agents now write the majority of our tests" (its unit and API test suite, not browser tests), and Next.js 16.4, released on 6 October 2026, hands upgrade work to coding agents along with verification steps. The question is no longer whether AI can write tests, but who checks the tests.
This article explains how agentic testing works, where Playwright's test agents fit, when self-healing hides bugs, and what a trustworthy AI test run should leave behind.
What is agentic E2E testing?
End-to-end (E2E) tests check a whole user journey in a browser: sign up, add to cart, pay, see the confirmation. Traditionally a developer writes each step as a script, or records clicks and replays them. Both break when the page changes.
A working definition: agentic E2E testing uses an AI agent to drive a real browser through user flows, generate or repair tests and report evidence, with a human or a fixed rule deciding what counts as a pass.
In practice an agentic test system does some or all of the following:
- Explores the site to find pages, forms and roles (visitor, member, admin)
- Plans which journeys to test for each role
- Executes each step in a real browser, deciding what to click or type from what it sees
- Checks results against expected outcomes
- Reports what passed, what failed and why, with screenshots and logs
The appeal is coverage. An agent can try paths nobody wrote a script for, and keep tests up to date as the interface changes. The risk is that a system clever enough to adapt is also clever enough to adapt its way around a real bug.
How are Playwright's test agents different from recorded tests?
Playwright, Microsoft's open-source browser testing framework, introduced Playwright Test Agents in version 1.56 on 6 October 2025. They are three agent definitions you add to a project with npx playwright init-agents, for use with clients such as VS Code, Claude Code or opencode:
| Agent | What it does (Playwright's description) | Output |
|---|---|---|
| Planner | "explores the app and produces a Markdown test plan" | A readable plan a person can edit |
| Generator | "transforms the Markdown plan into the Playwright Test files" | Ordinary Playwright test code |
| Healer | "executes the test suite and automatically repairs failing tests" | Fixed tests, or a skipped test if it believes the feature is broken |
Since version 1.62 (July 2026), Playwright also bundles its MCP server, which lets an AI agent control a browser directly through npx playwright mcp.
The difference from recorded tests is where the intelligence sits. A recorded test is a fixed list of clicks. Playwright's agents produce a plan first, then code you can read and commit, so the result is still a normal test suite that runs without AI in CI. That design choice matters: the plan is the point where a person can say "yes, these are the journeys that matter" before any code is generated.
When does "self-healing" hide real bugs?
Self-healing means a test repairs itself when it fails, usually by finding a new selector after a button moved or was renamed. That is useful when the failure is cosmetic. It is dangerous when the failure is the bug.
Consider three failures:
- The "Buy" button's label changed to "Purchase". Healing the selector is correct. Nothing is broken for users.
- The "Buy" button is now hidden behind a cookie banner on mobile. A healer might scroll, dismiss the banner and click. The test passes; real users who do not dismiss the banner cannot buy.
- The order confirmation now shows the wrong total. If the healer "fixes" the expected total to match what appears, the test passes and the bug ships.
Playwright's own documentation acknowledges the line: its healer can produce a skipped test when it believes the functionality itself is broken, rather than forcing a pass. That is the right instinct, and it shifts the problem to the review: someone has to notice skipped and healed tests, not just the green total.
Rules we apply to any self-healing system:
- Healing may change how a step finds an element, never what a test expects
- Every healed or skipped test is listed in the report, not hidden in the pass count
- Expected values (totals, statuses, permissions) are written by people or derived from the specification, not from what the page currently shows
- A test that needed healing twice in a row is flagged for a person to look at
What evidence should an AI test run leave behind?
A green tick from an AI agent is a claim, not proof. A trustworthy run leaves evidence that someone who was not there can check. Use this as a checklist:
| Evidence | Why it matters |
|---|---|
| The test plan, by role | Shows what was tested and what was not |
| Screenshots at key steps | Lets a person see the screen the agent judged |
| The checks and their results | Shows that the pass decision followed a rule |
| Console errors and failed network requests | Catches problems the screen hides |
| A list of healed, retried and skipped steps | Exposes the places where the system adapted |
| Reproduction steps for each defect | Lets a developer confirm and fix it |
| Environment details | Browser, viewport, build version and test account |
The single most important item is separation between the agent that acts and the check that decides. If the agent that clicked "Pay" also gets to decide whether the confirmation page "looks right", you are trusting its judgement twice. Machine-checkable assertions (the total equals the cart sum, the status is "Paid", the admin page returns 403 for a member) are harder to argue with.
Where do humans still need to judge the result?
Agents are good at volume and at noticing that something changed. People are still needed for:
- Choosing what matters. Which journeys are critical for this business, and which roles must never see which data.
- Judging defects. Whether a failure is a bug, an intended change or a test mistake.
- Usability and content. A page can pass every check and still confuse a customer.
- Approving healed tests. Accepting that a changed expectation is correct.
- Signing off a release. Deciding that the evidence is enough to ship.
This mirrors how we treat AI-written code in general; see how to review AI-written code. The agent does the work; a person reads the evidence and decides.
How we handle this on client projects
We develop and review code every day with Claude Code, an AI coding agent, and a person reviews every change before it ships. Tests written with AI help get the same treatment as any other code: we read what they expect before we trust what they report.
We are also building our own agentic browser E2E testing engine. In it, a browser agent produces the test plan, runs it, judges the results and writes the report. The engine is in development; we are not quoting results for it, and it is not a product you can buy today.
The principles in this article are the ones we hold ourselves to: pass decisions an agent can't quietly change, and evidence a person can read. If your app was built quickly with AI tools and you are not sure what has been tested, our vibe-coded app security checklist is a good place to start.
Frequently asked questions
Can AI write all my end-to-end tests?
AI can draft most of them, and teams such as Linear report that agents now write the majority of their tests (in Linear's case, unit and API tests rather than browser tests). A person should still choose the critical journeys, review the expected results and read the evidence from each run. Fully unattended test writing tends to test what the app does, not what it should do.
What is a self-healing test?
A self-healing test repairs itself when it fails, usually by finding a new way to locate an element that moved or changed. It saves time on cosmetic changes, but it should never change what the test expects, or it can turn a real bug into a pass.
Playwright MCP vs Playwright test agents: what's the difference?
Playwright MCP is a server that lets an AI agent control a browser directly, step by step. Playwright Test Agents (planner, generator and healer) use the agent to produce a test plan and ordinary Playwright test files, which then run in CI without AI.
Do agentic tests replace manual QA?
Not entirely. They cover more paths and keep up with interface changes better than scripts, but people still judge usability, decide whether failures are bugs, and sign off releases. Agentic testing changes manual QA from clicking through to reviewing evidence.
Sources
All checked October 2026.
- Linear, CI bottleneck reworked, 21 September 2026
- Next.js blog, Next.js 16.4, 6 October 2026
- Playwright, Release notes: version 1.56 (Playwright Test Agents) and version 1.62 (bundled MCP server)
- GitHub, Playwright v1.56.0 release, 6 October 2025
- Playwright, Test agents documentation
Want to know how your app would be tested?
Building or rebuilding a web app and want to know how it will be tested? Ask us how we would approach yours through the project request form, or email dwkim@nqsolution.kr. Our web development service page explains how we build and hand over.

