How to Write Your First AI-Driven E2E Test in 15 Minutes

2 views 0 likes 0 comments 17 minutesOriginalTutorial

A practical guide to replacing fragile UI selectors with natural language. Learn how to initialize the e2e framework, combine agent.act/assert with Playwright assertions, and leverage operation caching for near-zero cost regression testing.

#E2E Testing #AI Testing #Playwright #Frontend Testing #Automation Testing #TypeScript #Testing Framework
How to Write Your First AI-Driven E2E Test in 15 Minutes

How to Write AI-Driven E2E Tests with Natural Language: A Practical Guide

What's the most frustrating part of writing E2E (end-to-end) tests? After 8 years in backend development, it's rarely the business logic. The real pain is maintaining UI tests built with fragile locators in Selenium or Cypress. One frontend class rename or button relocation, and your tests fail, forcing you to spend half a day fixing selectors instead of debugging actual bugs.

Recently, I discovered a fundamentally different approach: describe your test goals in natural language and let an AI agent "operate" the page to verify the results. In this tutorial, I'll walk you through tester-army/e2e, a next-generation E2E testing framework. From installation to writing your first complete test case, you'll be able to spin up an AI-assisted E2E testing pipeline in your project by the end.


Prerequisites

Before you begin, ensure your environment meets these requirements:

  • Node.js >= 18 (LTS version recommended)
  • npm or pnpm (package manager of choice)
  • A Large Language Model API Key (supports OpenAI, Anthropic, local models, etc.)
  • Basic TypeScript syntax knowledge
  • Access to your target web application (local localhost or production URL)

Note: No prior Playwright experience is required. The framework abstracts the underlying engine so you can focus purely on test logic.


Quick Start: Initialize Your Project

Open your terminal, navigate to your project directory, and run a single command:

bash 复制代码
npx e2e init

This interactive setup will guide you through:

  1. Engine selection: web (browser testing via Playwright) or mobile (iOS/Android simulators)
  2. Model provider configuration: Input your API Key (multi-provider supported)
  3. Auto-generation: Creates a default config file and example test files

Why use interactive initialization? Different projects require vastly different testing engines and model providers. Interactive configuration prevents manual parameter errors and ensures out-of-the-box readiness.

Once initialized, your project structure will include:

复制代码
└── tests/
    └── example.e2e.ts  ← auto-generated example test

A e2e.config file will also appear in your root directory, storing your engine and model settings.


Understanding the First Test Case

Open the generated example file. You'll see something like this:

typescript 复制代码
// tests/checkout.e2e.ts
import { test, expect } from 'e2e';

test('a member upgrades to Pro', async ({ app, agent, screen }) => {
  await app.open('/settings/billing');

  // Agent uses natural language to describe actions
  await agent.act('upgrade the workspace to the Pro plan');
  await agent.assert('the invoice preview shows a prorated amount');

  // Traditional locator assertions used as a validation fallback
  await expect(screen.getByRole('status')).toContainText('Pro');
});

Here's a breakdown of the three core objects:

Object Purpose Explanation
app Navigation app.open() opens a specific URL, similar to Playwright's page.goto()
agent AI Agent agent.act() describes page operations in natural language; agent.assert() describes the expected outcome
screen Element Query Standard locator API fully compatible with Playwright, used for precise assertions

How the Agent Works (Smart & Cost-Effective)

This is the framework's most clever design: on the first run, agent.act() calls the LLM, parses your natural language instruction, and generates specific page operations (clicks, inputs, etc.). These steps are immediately cached.

On subsequent runs, if the UI structure remains unchanged, the framework directly replays the cached operation sequence without calling the LLM. Only when frontend updates invalidate the cache does it trigger a new model call. This guarantees intelligent test generation while strictly controlling API costs.


Practical Walkthrough: Writing Your First Real Test

Let's assume you have an internal admin system and want to test the scenario: "After logging in, can a user successfully create a support ticket?" We'll write it step-by-step.

Step 1: Create the Test File

Create tests/create-ticket.e2e.ts in the tests/ directory:

typescript 复制代码
import { test, expect } from 'e2e';

test('user creates a ticket after login', async ({ app, agent, screen }) => {
  // 1. Open the login page
  await app.open('http://localhost:3000/login');

  // 2. Let the agent perform the login operation
  await agent.act('enter username admin@example.com and password 123456, then click the login button');

  // 3. Wait for redirect to the dashboard after successful login
  await expect(screen.getByRole('heading', { name: 'Dashboard' })).toBeVisible();

  // 4. Let the agent navigate to the ticket page and create a ticket
  await agent.act('click Ticket Management in the sidebar, then click New Ticket');
  await agent.act('fill title as Test Ticket, description as This is an AI-generated ticket, set priority to Medium, then submit');

  // 5. Verify the ticket was created successfully
  await agent.assert('the page displays a ticket record titled Test Ticket');
  await expect(screen.getByText('Test Ticket')).toBeVisible();
});

Pro Tip for agent usage: Break complex workflows into multiple agent.act() calls instead of cramming everything into one prompt. This makes debugging significantly easier when a specific step fails.

Step 2: Run the Test

bash 复制代码
## Run a single test file
npx e2e run tests/create-ticket.e2e.ts

## Run all tests in the directory
npx e2e run tests/

The first execution will invoke the LLM to generate the operation sequence. Subsequent runs will replay the cache. As long as your frontend UI hasn't changed, you'll notice the tests execute blazingly fast.


Common Pitfalls & FAQ

1. What if the agent fails to execute an action?

Check the following:

  • Page readiness: Is the page fully loaded? Use await expect(screen.getByRole(...)).toBeVisible() before agent.act() to wait for target elements.
  • Prompt clarity: Is your natural language description unambiguous? Specify exact button labels or element names.
  • Cache invalidation: If recent UI changes broke the flow, delete the cache directory (usually .e2e/cache/) to force the agent to re-learn.

2. Will API costs skyrocket?

No. As long as the UI remains stable, subsequent runs incur zero model invocation costs. Model calls only trigger when UI changes invalidate the cache. Daily regression testing becomes essentially free.

3. Can I disable telemetry?

The framework anonymously collects CLI execution metrics by default. If you prefer not to share this, run:

bash 复制代码
npx e2e telemetry disable
## or set the environment variable
export E2E_TELEMETRY_DISABLED=1

4. Framework Stability Note

The framework is actively progressing toward v1.0. Minor API or configuration changes may occur in patch releases. Pin your version in package.json, and always run existing tests locally after upgrading to verify compatibility.


Summary & Next Steps

Here's what you've accomplished today:

  1. Initialized the project: npx e2e init for interactive engine & model setup
  2. Mastered the core objects: app (navigation), agent (natural language operations), screen (precise assertions)
  3. Wrote tests: Combined agent.act() + traditional expect() for reliable business logic coverage
  4. Executed tests: npx e2e run — learn once, replay thereafter
  5. Managed costs: Leveraged caching to achieve zero-cost daily regression

Recommended Next Step: If your team maintains multiple projects requiring E2E coverage, integrate e2e directly into your CI/CD pipeline to run tests automatically before PR merges. The framework also offers a @e2e-dev/github package that posts test results directly as PR comments, streamlining team collaboration.

Full documentation is available at e2e.tester.army/docs, and an offline copy ships in node_modules/e2e/docs after installation. The team is highly active on Discord if you run into roadblocks.

If traditional E2E maintenance has been slowing down your releases, give this AI-driven approach 15 minutes. Feel free to drop your questions or experiences in the comments below!

Last Updated:

Comments (0)

Post Comment

Loading...
0/500

No comments yet, be the first to comment!