For QA engineers and engineering leads

Cover the journeys you never had time to script.

Write the goal in plain language. The agents drive a real device and a real browser to reach it, and report every dead button, broken price and empty title they pass on the way. What passes becomes a replay you can schedule.

MobileWeb
Web
Mobile
One goal, both surfaces6 stages · 7 findings · 3 on mobile
Mobile · Confirmed screenProof for 01
Stage 04 · Cart

Opened the cart and tried the advertised code

MobileTRAIL10 applied. 10% came off the total.
WebTRAIL10 returned a 500 from the coupon service. The advertised discount never applied.
On mobile: 3 of 7 findings
2 friction findings1 differenceon mobile
Every stage, both sidesMobile above, Web below
01 Home
02 Shop
03 Product
04 Cart
05 Checkout
06 Confirmed
What you get

Four things you get from one run

01

One goal, both surfaces

The same sentence of intent given to a real device and a real browser. Two runs, one vocabulary, and a report that says which surface each defect was on.

02

A run you did not write

The agents work out the route themselves. Nothing is pinned to a selector or a resource id, so a redesign does not red the suite the next morning.

03

The hygiene nobody scripts

Dead buttons, NaN prices, placeholder copy still live, stale copyright years, empty titles. The things a script would never think to check.

04

A replay when it passes

Promote a passing run and it re-runs exactly, with no model call at all. Classic automation reliability, none of the upkeep.

From the run above
Stages walked, per surface
6
Defects raisedall four on the browser run
4
Test scripts written
0
The words on this page

Four terms do most of the work. None of them are ours to redefine in a meeting later.

Goal

The outcome you want reached, written the way you would say it out loud. "Check out as a guest." No script, no SDK.

Test Atlas

The living, editable map of your app that every run merges into, and that every later run starts from.

Deterministic replay

A passing run promoted to a flow that repeats exactly the same way, with no model call in it at all.

Defect and friction

Two different tiers. A defect is the app getting something wrong. Friction is the app costing someone effort while being technically correct.

The fair question

“We already have a suite. Why would we run this beside it?”

Because a suite asserts what somebody already thought to assert. The runs above were given one sentence, "buy a pack", on a phone and in a browser. The four defects that came back are a broken discount, a mislabelled buy button, a slow control and a dead page, and all four are on one surface only. None of them are what the person writing that test would have been checking.

How the run holds up

Goals, not selectors

You describe the outcome. Nothing in the run is pinned to a class name or a resource id, so a redesign does not break the suite the morning after it ships.

The same goal on both surfaces

A device run and a browser run of one sentence, reported in one vocabulary. A defect that only exists on one of them is the point, not an inconvenience.

Deterministic replay

Promote a run that passed and it replays exactly, with no model call at all. Classic automation reliability, none of the upkeep.

Test Atlas

Every run merges into one living map of your app. Edit it, run against it, and let each future run start from what the last one learned.

Point it at your app and see what it comes back with.

Same credits, same agents, whichever way you came in.