For product teams and UX designers

See your product through something that just used it for the first time.

Every screen the agent reached is read back three ways: is it usable, is it crafted, does it agree with the screens either side of it. It runs on the phone and in the browser, so the same journey can be judged where your customers actually are. Every finding points at the frame it came from.

MobileWeb
Mobile · Home100%
Web · Home
One goal, both surfaces5 stages · 8 findings · 7 on mobile
Mobile · Product2 raised at this stage
Stage 03 · Product

Opened one pack and looked for the way to buy it

MobileImage, thumbnails, title, price and description all arrive before any action does.
WebPrice, rating and description sit beside the image rather than under it.
On mobile: 7 of 8 findings
6 UI/UX findings1 friction findingon mobile
Every stage, both sidesMobile above, Web below
01 Home
02 Shop
03 Product
04 Checkout
05 Empty cart
What you get

Four things you get from one run

01

Three readings of one journey

Usability, visual craft, and consistency across the flow. Three different questions asked of the same evidence, so a finding is never only a matter of taste.

02

The phone and the browser

The same journey read on both, which is the only way to catch the screen that was designed once and adapted badly.

03

The frame it came from

Every finding carries the screen it was read from and the step it happened on. You can disagree with the agent while looking at exactly what it looked at.

04

What is done well, too

A report that only lists faults cannot be trusted about their weight, so the things the app gets right are named in the same voice.

From the run above
Screens read
5
Readings over each one
3
Read again at full resolutionthe worst ranked
2
The words on this page

Four terms do most of the work. None of them are ours to redefine in a meeting later.

The three readings

Usability asks whether a person could do it. Visual craft asks whether the screen was built with care. Flow consistency asks whether the screens agree with each other.

Contact sheet

The ordered set of screens a run actually reached, which is what gets read. Nothing is judged that the agent did not see.

Deep read

The second pass over the worst screens at full resolution, where the small things live.

Friction

Effort the app charged the person while working exactly as built. Searches that returned nothing, dead ends, repeated taps, replanning loops, recovery detours.

The fair question

“Is this not just an opinion generated by a model?”

It is an opinion with the frame attached. Each finding names the screen, the step, the surface and the reading that raised it, so the argument in the room is about the screen rather than about the tool.

How the run holds up

Three readings, one pass

Usability, visual craft and consistency across the flow. Three different questions asked of the same evidence, so a finding is never just an opinion about taste.

Both surfaces, one vocabulary

The phone run and the browser run are read the same way, so "worse on mobile" is a thing the report can actually say instead of a thing you argue about.

Evidence, not adjectives

Every finding carries the screenshot it was read from and the step it happened on. You can disagree with the agent while looking at the same frame it looked at.

Ranked, so you can stop reading

The screens are ordered by how much they cost the user, and only the worst get the expensive second read at full resolution.

Point it at your app and see what it comes back with.

Same credits, same agents, whichever way you came in.