Experiment 01 / A browser with a goal

Browser Sprint.

One goal. Two ways through the web. Watch the difference.

Same URL & goal · separate fresh sessions
Diffusion · fast default$0.04 in / $0.15 out per 1M tokensModel card ↗
Choose a challenge

Search + availability + format + budget + sorting. The cheapest book overall is a decoy.

01 / Typed decisions

Jev

Jev + Mercury for text entry
Jev model card ↗

Elapsed0.0s
Actions00
Pages00
AI model cost$0.00USD · AI calls only
practice.jev-on.test/bookshopREADY

Jev’s front-row seat

Same goal. Its own way there.

A fresh browser, real decisions. Watch every move right here.

READY FOR YOUR NEXT RUN
A fresh session for every run1120 × 780
Jev questions Inside the decision

Run a task to inspect the exact questions, choices, and page context sent to Jev. Each request appears here as it happens.

02 / Tool-calling agent

LLM + Playwright

Mercury 2.5 ↗
Playwright tools

Elapsed0.0s
Actions00
Pages00
AI model cost$0.00USD · AI calls only
practice.jev-on.test/bookshopREADY

LLM + Playwright’s front-row seat

Same goal. Its own way there.

A fresh browser, real decisions. Watch every move right here.

READY FOR YOUR NEXT RUN
A fresh session for every run1120 × 780
What are we comparing?

Jev uses the operation-and-target policy adapted from Jev Ultrafast. Mercury 2.5 supplies text when a field needs typing. LLM + Playwright lets the selected model choose each action and its arguments, then executes Playwright’s click, fill, select, and scroll tools.

Both receive the same visible page text, observed controls, goal, and recent action history. They start together in separate Chromium sessions with identical viewports and limits. Jev also uses Playwright to launch Chromium; this compares the decision and action paths. Screenshots are for you and are sent to neither model. Model speed depends on the task and provider. No artificial delays are added.

Each browser stops after 24 actions, 40 model calls, 10 pages, or 90 seconds. Each has a $0.05 AI spending threshold checked between calls; a call in progress can finish above it. AI Gateway supplies costs when available; token-based estimates and incomplete bills are labeled. Browser hosting is excluded. Timers include startup, actions, verification, and cleanup, and stop separately.

Public pages and ordinary HTML controls work best. Sign-ins, purchases, uploads, pop-up tabs, and embedded frames are outside this playground. Page text and your goal are sent through AI Gateway. Cloud runs are processed by Vercel Sandbox, including execution logs. The app does not offer saved run history. The practice bookshop uses real model decisions on a fictional site; preset results are independently checked.