Smart Eye benchmark · latest measured release

Proof that agents can see, act, and recover.

Latest measured release: v2.15.0 (2026-08-12). It passed 10/10 rounds across matched MCP and CLI routes, Killer Path, 5,200-node large-app stress, and five distinct safe local real-app profiles.

34/34 quality gate passed in each real-app round
2026-08-12 local run · v2.15.0 campaign
5/5real-app targets passed
2.213sfirst useful observation avg
2.941sfirst action evidence avg
1,732useful observation tokens avg
doctorchecks setup and gives the next command
openhands back a live target prefix
perceivereturns layout, refs, console, and next steps
actclicks, fills, reloads, and records dispatch
evidencereports DOM, console, network, and recovery deltas
reportpackages the long-session handoff

What the gate proves

JSON action evidence completeness
Failed action diagnosis and recovery
No-change action investigation
Executable next steps in handoffs
CSS source tracing handoff
Frame, modal, and HMR probes
Session report artifacts
Stale-ref recovery

Time to useful evidence

First observation avg
2.213s
First action evidence
2.941s
Golden path avg
5.538s
Total run avg
10.686s

The five real-app rounds are local test fixtures with profile-specific overlay, stale-ref, iframe, shadow DOM, SPA route, slow-network, auth-wall, large-table, hidden-template, and canvas probes. They are not external production-app evidence.

npm run benchmark:campaign -- --rounds 10 --types mcp,cli,killer,large-app,real-app,real-app,real-app,real-app,real-app,cli --real-app-targets dashboard,docs-app,auth-flow,data-table,canvas-heavy --settle-ms 0 --json

This page is the v2.15.0 mixed local campaign still. The README first screen now leads Cool → live-session one-step → win score (chrome-cdp-ex 10 PASS, Browser Use 8 PASS, Playwright 9 PASS) and steps / token / time charts. The ten-row engineer grid is in the PK board doc, not a README lab table.