Smart Eye benchmark · last measured campaign

Proof that agents can see, act, and recover.

Last measured release campaign: v2.12.0 (2026-07-12). It passed 10/10 rounds across matched MCP and CLI routes, Killer Path, 5000+ node large-app stress, and five distinct local real-app profiles. Current product release is v2.14.0 (distribution / MCP / skill packaging); regenerate a campaign before publishing new speed or token claims.

34/34 quality gate passed in each real-app round
2026-07-12 local run · v2.12.0 campaign
5/5real-app targets passed
2.225sfirst useful observation avg
2.902sfirst action evidence avg
1,564useful observation tokens avg
doctorchecks setup and gives the next command
openhands back a live target prefix
perceivereturns layout, refs, console, and next steps
actclicks, fills, reloads, and records dispatch
evidencereports DOM, console, network, and recovery deltas
reportpackages the long-session handoff

What the gate proves

JSON action evidence completeness
Failed action diagnosis and recovery
No-change action investigation
Executable next steps in handoffs
CSS source tracing handoff
Frame, modal, and HMR probes
Session report artifacts
Stale-ref recovery

Time to useful evidence

First observation avg
2.225s
First action evidence
2.902s
Golden path avg
5.353s
Total run avg
10.264s

The five real-app rounds are local test fixtures with profile-specific overlay, stale-ref, iframe, shadow DOM, SPA route, slow-network, auth-wall, large-table, hidden-template, and canvas probes. They are not external production-app evidence.

npm run benchmark:campaign -- --rounds 10 --types mcp,cli,killer,large-app,real-app,real-app,real-app,real-app,real-app,cli --real-app-targets dashboard,docs-app,auth-flow,data-table,canvas-heavy --settle-ms 0 --json

Publish competitor deltas only after rerunning with measured Playwright and generic-CDP baselines. The README promotion checklist blocks claims when the benchmark gate fails.