What it is
A reproducible evaluation harness for Jared's Jev Ultrafast research-browser session, with baseline cases, stress suites, QC grades, field note, and notebooks. For researchers evaluating browser agents.
How it uses Jev
Jev makes decisions in the research-browser agent; the harness drives an upstream checkout of browser-use/jev-ultrafast to run cases, then applies QC grades and generates reports. Jev's outputs are used in the agent's history and traces.
Technique worth stealing
CoS-locked QC grades and per-step Trace reports for reproducible evaluation.
Try it
Run python scripts/generate_report_v4.py --input fixtures/qc_rescored.json --output /tmp/jev_note_regen.html
Evidence
Each line is one question put to Jev about the README. ≥ 0.60 reads as yes, ≤ 0.40 as no; in between Jev is not making a call.
- Jev-centricunclear0.44
- Shows a System One patternunclear0.40
- Handles uncertaintyno0.05
- Measuredno0.19
- Runnableno0.11
- Worth recommendingunclear0.43
- Model replicano0.04
- Problem scopescore on a 0–2 scale1.00
- About Jevyes0.97
Signals by Jev jev-1.13.0, card written by DeepSeek V4.1 Flash from the README on 20 Sept 2026.