jevbooks

← All projects

jev-research-eval

jgridifier/jev-research-eval

Reproducible eval harness for Jev Ultrafast research-browser tasks.

Evaluation & Benchmarking100%Runner-up: Classification & Routing

What it is

A reproducible evaluation harness for Jared's Jev Ultrafast research-browser session, with baseline cases, stress suites, QC grades, field note, and notebooks. For researchers evaluating browser agents.

How it uses Jev

Jev makes decisions in the research-browser agent; the harness drives an upstream checkout of browser-use/jev-ultrafast to run cases, then applies QC grades and generates reports. Jev's outputs are used in the agent's history and traces.

Technique worth stealing

CoS-locked QC grades and per-step Trace reports for reproducible evaluation.

Try it

Run python scripts/generate_report_v4.py --input fixtures/qc_rescored.json --output /tmp/jev_note_regen.html

View on GitHub

judged by Jevjev-1.13.0

Evidence

Each line is one question put to Jev about the README. ≥ 0.60 reads as yes, ≤ 0.40 as no; in between Jev is not making a call.

  • Jev-centricunclear0.44
  • Shows a System One patternunclear0.40
  • Handles uncertaintyno0.05
  • Measuredno0.19
  • Runnableno0.11
  • Worth recommendingunclear0.43
  • Model replicano0.04
  • Problem scopescore on a 0–2 scale1.00
  • About Jevyes0.97

Signals by Jev jev-1.13.0, card written by DeepSeek V4.1 Flash from the README on 20 Sept 2026.