jevbooks

← All projects

system-one-bench

reachjalil/system-one-bench

Benchmark catalog for Jev evidence, findings, and reproducible decision benchmarks.

Evaluation & Benchmarking95%Runner-up: Agent Decisions

What it is

A public catalog of experiments measuring what Jev can help an agent do, for individual agent users and builders. It records each experiment's measurement, baseline, and conclusion limits.

How it uses Jev

Jev makes decisions like choosing a palette, selecting items, ranking context, checking claims, picking tools, and triaging logs. The state is the task context; Jev returns calibrated probabilities used to select or rank options in code.

Primitives:choicescorenoul

Technique worth stealing

Define typed questions over shared state to get calibrated probabilities for agent decisions.

Try it

See the practical guides and evidence records for reproducible decision benchmarks.

View on GitHub

judged by Jevjev-1.13.0

Evidence

Each line is one question put to Jev about the README. ≥ 0.60 reads as yes, ≤ 0.40 as no; in between Jev is not making a call.

  • Jev-centricyes0.62
  • Shows a System One patternno0.11
  • Handles uncertaintyno0.04
  • Measuredno0.08
  • Runnableno0.08
  • Worth recommendingno0.31
  • Model replicano0.11
  • Problem scopescore on a 0–2 scale1.05
  • About Jevyes0.95

Signals by Jev jev-1.13.0, card written by DeepSeek V4.1 Flash from the README on 20 Sept 2026.