jevbooks

← All projects

jev-decision-benchmarks

baibizhe/jev-decision-benchmarks

Independent JEV benchmarks on MetaTool, When2Call, and BFCL V4.

Evaluation & Benchmarking84%Runner-up: Agent Decisions

What it is

Evaluation reports of JEV on three agent decision benchmarks, comparing against Claude, Qwen, GPT and other models. For readers assessing JEV tool selection and abstention behavior.

How it uses Jev

JEV selects the right tool or tool combination in MetaTool, chooses next-step actions in When2Call, and decides tool call vs abstention in BFCL V4. It uses a Choice adapter on MetaTool and When2Call, and a binary adapter on BFCL, with results compared to published baselines.

Primitives:choice

Technique worth stealing

Reproducible bilingual benchmark tables with hash-checked source data and offline rebuild scripts.

Try it

python3 scripts/build_tables.py --check; python3 scripts/build_tables.py

View on GitHub

judged by Jevjev-1.13.0

Evidence

Each line is one question put to Jev about the README. ≥ 0.60 reads as yes, ≤ 0.40 as no; in between Jev is not making a call.

  • Jev-centricyes0.86
  • Shows a System One patternno0.12
  • Handles uncertaintyno0.08
  • Measuredyes0.96
  • Runnableno0.04
  • Worth recommendingno0.21
  • Model replicano0.07
  • Problem scopescore on a 0–2 scale1.31
  • About Jevyes0.97

Signals by Jev jev-1.13.0, card written by DeepSeek V4.1 Flash from the README on 20 Sept 2026.