jevbooks

← All projects

jev-ood-calibration

scienthoon/jev-ood-calibration

Independent calibration test of Jev on unseen rule-based tasks.

Calibration & Research80%Runner-up: Evaluation & Benchmarking

What it is

An independent calibration study of TypeSafe's Jev, measuring whether its probabilities stay honest on inputs it cannot have seen: 900 rule-generated support tickets plus three public benchmarks, with every raw response published.

How it uses Jev

Jev answers Choice, Score and Boolean questions on synthetic support tickets and public-benchmark items via Vercel AI Gateway. Returned probabilities are scored for ECE against a simulated noise floor, refit with a single temperature, and analysed for the per-type sign of miscalibration.

Primitives:choicescorenoul

Technique worth stealing

Publish every raw response before scoring, and compare ECE against a simulated noise floor rather than zero.

View on GitHub

judged by Jevjev-1.13.0

Evidence

Each line is one question put to Jev about the README. ≥ 0.60 reads as yes, ≤ 0.40 as no; in between Jev is not making a call.

  • Jev-centricyes0.76
  • Shows a System One patternno0.30
  • Handles uncertaintyyes0.88
  • Measuredyes0.79
  • Runnableno0.06
  • Worth recommendingno0.38
  • Model replicano0.03
  • Problem scopescore on a 0–2 scale1.51
  • About Jevyes0.99

Signals by Jev jev-1.13.0, card written by DeepSeek V4.1 Flash from the README on 20 Sept 2026.