What it is
An independent calibration study of TypeSafe's Jev, measuring whether its probabilities stay honest on inputs it cannot have seen: 900 rule-generated support tickets plus three public benchmarks, with every raw response published.
How it uses Jev
Jev answers Choice, Score and Boolean questions on synthetic support tickets and public-benchmark items via Vercel AI Gateway. Returned probabilities are scored for ECE against a simulated noise floor, refit with a single temperature, and analysed for the per-type sign of miscalibration.
Primitives:choicescorenoul
Technique worth stealing
Publish every raw response before scoring, and compare ECE against a simulated noise floor rather than zero.
Evidence
Each line is one question put to Jev about the README. ≥ 0.60 reads as yes, ≤ 0.40 as no; in between Jev is not making a call.
- Jev-centricyes0.76
- Shows a System One patternno0.30
- Handles uncertaintyyes0.88
- Measuredyes0.79
- Runnableno0.06
- Worth recommendingno0.38
- Model replicano0.03
- Problem scopescore on a 0–2 scale1.51
- About Jevyes0.99
Signals by Jev jev-1.13.0, card written by DeepSeek V4.1 Flash from the README on 20 Sept 2026.