What it is
A reproducible billing study of TypeSafe Jev, measuring token costs and answer stability for batched versus separate questions across state sizes and question counts, using synthetic support tickets.
How it uses Jev
Jev answers Noul, Choice, and Score questions on synthetic support ticket states. The study compares billing and answers when N questions are asked in one call versus N separate calls, using usage.input_tokens and response values.
Primitives:choicescorenoul
Technique worth stealing
Derive fixed per-request overhead F independently at each N and question subset to verify exact linear billing.
Try it
export TYPESAFE_API_KEY=...; python bench.py run; python bench.py report
Evidence
Each line is one question put to Jev about the README. ≥ 0.60 reads as yes, ≤ 0.40 as no; in between Jev is not making a call.
- Jev-centricyes0.62
- Shows a System One patternunclear0.59
- Handles uncertaintyno0.03
- Measuredyes0.63
- Runnableno0.07
- Worth recommendingno0.34
- Model replicano0.02
- Problem scopescore on a 0–2 scale1.01
- About Jevyes0.99
Signals by Jev jev-1.13.0, card written by DeepSeek V4.1 Flash from the README on 23 Sept 2026.