What it is
An independent, reproducible audit of TypeSafe AI's Jev model on Spanish, measuring accuracy, calibration and token cost across four human-labelled datasets, plus a CLI to run the same comparison on your own labelled data.
How it uses Jev
Jev answers typed questions (Choice, Noul) from a state. The audit compares three arms: English state with English instructions (A), Spanish state with English instructions (B), and Spanish state with Spanish instructions (C). Results are used to measure accuracy, calibration, and token cost.
Primitives:choicenoul
Technique worth stealing
Paired design: same items across arms, isolating the effect of state language and instruction language.
Try it
uv tool install jev-acento; acento run --provider typesafe
Evidence
Each line is one question put to Jev about the README. ≥ 0.60 reads as yes, ≤ 0.40 as no; in between Jev is not making a call.
- Jev-centricunclear0.51
- Shows a System One patternno0.09
- Handles uncertaintyno0.04
- Measuredyes0.97
- Runnableunclear0.42
- Worth recommendingno0.37
- Model replicano0.03
- Problem scopescore on a 0–2 scale1.08
- About Jevyes0.99
Signals by Jev jev-1.13.0, card written by DeepSeek V4.1 Flash from the README on 21 Sept 2026.