jevbooks

← All projects

jev-acento

marcosmartinez/jev-acento · Homepage

Audit of Jev on Spanish: accuracy, calibration, token cost.

Evaluation & Benchmarking81%Runner-up: Calibration & ResearchFunnel

What it is

An independent, reproducible audit of TypeSafe AI's Jev model on Spanish, measuring accuracy, calibration and token cost across four human-labelled datasets, plus a CLI to run the same comparison on your own labelled data.

How it uses Jev

Jev answers typed questions (Choice, Noul) from a state. The audit compares three arms: English state with English instructions (A), Spanish state with English instructions (B), and Spanish state with Spanish instructions (C). Results are used to measure accuracy, calibration, and token cost.

Primitives:choicenoul

Technique worth stealing

Paired design: same items across arms, isolating the effect of state language and instruction language.

Try it

uv tool install jev-acento; acento run --provider typesafe

View on GitHub

judged by Jevjev-1.13.0

Evidence

Each line is one question put to Jev about the README. ≥ 0.60 reads as yes, ≤ 0.40 as no; in between Jev is not making a call.

  • Jev-centricunclear0.51
  • Shows a System One patternno0.09
  • Handles uncertaintyno0.04
  • Measuredyes0.97
  • Runnableunclear0.42
  • Worth recommendingno0.37
  • Model replicano0.03
  • Problem scopescore on a 0–2 scale1.08
  • About Jevyes0.99

Signals by Jev jev-1.13.0, card written by DeepSeek V4.1 Flash from the README on 21 Sept 2026.