jevbooks

← All projects

jev-measured

WallerChen/jev-measured

Measured cost, latency, and raw output from the live Jev API across eight use cases.

What it is

A reproducible benchmark measuring cost, latency, and raw output from the live Jev API (TypeSafe AI's System One decision model) across eight realistic use cases, for developers evaluating Jev.

How it uses Jev

Jev makes decisions such as support-ticket triage, phishing detection, lead scoring, content moderation, RAG reranking, LLM model routing, agent tool selection, and agent output guardrails. Each call sends a state plus typed questions; the API returns calibrated probabilities used in code for routing, scoring, or gating.

Primitives:choicescorenoul

Technique worth stealing

Measure the network floor before quoting latency, and ask all questions in one call since they share one state and run in parallel.

Try it

Bring your own API key and re-run the benchmark scripts in bench/ to reproduce the measurements.

View on GitHub

judged by Jevjev-1.13.0

Evidence

Each line is one question put to Jev about the README. ≥ 0.60 reads as yes, ≤ 0.40 as no; in between Jev is not making a call.

  • Jev-centricyes0.84
  • Shows a System One patternno0.15
  • Handles uncertaintyno0.05
  • Measuredyes0.69
  • Runnableno0.09
  • Worth recommendingno0.28
  • Model replicano0.03
  • Problem scopescore on a 0–2 scale1.02
  • About Jevyes0.98

Signals by Jev jev-1.13.0, card written by DeepSeek V4.1 Flash from the README on 20 Sept 2026.