What it is
A public field guide to Jev 1.13.0's framing sensitivity and failures, based on 11,621 text-study requests across synthetic tasks. For AI researchers and practitioners choosing how to interact with Jev.
How it uses Jev
Jev makes decisions by answering explicit typed questions in each task, such as choosing a Snake move, a driving action, or a numerical answer. The study records each response and its token usage, then evaluates accuracy against known correct answers offline.
Primitives:choicescore
Technique worth stealing
Describing tasks in plain language with explicit priorities improved Snake performance over raw board histories.
Try it
Explore the City lab and Snake lab demos online, or read the reports and raw CSVs in the repository.
Evidence
Each line is one question put to Jev about the README. ≥ 0.60 reads as yes, ≤ 0.40 as no; in between Jev is not making a call.
- Jev-centricunclear0.43
- Shows a System One patternno0.18
- Handles uncertaintyno0.05
- Measuredyes0.66
- Runnableno0.04
- Worth recommendingunclear0.50
- Model replicano0.03
- Problem scopescore on a 0–2 scale1.02
- About Jevyes0.99
Signals by Jev jev-1.13.0, card written by DeepSeek V4.1 Flash from the README on 20 Sept 2026.