jevbooks

← All projects

jev_playground

JYeswak/jev_playground

Measure Jev's real capabilities before building on it.

Evaluation & Benchmarking79%Runner-up: Verification & GuardrailsQuestion health measured

What it is

A playground that tests Jev against real data, documenting graded findings, ruled-out candidates, and recipes with stop-conditions. For developers evaluating Jev's cost-benefit before integration.

How it uses Jev

The README does not describe specific Jev decisions, state inputs, or how results are used in code. It reports that Jev was tested on surfaces like tool-call harm, phishing, routing, and compaction, but no implementation details are given.

Technique worth stealing

Use held-out data and base rates to avoid overestimating deployability.

Try it

Run `node work/omp-harm-rule/verify-claim.mjs` or `bash foundation/gates.sh`.

View on GitHub

judged by Jevjev-1.13.0

Evidence

Each line is one question put to Jev about the README. ≥ 0.60 reads as yes, ≤ 0.40 as no; in between Jev is not making a call.

  • Jev-centricno0.29
  • Shows a System One patternno0.28
  • Handles uncertaintyno0.10
  • Measuredyes0.94
  • Runnableno0.09
  • Worth recommendingno0.19
  • Model replicano0.15
  • Problem scopescore on a 0–2 scale1.46
  • About Jevyes0.95

Signals by Jev jev-1.13.0, card written by DeepSeek V4.1 Flash from the README on 20 Sept 2026.