jevbooks

← All projects

jev-synthetic-survey

jjd-lab/jev-synthetic-survey · Homepage

Jev vs GPT-4.1 as synthetic survey respondents on Twin-2K-500.

Evaluation & Benchmarking46%Runner-up: Calibration & ResearchCheck-up

What it is

Runs Jev and GPT-4.1 as 300 synthetic survey respondents on Twin-2K-500, comparing probability elicitation methods for survey research.

How it uses Jev

Jev answers yes/no questions as Noul (single probability) and multi-option as Choice. Its probability vector is compared to GPT-4.1's verbalized probabilities on distribution gap, calibration, Brier, and accuracy.

Primitives:choicenoul

Technique worth stealing

Asking yes/no items as Noul instead of Choice moves results more than model choice.

View on GitHub

judged by Jevjev-1.13.0

Evidence

Each line is one question put to Jev about the README. ≥ 0.60 reads as yes, ≤ 0.40 as no; in between Jev is not making a call.

  • Jev-centricyes0.82
  • Shows a System One patternno0.33
  • Handles uncertaintyno0.04
  • Measuredyes0.75
  • Runnableno0.04
  • Worth recommendingunclear0.43
  • Model replicano0.04
  • Problem scopescore on a 0–2 scale0.83
  • About Jevyes0.98

Signals by Jev jev-1.13.0, card written by DeepSeek V4.1 Flash from the README on 21 Sept 2026.