JevBench tests whether typed probabilistic decision models answer related questions coherently: a question and its negation, two options and their merger, a menu and its sub-menu. It measures consistency, not accuracy or speed, and needs no gold labels.
Every test transforms only the question side of a request (never the input it is about) in a way whose effect on the answer is fixed by probability or choice theory, and measures in probability units how far the model's answers miss that law. A test passes within 0.05. The 50 laws fall into five dimensions, by the operation that links the two answers compared.
The same events described differently (a paraphrase, a translation, reordered or renamed options) get the same answer.
A question's answer does not depend on the other questions asked with it, their order or their number.
Answers add up: a question and its negation sum to one; merging options merges their probabilities.
A conjunction is never more probable than its parts; a stricter threshold is never more probable.
Adding an irrelevant option or removing one leaves the other options' relative probabilities unchanged.
JevBench-mini (1,200 tests on 240 cases in 12 domains). Overall: the mean of the five dimensions, with its 95% interval. Range: the ranks the intervals allow; models whose ranges overlap are not separated. Accuracy is reported beside coherence: the model's unchanged answers against the gold answers of the cases. Click a row for its scores on the eleven groups, its profile and its full report.
Radar charts of the scores, on the five dimensions or the eleven groups. Pick up to six models to compare.
A model that ignores its input and answers every question uniformly keeps every law that does not depend on what the answers mean. It scores on coherence, above every model we tested, at chance accuracy (). Accuracy alone does not tell the decoders apart (), coherence does (). Read the two together.
Reference models (crosses): uniform answers uniformly; random gives unrelated answers to different questions; Luce is a toy word-overlap scorer, also with three biases.
From the technical report; every number on this page is computed from the published reports.
Every model was installed from its developers' release at a pinned revision and served with their own server, except where noted; the exact commands are in bench/models. None varied across three repetitions of the same request.
Any server implementing the Jev interface (POST /v1/systemone) or any Python
function (state, questions) → answers can be tested. The answer format is in
Connecting your model.
pip install jevbench # check that your server works with JevBench jevbench check --url http://127.0.0.1:8000/v1/systemone \ --model my-model # evaluate it on jevbench-mini jevbench eval --url http://127.0.0.1:8000/v1/systemone \ --model my-model --out my-model.json
import jevbench as jb
model = jb.systemone("http://127.0.0.1:8000/v1/systemone",
model="my-model")
jb.check(model) # one request
report = jb.evaluate(model) # jevbench-mini
print(report.table("group"))
report.save("my-model.json")
@misc{feng2026jevbench,
title = {JevBench: Metamorphic Coherence Testing for Typed Probabilistic Decision Models},
author = {Feng, Chen},
year = {2026},
howpublished = {\url{https://jevbench.github.io}},
note = {Technical report. ML Lab, Queen's University Belfast}
}