JevBench
Benchmark · Technical report · Python package

Do a decision model's answers fit together?

JevBench tests whether typed probabilistic decision models answer related questions coherently: a question and its negation, two options and their merger, a menu and its sub-menu. It measures consistency, not accuracy or speed, and needs no gold labels.

01What it measures

Every test transforms only the question side of a request (never the input it is about) in a way whose effect on the answer is fixed by probability or choice theory, and measures in probability units how far the model's answers miss that law. A test passes within 0.05. The 50 laws fall into five dimensions, by the operation that links the two answers compared.

REP

Representation

The same events described differently (a paraphrase, a translation, reordered or renamed options) get the same answer.

BAT

Batch independence

A question's answer does not depend on the other questions asked with it, their order or their number.

MEA

Probability measure

Answers add up: a question and its negation sum to one; merging options merges their probabilities.

LOG

Logical order

A conjunction is never more probable than its parts; a stricter threshold is never more probable.

CHO

Choice-set conditioning

Adding an irrelevant option or removing one leaves the other options' relative probabilities unchanged.

02Leaderboard

JevBench-mini (1,200 tests on 240 cases in 12 domains). Overall: the mean of the five dimensions, with its 95% interval. Range: the ranks the intervals allow; models whose ranges overlap are not separated. Accuracy is reported beside coherence: the model's unchanged answers against the gold answers of the cases. Click a row for its scores on the eleven groups, its profile and its full report.

03Profiles

Radar charts of the scores, on the five dimensions or the eleven groups. Pick up to six models to compare.

By family

04Coherence is not accuracy

A model that ignores its input and answers every question uniformly keeps every law that does not depend on what the answers mean. It scores on coherence, above every model we tested, at chance accuracy (). Accuracy alone does not tell the decoders apart (), coherence does (). Read the two together.

Reference models (crosses): uniform answers uniformly; random gives unrelated answers to different questions; Luce is a toy word-overlap scorer, also with three biases.

05Findings

From the technical report; every number on this page is computed from the published reports.

    06Models

    Every model was installed from its developers' release at a pinned revision and served with their own server, except where noted; the exact commands are in bench/models. None varied across three repetitions of the same request.

    07Get started

    Any server implementing the Jev interface (POST /v1/systemone) or any Python function (state, questions) → answers can be tested. The answer format is in Connecting your model.

    pip install jevbench
    
    # check that your server works with JevBench
    jevbench check --url http://127.0.0.1:8000/v1/systemone \
                   --model my-model
    
    # evaluate it on jevbench-mini
    jevbench eval  --url http://127.0.0.1:8000/v1/systemone \
                   --model my-model --out my-model.json
    import jevbench as jb
    
    model = jb.systemone("http://127.0.0.1:8000/v1/systemone",
                         model="my-model")
    jb.check(model)              # one request
    report = jb.evaluate(model)  # jevbench-mini
    print(report.table("group"))
    report.save("my-model.json")

    08Cite

    @misc{feng2026jevbench,
      title        = {JevBench: Metamorphic Coherence Testing for Typed Probabilistic Decision Models},
      author       = {Feng, Chen},
      year         = {2026},
      howpublished = {\url{https://jevbench.github.io}},
      note         = {Technical report. ML Lab, Queen's University Belfast}
    }