← Models & interaction guides
GUIDES / JevBench

Reading JevBench: Public 231, Hard 111 and SPX-CD results

Compare SPX-CD-Pro, Flash and Jev on Public 231 and Hard 111, with effort settings and ECE, without confusing public accuracy with the full leaderboard.

Surd AI · Research & engineeringEnglish

A JevBench search can return accuracy, Intelligence, Calibration and composite scores. These answer different questions. Correct answers on Public 231 are not the full JevBench composite score.

Who maintains JevBench?

JevBench is an independent typed-decision evaluation, not an official TypeSafe AI benchmark. Jev is one evaluated system. Its public splits contain Easy 48, Original 72 and Hard 111, totaling 231. Hard 111 is already inside Public 231. Project and datasets · Third-party attribution

Public 231 and Hard 111 results

These records were checked against the Surd AI release page on October 8, 2026. SPX-CD entries are self-evaluations; Jev 1.13.0 is the published reference used there for the same public group. This table does not claim certification or ranking on the complete third-party board. Release results

Model and setting Public 231 Hard 111 Hard ECE ↓
SPX-CD-Pro · effort=1 207/231 · 89.61% 88/111 · 79.28% 4.73%
SPX-CD-Pro · effort=2 208/231 · 90.04% 89/111 · 80.18% 9.80%
SPX-CD-Flash · effort=1 204/231 · 88.31% 86/111 · 77.48% 5.18%
SPX-CD-Flash · effort=2 206/231 · 89.18% 86/111 · 77.48% 7.06%
Jev 1.13.0 · published reference 200/231 · 86.58% 81/111 · 72.97% Not provided

Pro answers one more Hard item correctly at effort=2. Flash has the same Hard correct count at both settings. Both have lower Hard ECE at effort=1. These records therefore do not show a simultaneous improvement in accuracy and calibration with additional evaluations. Missing ECE is not zero error.

Why can the composite be lower than public accuracy?

The v1.4 method also considers sealed items, probability calibration, speed and cost, combining axes through rules including a harmonic mean. A high public accuracy need not produce a high composite. v1.4 method

Our suggested reading order is version, coverage, raw measurements, then aggregation. A composite in the sixties does not imply that the model answers only sixty percent of the public questions correctly.

Connect probability metrics to your decision threshold

ECE summarizes the discrepancy between confidence buckets and observed correctness. Its interpretation depends on sample size, question distribution and binning. If an application acts only above a particular probability threshold, report coverage and error rate at that threshold too.

A support-routing mistake and an irreversible operation have different costs. Use benchmark evidence to select candidates, then choose the operating threshold on independently labeled application examples.

Frequently asked questions

Can I add Public 231 and Hard 111?

No. That would count the hard subset twice.

Is effort=2 always better?

This table does not establish that. Compare correctness, probability quality and serving cost while preserving the exact test setting.

Is Decision Index the same evaluation?

No. Read the Decision Index guide and the CMDB guide, then test your own tasks in the Console.

SIMPLEX CD / API

Try a decision with your own input.

Create an account, explore decision models in the Console, or get started with the API documentation.