Completing roughly 150,000 requests does not necessarily produce the current leaderboard’s Full score. Decision Index editions have changed the suite and scoring scope. Every reported result needs an edition and a clear distinction between public index and Full score.
What is Decision Index?
Decision Index evaluates typed decision engines and provides reproduction tools alongside a Hugging Face board. The documentation checked on October 8, 2026 identifies 0.3 as current and retains reproducible earlier editions. Reproduction project · Hugging Face board
What does the roughly 150,000-request count mean?
The 0.2.1 documentation lists 119,898 scoreable requests plus 30,419 in its added portion: 150,317 in total. This counts requests, not independent source datasets. Edition details
A useful evaluation record preserves the suite manifest, answered and failed counts, and aggregation method. That lets you distinguish model changes from incomplete runs. A large request count alone does not tell you whether a final index is an ordinary accuracy percentage.
Versioned SPX-CD and Jev records
The Surd AI release retains these Decision Index 0.2.1 results. SPX-CD rows are self-evaluations at effort=1; Jev is the published reference cited there. This is not a Decision Index 0.3 Full-score ranking. Model release
| Model | 0.2.1 index ↑ | Setting |
|---|---|---|
| SPX-CD-Pro | 58.30 | effort=1 |
| Jev 1.13.0 | 57.91 | Published reference |
| SPX-CD-Flash | 54.65 | effort=1 |
Pro is 0.39 points above Jev in these records, while Flash is 3.26 below. A more favorable ordering on another evaluation cannot establish superiority on every decision task. No SPX-CD effort=2 value for this index is provided in the release table.
Public index versus the 0.3 Full score
The current documentation assigns 20% to the public component, 50% to private tests of related skills and 30% to private decisions in new domains, with scale alignment for the private components. The public kit alone does not compute the Full score. 0.3 scoring and submission
Label a public-only run as such. Do not place it in the Full-score column or relabel a 0.2.1 number as 0.3.
Why coverage and chance baselines matter
Scoring only answered questions can reward a system for skipping difficult ones. A coverage adjustment preserves the effect of unfinished work. A chance baseline recognizes that guessing between two options differs from guessing among many. These principles explain why a normalized index needs its own interpretation.
For your own application, freeze error handling before running the evaluation. Decide how timeouts, unsupported types and malformed outputs affect the denominator, instead of selecting a favorable rule after seeing results.
Frequently asked questions
Does 58.30 mean 58.30% accuracy?
No. It is a versioned aggregate index whose meaning depends on task metrics and the aggregation rules.
Why can JevBench, CMDB and Decision Index rank models differently?
They differ in tasks, modalities, coverage, configuration and scoring. Use disagreement to guide further testing rather than averaging unrelated numbers into a supposed overall accuracy.
Where should I read next?
See the JevBench guide, the CMDB guide, or the guide directory for image and audio evaluation topics.