Text accuracy, image accuracy, valid responses, ECE and I/C are different measurements. Read CMDB-1500 by modality and denominator first, then examine probabilities and your application’s error costs.
What is CMDB-1500?
CMDB stands for Comprehensive Multimodal Decision Benchmark. Surd AI’s CMDB-1500 contains 1,200 text questions and 300 image questions for evaluating choices and probabilities. Surd AI maintains the board and evaluates its own SPX-CD models, so retain that provenance. CMDB introduction and leaderboard
| View | Fixed denominator | Question it helps answer |
|---|---|---|
| Text | 1,200 | Which system answers more of the same text tasks correctly? |
| Image | 300 | How does it perform on this visual task set? |
| All tasks | 1,500 | What is the result for a system evaluated on both modalities? |
| Probability and I/C | See the associated method and coverage | How are confidence quality and normalized indices measured? |
SPX-CD and Jev text results
The following text records were checked on October 8, 2026. SPX-CD uses effort=2 in this evaluation. These are results on a fixed suite, not promised accuracy on arbitrary applications. Full results and method
| Model | Text accuracy | Correct / 1,200 |
|---|---|---|
| SPX-CD-Pro | 79.33% | 952 |
| SPX-CD-Flash | 77.58% | 931 |
| Jev | 72.17% | 866 |
Pro answers 86 more text items correctly than Jev; Flash answers 65 more. Jev’s cited record covers the 1,200 text questions. It does not establish an image result and should not be inserted into a mixed-modality ranking as though it did.
Why valid coverage matters
Imagine a model returning 1,000 usable responses to 1,200 questions, with 900 correct. Accuracy among usable answers is 90%; accuracy over the fixed task denominator is 75%. These describe different outcomes. The example illustrates denominators, not an actual model result.
Also inspect the scope of the coverage column. A multimodal model may have 1,500 valid responses across text and images while its text accuracy still uses 1,200 questions.
Accuracy, ECE and I/C are not interchangeable
The CMDB page describes ECE using ten probability bins over valid choice and yes/no questions with native probabilities, and notes that modality coverage can differ. Calibration methodology
Lower ECE does not mean every answer is correct. High accuracy does not establish that confident predictions are reliable. Interpret I/C through this board’s methodology rather than equating similarly named axes across evaluations.
For an automated workflow, add a confusion matrix showing allowed cases incorrectly blocked and blocked cases incorrectly allowed. Set thresholds according to the consequences of those errors rather than a composite score alone.
Frequently asked questions
Can I convert this result into Image JevBench points?
No direct conversion follows from these tables. Match dataset, version and scoring method. Read the image benchmark guide.
Does this accuracy table establish which production API is fastest?
No. Measure end-to-end time using matched input length, image resolution, question count and concurrency.
How should I run my own evaluation?
Use the leaderboard to select candidates, then freeze labels and settings with the decision API evaluation guide. Registered users can explore application examples in the Console.