RESEARCH / MULTIMODAL DECISIONS
CMDB-1500Multimodal Decision Model Benchmark
Comprehensive Multimodal Decision Benchmark 1500
1,200 text questions and 300 image questions, measuring decisions and probability estimates.
View exact text results
| Rank | Model / version | Accuracy · 0–100% | Correct / fixed total | Valid / evaluated scope |
|---|---|---|---|---|
| 01 | GPT 6.1 Sol · low | 80.00% | 960 / 1,200 | 1,500 / 1,500 |
| 02 | SPX-CD-Pro (2026-10-04) | 79.33% | 952 / 1,200 | 1,500 / 1,500 |
| 03 | Claude Sonnet 5.5 · low | 78.00% | 936 / 1,200 | 1,500 / 1,500 |
| 04 | StartLux-Decision-27B · BF16 | 77.92% | 935 / 1,200 | 1,500 / 1,500 |
| 05 | SPX-CD-Flash | 77.58% | 931 / 1,200 | 1,500 / 1,500 |
| 06 | GPT 6 Luna · low | 76.58% | 919 / 1,200 | 1,500 / 1,500 |
| 07 | GLM 5.3 Flash · low | 75.75% | 909 / 1,200 | 1,500 / 1,500 |
| 08 | StartLux-Decision-35B-A3B · BF16 | 75.67% | 908 / 1,200 | 1,500 / 1,500 |
| 09 | SPX-CD-Omni | 74.67% | 896 / 1,200 | 1,500 / 1,500 |
| 10 | Mercury Decide | 73.92% | 887 / 1,200 | 1,200 / 1,200 |
| 11 | pplx-decider 27B | 72.83% | 874 / 1,200 | 1,500 / 1,500 |
| 12 | Jev | 72.17% | 866 / 1,200 | 1,200 / 1,200 |
| 13 | Cloudflare Clef | 71.50% | 858 / 1,200 | 1,485 / 1,500 |
| 14 | JPT · 9B | 71.33% | 856 / 1,200 | 1,500 / 1,500 |
| 15 | Qwen 3.8 Flash | 71.08% | 853 / 1,200 | 1,500 / 1,500 |
| 16 | Kimi K2.6 | 71.00% | 852 / 1,200 | 1,500 / 1,500 |
| 17 | Mapika · Decider 4B v2.1 | 69.83% | 838 / 1,200 | 1,200 / 1,200 |
| 18 | Cloudflare Clef-Flash | 69.67% | 836 / 1,200 | 1,485 / 1,500 |
| 19 | Imajev-9B | 68.83% | 826 / 1,200 | 1,478 / 1,500 |
| 20 | Kev · 9B | 68.67% | 824 / 1,200 | 1,200 / 1,200 |
| 21 | DeepSeek V4.1 Flash | 68.58% | 823 / 1,200 | 1,500 / 1,500 |
| 22 | JPT · 4B | 67.67% | 812 / 1,200 | 1,500 / 1,500 |
| 23 | Decision Lux · 9B | 67.42% | 809 / 1,200 | 1,200 / 1,200 |
| 24 | Jet · v6.2 4B | 66.92% | 803 / 1,200 | 1,200 / 1,200 |
| 25 | Open-Jev · Zefan 9B | 66.67% | 800 / 1,200 | 1,199 / 1,200 |
| 26 | Nimble · v2 9B | 66.58% | 799 / 1,200 | 1,200 / 1,200 |
| 27 | Hopper G · 4B v1.3 | 65.92% | 791 / 1,200 | 1,200 / 1,200 |
| 28 | Malkuth · 4B | 65.58% | 787 / 1,200 | 1,200 / 1,200 |
| 29 | Hopper · 4B | 65.50% | 786 / 1,200 | 1,200 / 1,200 |
| 30 | Winnow · 12B Q8 | 65.33% | 784 / 1,200 | 1,410 / 1,500 |
| 31 | JevK5 | 64.83% | 778 / 1,200 | 1,200 / 1,200 |
| 32 | Lev · 4B | 64.58% | 775 / 1,200 | 1,200 / 1,200 |
| 33 | Mica · 4B BF16 GGUF | 64.25% | 771 / 1,200 | 1,200 / 1,200 |
| 34 | Kev · 4B | 64.17% | 770 / 1,200 | 1,200 / 1,200 |
| 35 | Jev-Omni · 12B | 63.58% | 763 / 1,200 | 1,493 / 1,500 |
| 36 | APUS OpenJev · 9B 3000 High | 61.42% | 737 / 1,200 | 1,085 / 1,200 |
| 37 | OpenDecider · Small | 61.25% | 735 / 1,200 | 1,200 / 1,200 |
| 38 | APUS OpenJev · 9B 5949 High | 61.08% | 733 / 1,200 | 1,085 / 1,200 |
| 39 | OpenDecider · Small TD | 60.67% | 728 / 1,200 | 1,200 / 1,200 |
| 40 | APUS OpenJev · 4B 5949 High | 60.08% | 721 / 1,200 | 1,085 / 1,200 |
| 41 | Ateve Jev | 59.75% | 717 / 1,200 | 1,087 / 1,200 |
| 42 | Bocha Jev | 58.92% | 707 / 1,200 | 1,087 / 1,200 |
| 43 | Winnow · E4B Q8 | 58.92% | 707 / 1,200 | 1,410 / 1,500 |
| 44 | Reflex · Stable 4B | 58.50% | 702 / 1,200 | 1,387 / 1,500 |
| 45 | Plumb · v5.2 4B | 58.08% | 697 / 1,200 | 1,085 / 1,200 |
| 46 | FLock THIS/THAT | 57.67% | 692 / 1,200 | 1,156 / 1,200 |
| 47 | Kev · 0.8B | 51.75% | 621 / 1,200 | 1,200 / 1,200 |
| 48 | Bosun v3.1 · 1.7B | 49.75% | 597 / 1,200 | 1,200 / 1,200 |
| 49 | OpenDecider · Nano | 46.75% | 561 / 1,200 | 1,143 / 1,200 |
| 50 | CalDec · Laya | 46.08% | 553 / 1,200 | 1,189 / 1,200 |
| 51 | Bosun v3.1 · 0.6B | 45.58% | 547 / 1,200 | 1,200 / 1,200 |
| 52 | Laya · Typed Decisions | 44.58% | 535 / 1,200 | 1,189 / 1,200 |
| 53 | CalDec · GLiNER | 43.17% | 518 / 1,200 | 1,200 / 1,200 |
| 54 | Laya · English | 41.00% | 492 / 1,200 | 1,189 / 1,200 |
| 55 | Laya · Multilingual | 40.25% | 483 / 1,200 | 1,192 / 1,200 |
| 56 | GLiNER2.5-Decide | 38.58% | 463 / 1,200 | 1,200 / 1,200 |
| 57 | Julia-1 | 35.75% | 429 / 1,200 | 1,079 / 1,200 |
| 58 | CLM v0.1 · 8B | 34.92% | 419 / 1,200 | 1,200 / 1,200 |
View exact image results
| Rank | Model / version | Accuracy · 0–100% | Correct / fixed total | Valid / evaluated scope |
|---|---|---|---|---|
| 01 | Claude Sonnet 5.5 · low | 92.67% | 278 / 300 | 1,500 / 1,500 |
| 02 | GPT 6.1 Sol · low | 92.00% | 276 / 300 | 1,500 / 1,500 |
| 03 | GLM 5.3 Flash · low | 90.67% | 272 / 300 | 1,500 / 1,500 |
| 04 | StartLux-Decision-35B-A3B · BF16 | 89.67% | 269 / 300 | 1,500 / 1,500 |
| 05 | GPT 6 Luna · low | 89.33% | 268 / 300 | 1,500 / 1,500 |
| 06 | SPX-CD-Pro (2026-10-04) | 88.67% | 266 / 300 | 1,500 / 1,500 |
| 07 | SPX-CD-Flash | 87.33% | 262 / 300 | 1,500 / 1,500 |
| 08 | StartLux-Decision-27B · BF16 | 87.00% | 261 / 300 | 1,500 / 1,500 |
| 09 | pplx-decider 27B | 86.33% | 259 / 300 | 1,500 / 1,500 |
| 10 | JPT · 9B | 85.67% | 257 / 300 | 1,500 / 1,500 |
| 11 | Imajev-9B | 83.67% | 251 / 300 | 1,478 / 1,500 |
| 12 | Qwen 3.8 Flash | 83.00% | 249 / 300 | 1,500 / 1,500 |
| 13 | SPX-CD-Omni | 82.67% | 248 / 300 | 1,500 / 1,500 |
| 14 | Kimi K2.6 | 81.33% | 244 / 300 | 1,500 / 1,500 |
| 15 | JPT · 4B | 81.00% | 243 / 300 | 1,500 / 1,500 |
| 16 | Reflex · Stable 4B | 80.33% | 241 / 300 | 1,387 / 1,500 |
| 17 | DeepSeek V4.1 Flash | 79.67% | 239 / 300 | 1,500 / 1,500 |
| 18 | Jev-Omni · 12B | 79.00% | 237 / 300 | 1,493 / 1,500 |
| 19 | Winnow · 12B Q8 | 77.33% | 232 / 300 | 1,410 / 1,500 |
| 20 | Winnow · E4B Q8 | 66.33% | 199 / 300 | 1,410 / 1,500 |
| 21 | Cloudflare Clef | 42.33% | 127 / 300 | 1,485 / 1,500 |
| 22 | Cloudflare Clef-Flash | 41.67% | 125 / 300 | 1,485 / 1,500 |
View exact overall results
| Rank | Model / version | Accuracy · 0–100% | Correct / fixed total | Valid / evaluated scope |
|---|---|---|---|---|
| 01 | GPT 6.1 Sol · low | 82.40% | 1,236 / 1,500 | 1,500 / 1,500 |
| 02 | SPX-CD-Pro (2026-10-04) | 81.20% | 1,218 / 1,500 | 1,500 / 1,500 |
| 03 | Claude Sonnet 5.5 · low | 80.93% | 1,214 / 1,500 | 1,500 / 1,500 |
| 04 | StartLux-Decision-27B · BF16 | 79.73% | 1,196 / 1,500 | 1,500 / 1,500 |
| 05 | SPX-CD-Flash | 79.53% | 1,193 / 1,500 | 1,500 / 1,500 |
| 06 | GPT 6 Luna · low | 79.13% | 1,187 / 1,500 | 1,500 / 1,500 |
| 07 | GLM 5.3 Flash · low | 78.73% | 1,181 / 1,500 | 1,500 / 1,500 |
| 08 | StartLux-Decision-35B-A3B · BF16 | 78.47% | 1,177 / 1,500 | 1,500 / 1,500 |
| 09 | SPX-CD-Omni | 76.27% | 1,144 / 1,500 | 1,500 / 1,500 |
| 10 | pplx-decider 27B | 75.53% | 1,133 / 1,500 | 1,500 / 1,500 |
| 11 | JPT · 9B | 74.20% | 1,113 / 1,500 | 1,500 / 1,500 |
| 12 | Qwen 3.8 Flash | 73.47% | 1,102 / 1,500 | 1,500 / 1,500 |
| 13 | Kimi K2.6 | 73.07% | 1,096 / 1,500 | 1,500 / 1,500 |
| 14 | Imajev-9B | 71.80% | 1,077 / 1,500 | 1,478 / 1,500 |
| 15 | DeepSeek V4.1 Flash | 70.80% | 1,062 / 1,500 | 1,500 / 1,500 |
| 16 | JPT · 4B | 70.33% | 1,055 / 1,500 | 1,500 / 1,500 |
| 17 | Winnow · 12B Q8 | 67.73% | 1,016 / 1,500 | 1,410 / 1,500 |
| 18 | Jev-Omni · 12B | 66.67% | 1,000 / 1,500 | 1,493 / 1,500 |
| 19 | Cloudflare Clef | 65.67% | 985 / 1,500 | 1,485 / 1,500 |
| 20 | Cloudflare Clef-Flash | 64.07% | 961 / 1,500 | 1,485 / 1,500 |
| 21 | Reflex · Stable 4B | 62.87% | 943 / 1,500 | 1,387 / 1,500 |
| 22 | Winnow · E4B Q8 | 60.40% | 906 / 1,500 | 1,410 / 1,500 |
The multimodal filter shows models evaluated on image questions.
Pooled ECE calibration leaderboard
Taller bars mean lower error ↑View exact ECE results
| Rank | Model / version | ECE ↓ |
|---|---|---|
| 01 | SPX-CD-Pro | 2.46% |
| 02 | SPX-CD-Flash | 2.49% |
| 03 | StartLux-Decision-27B · BF16 | 2.65% |
| 04 | StartLux-Decision-35B-A3B · BF16 | 2.71% |
| 05 | pplx-decider 27B | 3.08% |
| 06 | Reflex · Stable 4B | 3.33% |
| 07 | JPT · 9B | 3.47% |
| 08 | Plumb · v5.2 4B | 3.53% |
| 09 | Imajev-9B | 3.74% |
| 10 | GLiNER2.5-Decide | 3.98% |
| 11 | Mapika · Decider 4B v2.1 | 4.08% |
| 12 | Kev · 9B | 4.09% |
| 13 | OpenDecider · Small TD | 4.21% |
| 14 | Decision Lux · 9B | 4.36% |
| 15 | Malkuth · 4B | 4.43% |
| 16 | Winnow · E4B Q8 | 4.56% |
| 17 | OpenDecider · Small | 4.58% |
| 18 | Laya · Typed Decisions | 4.86% |
| 19 | Jev | 4.99% |
| 20 | OpenDecider · Nano | 5.00% |
| 21 | JevK5 | 5.05% |
| 22 | SPX-CD-Omni | 5.11% |
| 23 | Kev · 0.8B | 5.27% |
| 24 | Nimble · v2 9B | 5.28% |
| 25 | Kev · 4B | 5.63% |
| 26 | JPT · 4B | 5.97% |
| 27 | Ateve Jev | 6.09% |
| 28 | Hopper · 4B | 6.23% |
| 29 | Bocha Jev | 6.57% |
| 30 | Lev · 4B | 6.78% |
| 31 | Cloudflare Clef-flash | 7.90% |
| 32 | Jet · v6.2 4B | 7.91% |
| 33 | Jev-Omni · 12B | 8.08% |
| 34 | Open-Jev · Zefan 9B | 8.62% |
| 35 | Mica · 4B BF16 GGUF | 8.66% |
| 36 | Winnow · 12B Q8 | 9.07% |
| 37 | CalDec · Laya | 9.24% |
| 38 | Hopper G · 4B v1.3 | 9.29% |
| 39 | Cloudflare Clef | 9.40% |
| 40 | CalDec · GLiNER | 9.40% |
| 41 | Bosun v3.1 · 0.6B | 10.67% |
| 42 | Mercury Decide | 10.77% |
| 43 | APUS OpenJev · 9B 3000 High | 12.68% |
| 44 | Bosun v3.1 · 1.7B | 13.01% |
| 45 | Laya · English | 14.42% |
| 46 | APUS OpenJev · 4B 5949 High | 14.86% |
| 47 | APUS OpenJev · 9B 5949 High | 15.45% |
| 48 | Laya · Multilingual | 19.74% |
| 49 | FLock THIS/THAT | 30.09% |
| 50 | CLM v0.1 · 8B | 31.16% |
| 51 | Julia-1 | 40.55% |
ECE uses 10 probability bins and valid choice and yes/no responses with native probabilities. Labels show the actual error. Text and image coverage varies by model.
I/C index leaderboard
0–100; higher is better ↑
View exact I/C results
| Rank | Model / version | I ↑ | C ↑ |
|---|---|---|---|
| 01 | SPX-CD-Pro | 65.41 | 83.58 |
| 02 | Mercury Decide | 64.54 | 74.53 |
| 03 | StartLux-Decision-27B · BF16 | 64.48 | 83.63 |
| 04 | Winnow · 12B Q8 | 60.60 | 68.72 |
| 05 | StartLux-Decision-35B-A3B · BF16 | 60.05 | 82.74 |
| 06 | JPT · 9B | 59.70 | 83.87 |
| 07 | Cloudflare Clef | 59.42 | 79.68 |
| 08 | Cloudflare Clef-flash | 57.46 | 82.20 |
| 09 | pplx-decider 27B | 57.36 | 80.56 |
| 10 | Plumb · v5.2 4B | 55.90 | 73.01 |
| 11 | Jev | 55.40 | 76.52 |
| 12 | Imajev-9B | 54.76 | 82.45 |
| 13 | Open-Jev · Zefan 9B | 54.17 | 78.80 |
| 14 | SPX-CD-Flash | 53.79 | 83.25 |
| 15 | APUS OpenJev · 4B 5949 High | 53.64 | 58.56 |
| 16 | APUS OpenJev · 9B 3000 High | 53.04 | 64.91 |
| 17 | Malkuth · 4B | 52.30 | 81.90 |
| 18 | Jev-Omni · 12B | 52.30 | 64.26 |
| 19 | JPT · 4B | 52.26 | 79.91 |
| 20 | APUS OpenJev · 9B 5949 High | 51.79 | 57.45 |
| 21 | Jet · v6.2 4B | 51.40 | 80.01 |
| 22 | Ateve Jev | 48.07 | 75.22 |
| 23 | Hopper G · 4B v1.3 | 47.68 | 74.10 |
| 24 | Hopper · 4B | 47.61 | 79.12 |
| 25 | FLock THIS/THAT | 47.59 | 43.08 |
| 26 | JevK5 | 46.76 | 79.66 |
| 27 | SPX-CD-Omni | 46.66 | 80.19 |
| 28 | Lev · 4B | 44.66 | 72.91 |
| 29 | Decision Lux · 9B | 44.15 | 78.17 |
| 30 | Kev · 4B | 44.02 | 79.79 |
| 31 | Mapika · Decider 4B v2.1 | 43.94 | 79.68 |
| 32 | Bocha Jev | 42.46 | 73.37 |
| 33 | Kev · 9B | 41.59 | 80.70 |
| 34 | Nimble · v2 9B | 39.76 | 83.89 |
| 35 | Mica · 4B BF16 GGUF | 39.66 | 72.19 |
| 36 | OpenDecider · Small | 39.41 | 81.38 |
| 37 | Winnow · E4B Q8 | 34.48 | 76.40 |
| 38 | OpenDecider · Small TD | 33.98 | 79.02 |
| 39 | Reflex · Stable 4B | 31.97 | 77.79 |
| 40 | Bosun v3.1 · 1.7B | 29.54 | 66.11 |
| 41 | Laya · English | 25.25 | 67.96 |
| 42 | Bosun v3.1 · 0.6B | 22.30 | 67.78 |
| 43 | OpenDecider · Nano | 19.82 | 76.75 |
| 44 | Laya · Typed Decisions | 18.68 | 78.89 |
| 45 | Laya · Multilingual | 18.43 | 53.67 |
| 46 | CalDec · Laya | 15.85 | 78.49 |
| 47 | Kev · 0.8B | 13.45 | 74.45 |
| 48 | Julia-1 | 10.34 | 27.00 |
| 49 | CLM v0.1 · 8B | 5.01 | 55.44 |
| 50 | CalDec · GLiNER | 0.00 | 67.28 |
| 51 | GLiNER2.5-Decide | 0.00 | 74.53 |
I · Intelligence: measures decision accuracy relative to random guessing.C · Calibration: measures how well predicted probabilities match the correct answers.
I/C uses the same 900 text questions for every model, applying the JevBench v1.5 formulas to CMDB results. ECE and I/C compare models that provide native candidate probabilities.
Want your model evaluated on CMDB-1500?
From text to images,
make structured decisions.
CMDB-1500 combines language, image, and action selection tasks in four formats: single choice, yes/no, ordered scoring, and multi-select. The dataset contains 1,500 questions and 322 distinct source images.
Select a category to see its share; select it again to reset.
Questions are sampled across domains, with answer options and source images preserved. Accuracy is calculated over 1,200 text questions, 300 image questions, or all 1,500 questions. Refusals and missing answers count as incorrect. Models evaluated only on text appear in the text leaderboard.
What do these
questions test?
The tasks cover language understanding, business decisions, image understanding, tool use, and action selection.
Language, knowledge & judgment450 questions
Questions from JevBench test knowledge, intent recognition, evidence verification, language understanding, sentiment, and response quality.
ARC-Challenge · MMLU · Banking77 · CLINC150 · MASSIVE · BoolQ · MNLI · ChaosNLI · FEVER · PAWS · SST-5 · STS-B · StrategyQA · Civil Comments · Measuring Hate Speech · SMS Spam · HelpSteer2
Professional & business decisions300 questions
Questions from Atlan Decision Bench test tool selection, task completion checks, citation verification, SQL selection, and business routing.
The collection includes tasks adapted from BFCL, AgentDojo, τ-bench, MT-Bench, and Spider. Models choose an answer from the supplied context.
Visual decisions300 questions
Thirty questions from each of ten visual benchmarks test visual mathematics, image and text understanding, scientific diagrams, and fashion classification. Questions that require multiple images retain all source images.
MathVista · MMMU · MMMU-Pro · AI2D · ScienceQA · Fashion200k · GQA · VQAv2 · TextVQA · VizWiz
GQA, VQAv2, TextVQA, and VizWiz contribute yes/no question subsets.
Safety & adversarial tasks100 questions
Questions from Aegis, PhishNChips, and JevAdvBench test content safety, phishing email and URL detection, and decisions under adversarial perturbations.
Adversarial questions retain their source labels, some of which are not human annotations.
Game actions100 questions
Choose an action from a given game state. OpenJev provides scenarios from racing, Snake, tic-tac-toe, platformers, runner games, and ViZDoom.
Scores measure agreement with reference actions at fixed game states.
Multi-select decisions100 questions
Fifty questions each from MultiRC and GoEmotions ask models to select all correct reading-comprehension answers or identify multiple emotions in a text.
Embodied actions50 questions
ALFRED , ALFWorld, and VirtualHome provide 50 questions about choosing the next action from a task, observations, and action history.
Reference answers come from expert trajectories and evaluate individual action choices.
Chinese transcript selection50 questions
Select the transcript closest to the reference from ChineseHP AISHELL-4 candidates. Answers are determined by character edit distance.
Inputs contain transcript candidates, without audio.
Tool use25 questions
Adapted from BFCL v3, these questions ask whether a user request requires a tool call.
Spatial & complex rules25 questions
Nineteen complex-rule questions and six spatial questions from THIS-THAT test decisions under conflicting rules and spatial relationships. Reference answers come from rules or simulators.
The same candidates.
Different ways to decide.
SPX-CD · Native decision models
SPX-CD-Pro, Flash, and Omni return structured choices and candidate probabilities. All three are evaluated with effort=2.
SPX-CD-Omni is open source on Hugging Face, with LoRA weights, inference code, and evaluation results.
Other native decision models
Jev, StartLux, Mercury Decide, pplx-decider, JPT, Winnow, and Kev use decision or probability interfaces. The leaderboard distinguishes model sizes, versions, and quantization formats.
General reasoning models
GPT, Claude, GLM, Qwen, Kimi, and DeepSeek are included in the accuracy comparison as general reasoning models.
Text and image accuracy are 79.33% / 88.67% for Pro and 77.58% / 87.33% for Flash.
Can you trust
the model’s confidence?
When a model assigns 80% probability to a set of answers, are about 80% actually correct? ECE measures the gap between predicted confidence and observed accuracy; lower is better. Pooled ECE is 2.46% for Pro and 2.49% for Flash.
Explore CMDB-1500 on Hugging Face ↗