A visual decision can be narrower than describing an image: select the next action in a screenshot, check an inspection condition or assign probabilities to supplied descriptions. Image support is a starting requirement. The useful comparison is how reliably a model solves the visual tasks your application actually presents.
ImageJev, Image JevBench and Imajev: identify the project
When researching ImageJev, check the complete name. Image JevBench is Benchmark Heaven’s image decision evaluation. Imajev is a separate model project accepting images and typed questions. Similar names do not identify one product. This guide was checked on October 8, 2026. Image JevBench · Imajev project
Match the benchmark version before comparing scores
The checked Image JevBench v0.3.0 page describes a fresh draw of 300 public and 1,200 sealed items. It separates carried v0.1.5 results from the new ranking scale. Its Capability measure combines Intelligence and Calibration. Version and method
For an application comparison, record dataset version, model revision, inference settings and coverage. A changed image, option set or scoring rule can change the meaning of a result even when the model name stays the same. Investigate the protocol before attributing a score change to improved weights.
Evaluate task accuracy, probabilities and service time separately
| Measurement | Practical question |
|---|---|
| Task accuracy | Did the model choose the correct option using evidence in the image? |
| Probability quality | Are confident errors concentrated in OCR, small objects or spatial relationships? |
| Coverage | How many supplied images produced usable decisions? |
| End-to-end time | How long did upload, processing, inference and response delivery take? |
Consider a screenshot with visually similar submit and cancel buttons. Recognizing both buttons does not establish which action fits the current goal. A useful application test keeps the screen fixed and changes the task. This isolates instruction following from object recognition; it is a proposed evaluation design, not a reported benchmark result.
Use the CMDB image track as a separate source of evidence
CMDB-1500 includes 1,200 text questions and 300 image questions. Its image track uses its own question set and denominator. Choose the image view in the CMDB leaderboard to inspect results and coverage.
When comparing SPX-CD with another system, preserve the images, resolution, candidate descriptions and grading rule. An untested image track should remain unreported rather than being inferred from text performance. The CMDB guide explains how to read mixed text and image results.
Frequently asked questions
Is an Image JevBench score an accuracy percentage?
Read the column definition. A transformed or composite score cannot be relabeled as the percentage of questions answered correctly.
Does this guide report an official SPX-CD Image JevBench result?
No verified SPX-CD entry in that evaluation is included here. The Surd AI visual results referenced here belong to CMDB and retain that name.
Where can I try image decisions?
Read the Simplex CD release and the image-input and routing sections of the API documentation. Record both the requested and returned model identifiers in your test.