When software needs to choose an action, an option identifier and a probability distribution are often more useful than an explanation that needs another parsing step. Jev and Simplex CD serve this structured decision workflow. Choosing between them requires a task definition, comparable evidence and a policy for using uncertain answers.
SPX-CD-Pro and Flash have higher accuracy than Jev in the published CMDB text results below. That finding applies to this test scope; it does not establish a winner for every application or for latency under load.
What are Jev and Simplex CD?
Jev is TypeSafe AI’s decision model. Its documented interface evaluates typed questions against a state. Simplex CD belongs to the Simplex model family developed by Surd AI; SPX-CD stands for Simplex Calibrated Decision Model. The platform offers Flash and Pro variants. Jev documentation · Simplex CD release
Both fit a design in which a model makes a judgment and application code controls the resulting action. A routing decision does not itself grant permission to execute a tool or update an account.
CMDB: compare the same 1,200 text questions
CMDB-1500 contains 1,200 text questions and 300 image questions. This table compares the shared text portion rather than mixing text-only accuracy with overall multimodal accuracy.
| Model | Text accuracy | Correct / text total | Valid / evaluated scope |
|---|---|---|---|
| SPX-CD-Pro | 79.33% | 952 / 1,200 | 1,500 / 1,500 |
| SPX-CD-Flash | 77.58% | 931 / 1,200 | 1,500 / 1,500 |
| Jev | 72.17% | 866 / 1,200 | 1,200 / 1,200 |
Pro answers 86 more text questions correctly than Jev; Flash answers 65 more. The differences calculated from counts are approximately 7.17 and 5.42 percentage points. Subtracting rounded display values can differ by 0.01.
Surd AI maintains this evaluation and evaluates its own models. SPX-CD uses effort=2 in this test. Valid-response counts cover each model’s evaluated scope. The Jev record does not include image questions; that omission alone is not a claim about every current Jev product’s capabilities. This article uses the records inspected on October 7, 2026. See the CMDB leaderboard and methodology for model details.
JevBench and Decision Index give a more nuanced picture
The release page also records SPX-CD effort=1 results alongside a published Jev 1.13.0 reference.
| Model | Decision Index 0.2.1 ↑ | JevBench Public 231 | JevBench Hard 111 |
|---|---|---|---|
| SPX-CD-Pro · effort=1 | 58.30 | 89.61% (207 / 231) | 79.28% (88 / 111) |
| SPX-CD-Flash · effort=1 | 54.65 | 88.31% (204 / 231) | 77.48% (86 / 111) |
| Jev 1.13.0 · published reference | 57.91 | 86.58% (200 / 231) | 72.97% (81 / 111) |
Flash is below Jev on Decision Index while Pro is slightly above it. Hard 111 is a subset of Public 231, not an independent set to add to the total. These records have different provenance and configurations; they are not a same-day production load comparison. Sources and effort=2 results are linked from the release evaluation section.
What should you compare beyond accuracy?
Error costs. A misrouted support ticket and an incorrectly authorized operation have different consequences. Report the confusion matrix, automated coverage, incorrect actions and review rate for your application.
Probability behavior. A concentrated distribution can still be wrong. Validate action thresholds on labeled application data rather than interpreting a reported confidence as a verified per-request success rate.
Actual response time. A single question, a batch of 64 and simultaneous users impose different workloads. Fix input length, question count, option count and concurrency; measure client P50/P95, success rate and billed usage. This article makes no speed ranking without matched measurements.
When should you evaluate Simplex CD?
Start with its examples if your workflow needs image-conditioned decisions or multiple selections. If you already use Jev, preserve your original questions and option descriptions, then shadow-test explicit Flash and Pro model IDs before changing production behavior.
Model names are not a substitute for measurement. Compare the cost of meeting your required quality level. Check the current API documentation for image and automatic routing: a returned alias alone does not identify the physical inference backend. API documentation
Frequently asked questions
Is Simplex CD a drop-in replacement for Jev?
Both support structured decision workflows, but endpoints, model identifiers, output semantics, errors and thresholds need verification. Use the migration checklist.
Does JevBench predict my application’s results?
It measures a defined task set. Your inputs, options and error costs may differ. Keep an independent application test set.
Which model should I try first?
Pro is a reasonable candidate if the CMDB text results match your priorities. To compare quality and resource costs, evaluate Flash, Pro and your existing Jev configuration on identical examples. Start in the Playground, then use the decision API evaluation guide.