Models & interaction guides
Explore JevBench, Image JevBench, AudioJevBench, CMDB-1500 and Decision Index guides, alongside decision APIs, robot interaction and full-duplex speech.
Benchmarks & leaderboards
JevBench, Image JevBench, AudioJevBench, CMDB-1500 and Decision Index: datasets, versions, probabilities and results.
Image JevBench, Imajev and CMDB: evaluating image decisions
Separate Image JevBench from the Imajev model project, compare versions and image coverage, and evaluate visual decision accuracy, calibration and latency.
Read guide ↗Audio benchmarksAudioJev and AudioJevBench: how to evaluate audio decisions
Distinguish AudioJev projects from AudioJevBench, compare direct audio and transcription pipelines, and read coverage, probability and latency evidence.
Read guide ↗JevBenchReading JevBench: Public 231, Hard 111 and SPX-CD results
Compare SPX-CD-Pro, Flash and Jev on Public 231 and Hard 111, with effort settings and ECE, without confusing public accuracy with the full leaderboard.
Read guide ↗CMDB-1500CMDB-1500 leaderboard guide: text, images and calibration
Read CMDB-1500 text and image accuracy, fixed denominators, valid coverage, ECE and I/C, with links to SPX-CD, Jev and other decision model results.
Read guide ↗Decision IndexDecision Index 0.2.1 vs 0.3: public results and the Full score
Understand the 150,317-request Decision Index 0.2.1 evaluation, public versus Full scores, and versioned SPX-CD and Jev results on Hugging Face.
Read guide ↗Interaction & speech
Concepts, engineering and published research. Interaction models are not publicly available yet.
What is a robot interaction model? From commands to collaboration
Understand robot interaction models, speech models, visual understanding and control, with practical checks for interruptions, task state and collaboration.
Read guide ↗Full-duplex interactionWhat is a full-duplex interaction model? Listening while speaking
Understand full-duplex interaction, streaming speech, user interruptions and task cancellation, with practical measurements for real-time conversational systems.
Read guide ↗Voice interaction basicsWhen should a voice agent respond? VAD, turn-taking and interruptions
Learn the differences between voice activity, end-of-turn and interruption detection, how to measure timing, and what Simplex MicroCF’s public results show.
Read guide ↗Decision models & APIs
Simplex CD vs Jev: choosing a decision model
Compare Jev, SPX-CD-Flash and SPX-CD-Pro using shared CMDB text tasks, JevBench results and practical API evaluation criteria.
Read guide ↗Evaluation methodsHow to evaluate a decision API: accuracy, probability and latency
Build a fair evaluation for Jev and Simplex CD. Measure error costs, calibration, batch throughput and client latency before choosing action thresholds.
Read guide ↗Integration guideMigrating from Jev to Simplex CD: a decision API checklist
Evaluate a Jev-to-Simplex-CD migration without changing the task. Check model IDs, output contracts, thresholds, retries and shadow-test results.
Read guide ↗Start with evidence.
Inspect the published benchmark or run a decision with your own input.