An “okay” can express agreement or hesitation. A transcript may omit tone, speaker identity and background sounds. Before comparing audio decision scores, establish whether the system heard the waveform or only received recognized words.
AudioJev names a model project; AudioJevBench names an evaluation
More than one project uses AudioJev. The shlv/AudioJev card describes a Qwen2.5-Omni-3B-derived candidate probability model. mocomoco-inc/AudioJev-AudioDecisionModel is a separate project. Identify the full repository when citing weights, results or licensing. shlv model card · mocomoco model card
The AudioJevBench v0.1 page checked on October 8, 2026 reports 988 scored items: 112 public and 876 sealed. It separates full-coverage, public-only and partial systems. Its AudioJev row identifies mocomoco edge ONNX, so that row should not be assigned to the shlv model. AudioJevBench
Compare direct audio with transcription as different systems
Suppose the question asks whether a horn sounded. If a transcription pipeline does not retain that event, the text decision model has lost necessary evidence. For a spoken ticket number, recognition of the words may instead be the main bottleneck.
Our suggested evaluation treats raw-audio input and transcription followed by a text decision as separate systems. Include recognition time and cost in the latter. Diagnose whether the error arose from missing evidence in the transcript or an incorrect judgment after transcription.
| Test condition | What to record |
|---|---|
| Audio | Language, duration, noise and overlapping speakers |
| Decision | Goal, candidate descriptions and fallback option |
| Coverage | Complete suite, public subset or one task family |
| Probabilities | Confident errors and threshold-specific outcomes |
| Service time | Transfer, preprocessing, recognition if used, and inference |
These are proposed application checks. Two rows with different coverage do not become directly comparable merely because both display percentages.
Candidate-order stability and probability calibration need separate checks
Keeping the same semantic choice after an option shuffle measures stability. Calibration asks whether a group of predictions assigned roughly 80% confidence is correct roughly 80% of the time. Test both properties rather than substituting one for the other.
Use independently labeled audio, repeat candidate permutations and inspect confidence buckets on held-out examples. Resolve ambiguous annotation rules before treating a probability discrepancy as a model defect.
Audio decisions do not establish full-duplex interaction quality
Offline answers leave timing and task recovery untested. A live system also needs to respond at the appropriate moment, stop playback and recover after interruption. Surd AI’s interaction models remain research previews. This guide announces no new audio API and claims no unverified AudioJevBench result. Read the full-duplex guide and turn-taking guide for those evaluation questions.
Frequently asked questions
Can an AudioJev score be used as a Jev score?
No. Establish the project, revision, input modality and evaluation before assigning a result to a model.
Can public-only results outrank full-coverage results directly?
Prefer a shared test set, or retain the coverage distinction. Public results do not establish performance on unseen sealed items.
Do Simplex CD image results establish audio capability?
No. Each modality requires evidence. Refer to the API documentation for currently available Simplex CD inputs.