← Models & interaction guides
GUIDES / Audio benchmarks

AudioJev and AudioJevBench: how to evaluate audio decisions

Distinguish AudioJev projects from AudioJevBench, compare direct audio and transcription pipelines, and read coverage, probability and latency evidence.

Surd AI · Research & engineeringEnglish

An “okay” can express agreement or hesitation. A transcript may omit tone, speaker identity and background sounds. Before comparing audio decision scores, establish whether the system heard the waveform or only received recognized words.

AudioJev names a model project; AudioJevBench names an evaluation

More than one project uses AudioJev. The shlv/AudioJev card describes a Qwen2.5-Omni-3B-derived candidate probability model. mocomoco-inc/AudioJev-AudioDecisionModel is a separate project. Identify the full repository when citing weights, results or licensing. shlv model card · mocomoco model card

The AudioJevBench v0.1 page checked on October 8, 2026 reports 988 scored items: 112 public and 876 sealed. It separates full-coverage, public-only and partial systems. Its AudioJev row identifies mocomoco edge ONNX, so that row should not be assigned to the shlv model. AudioJevBench

Compare direct audio with transcription as different systems

Suppose the question asks whether a horn sounded. If a transcription pipeline does not retain that event, the text decision model has lost necessary evidence. For a spoken ticket number, recognition of the words may instead be the main bottleneck.

Our suggested evaluation treats raw-audio input and transcription followed by a text decision as separate systems. Include recognition time and cost in the latter. Diagnose whether the error arose from missing evidence in the transcript or an incorrect judgment after transcription.

Test condition What to record
Audio Language, duration, noise and overlapping speakers
Decision Goal, candidate descriptions and fallback option
Coverage Complete suite, public subset or one task family
Probabilities Confident errors and threshold-specific outcomes
Service time Transfer, preprocessing, recognition if used, and inference

These are proposed application checks. Two rows with different coverage do not become directly comparable merely because both display percentages.

Candidate-order stability and probability calibration need separate checks

Keeping the same semantic choice after an option shuffle measures stability. Calibration asks whether a group of predictions assigned roughly 80% confidence is correct roughly 80% of the time. Test both properties rather than substituting one for the other.

Use independently labeled audio, repeat candidate permutations and inspect confidence buckets on held-out examples. Resolve ambiguous annotation rules before treating a probability discrepancy as a model defect.

Audio decisions do not establish full-duplex interaction quality

Offline answers leave timing and task recovery untested. A live system also needs to respond at the appropriate moment, stop playback and recover after interruption. Surd AI’s interaction models remain research previews. This guide announces no new audio API and claims no unverified AudioJevBench result. Read the full-duplex guide and turn-taking guide for those evaluation questions.

Frequently asked questions

Can an AudioJev score be used as a Jev score?

No. Establish the project, revision, input modality and evaluation before assigning a result to a model.

Can public-only results outrank full-coverage results directly?

Prefer a shared test set, or retain the coverage distinction. Public results do not establish performance on unseen sealed items.

Do Simplex CD image results establish audio capability?

No. Each modality requires evidence. Refer to the API documentation for currently available Simplex CD inputs.

SIMPLEX CD / API

Try a decision with your own input.

Create an account, explore decision models in the Console, or get started with the API documentation.