← Models & interaction guides
GUIDES / Evaluation methods

How to evaluate a decision API: accuracy, probability and latency

Build a fair evaluation for Jev and Simplex CD. Measure error costs, calibration, batch throughput and client latency before choosing action thresholds.

Surd AI · Research & engineeringEnglish

A higher benchmark score does not guarantee fewer errors in your application. Your options may be ambiguous, your inputs may be long, and one class of mistake may matter much more than another. Evaluate a decision API by starting with the action your software will take.

What is a decision model API?

A decision API maps context and a defined question to a value your program can use: a category, a judgment about a condition, or a position on an ordered rubric. Fixed outputs suit routing and filtering. Generated text suits explanations and open-ended composition. An application can use both.

Make each question express one judgment. Separating importance, urgency and eligibility can make errors easier to diagnose than asking for all three in one vague score. Your code can combine those judgments using an explicit policy.

Fix the task before comparing models

For a support router, define billing, technical and other with descriptions that explain their boundaries. Decide how mixed-topic tickets should be handled before labeling them.

Evaluation input What to record
Dataset Source, deduplication, time period and language mix
Reference labels Labeling rules, ambiguity and review process
Request Identical context, question, options and ordering
Configuration Model ID, date, service defaults and image use
Error costs Recoverable mistakes and cases needing review

Use a validation set to choose thresholds and a separate held-out test set for the final report. Repeatedly tuning against the final test questions makes your estimate optimistic.

Measure error direction and automated coverage

For allow, review and block judgments, overall accuracy hides the difference between an unsafe release and an unnecessary rejection. Report each direction in a confusion matrix.

When uncertain requests go to a human, report both the fraction handled automatically and the error rate within that fraction. A system that handles only a few easy cases may look accurate without saving much work.

Keep three views: task metrics across all requests, automated coverage and errors among automated actions. Break them down by language, input length, option count and task type.

Check what a probability field means

For many selected answers assigned probabilities near 0.9, inspect how often that group is actually correct. This checks a pattern across labeled examples; one answer’s probability cannot verify its own correctness.

Expected calibration error, or ECE, depends on sample selection, bins and calculation details. Fix those details for a comparison. Lower ECE alone does not establish higher accuracy or lower business risk.

An API’s confidence field may also differ from the selected option’s probability. Under TypeSafe’s documented Choice definition, a top probability of 0.6 among three candidates gives confidence 0.4. Verify field semantics before transferring a threshold. TypeSafe confidence definition

Separate batch amortization from response time

Suppose a request completes 64 questions in eight seconds. That is 125 milliseconds per question when amortized, but the caller waits eight seconds for the batch. An amortized time is not a single-request latency. These numbers are an arithmetic example, not a measured provider result.

Metric What it tells you
Client P50 / P95 Typical and slower response times
Successful question throughput Valid completed work per second
429, timeout and failure rates Whether load exceeds available capacity
inference_ms Server timing, interpreted using the API’s definition
Billed usage Cost per successfully completed task

Measure one question, then representative batches. Test each at a single-user load and at expected concurrency. Record the sizes of the shared context, questions and options. A short single-question request is not comparable to a large long-context batch.

Simplex CD reports inference_ms for inference-server decision duration, which differs from client round-trip time. If a response does not provide a trustworthy timing field, record it as unknown rather than zero. See the API timing documentation.

Apply this to Jev, Flash and Pro

Use published comparisons to shortlist candidates, then evaluate your actual inputs. Start with explicit model IDs; test automatic routing as a separate policy. Count failed requests instead of silently dropping them.

If Pro prevents enough costly errors to justify its resource cost, it may be the right choice. If Flash already meets your requirements, a higher leaderboard position alone is not a reason to change. Compare the total cost of reaching the same application target.

Frequently asked questions

How large should my test set be?

Cover important classes and realistic edge cases first. Distinguishing small gains or measuring rare errors requires more observations and uncertainty analysis. A handful of selected examples cannot establish a reliable ranking.

Can I use 0.8 as an automatic-action threshold?

First establish whether the value is a probability or another confidence statistic. Choose thresholds using a separate validation set and the consequences of each action.

Does shared context guarantee faster batches?

No. Savings depend on actual computation reuse, cache hits and scheduling. Compare end-to-end time and valid throughput for the same amount of work.

Start with a labeled example in the Simplex CD Playground, then save requests and responses so the experiment can be repeated.

SIMPLEX CD / API

Try a decision with your own input.

Create an account, explore decision models in the Console, or get started with the API documentation.