If you already use Jev for classification or action selection, preserve the existing task when evaluating Simplex CD. Changing the model, question wording, options and business policy together makes it difficult to explain any difference in results.
This checklist prepares a reversible evaluation. It does not assume that every field, limit or probability behavior is identical across services.
Preserve the state, questions and option meanings
TypeSafe’s quick start organizes requests around state, model and named questions. Simplex CD uses a similar basic organization, but the destination endpoint, credentials and detailed contract still need checking. TypeSafe quick start · Simplex CD API docs
Keep question names, candidate keys and descriptions stable. Renaming billing to payments can affect application code; removing an option description can affect interpretation. Verify the original task before optimizing its wording.
| Migration item | What to verify |
|---|---|
| Endpoint | Simplex CD currently documents https://api.surdai.com/v1/systemone |
| Authentication | Use a destination-platform token held on your server |
| Model | Test spx-cd-flash and spx-cd-pro separately |
| Task semantics | Preserve the meaning of state, instructions and criteria |
| Response handling | Check answers, choice, probabilities and error responses |
| Action policy | Revalidate thresholds, review rules and authorization |
Start with one explicit choice question
This request-body example follows the Simplex CD documentation. It routes a support request and includes an other option for inputs outside the defined categories.
{
"model": "spx-cd-flash",
"state": {
"message": "I was charged twice. Please check my bill."
},
"questions": {
"department": {
"type": "choice",
"instructions": "Choose the department for the primary request.",
"criteria": {
"billing": "Invoices, charges and refunds",
"technical": "Product failures and integration issues",
"other": "Neither category, or insufficient information"
}
}
}
}
Use the Playground to inspect the response, then repeat the same input with Pro. A human can label the expected answer billing; record the actual response rather than treating the expectation as a successful test.
Application code should check that department exists, choice belongs to its candidate set and required probability fields are valid. An error or missing field should follow the failure policy rather than silently becoming other or a zero-risk judgment.
Revalidate thresholds after switching providers
The same selected category can come with a different probability distribution. Preserve distributions and inspect how thresholds affect automated coverage and error rates on application validation data.
A practical sequence is to log both models’ answers while existing rules keep controlling actions. Review disagreements, choose thresholds on validation data, then evaluate a held-out test set. The evaluation guide describes these metrics.
Validate multi-select, images and scores separately
Simplex CD exposes choice, multi_choice, noul and score. Add new task types only after checking the existing text workflow. Capabilities and examples
For multi-select, specify the permitted number of selections and whether an empty result is valid. For ordered scoring, check the order of criteria and the meaning of the returned score. Image inputs need their own format, size and routing checks. A successful text Choice test does not verify these other contracts.
Do not invent an effort request field from benchmark tables. The tables describe evaluation configurations; supported fields and current service defaults belong to the API documentation.
Bound concurrency and measure actual cost
Begin with controlled traffic. Log question count, input size, client duration, valid responses and billed usage. For batches, report the full response time as well as completed question count; amortized time is not the caller’s waiting time.
Honor 429 responses with bounded backoff. A timeout may occur after the service has already performed the work. Avoid unbounded retries and prevent downstream business actions from executing twice. Keep the original-service fallback during rollout, with explicit retry and fallback limits.
Keep a migration acceptance record
| Check | Acceptance criteria come from |
|---|---|
| Choice values and candidate keys | Application contract and offline tests |
| Business error rate | The cost of each type of mistake |
| Review rate | Operational capacity |
| P95 and success rate | Product response-time requirements |
| Cost per valid completion | Billed usage and completed tasks |
| Fallback and repeated actions | Error handling and idempotency policy |
Frequently asked questions
Can I migrate by changing only the base URL?
That does not establish compatibility. Change credentials, select a supported model ID and verify fields, limits and output semantics.
Should I rewrite every question?
Preserve questions for the first comparison. Separate model replacement from prompt improvement so that differences remain interpretable.
Can I shadow-test alongside production?
Yes, after checking permission to use the data, extra cost and capacity. Shadow results should not trigger business actions.
Review the Jev and Simplex CD comparison, then start with one representative task you are authorized to use for testing.