Choose what runs the judgment
The same decision can run in a hosted service, on your own GPU, or through a small encoder. Choose the evidence, deployment, and evaluation contract before choosing a model name.
Documentation snapshot: October 1, 2026. This guide has not run a cross-provider benchmark.
A shared interface is a starting point
Jev, Clef, Kev, and Laya take evidence plus typed questions and return bounded judgments. Choice selects an option, Score reports a position on described levels, and Noul reports the probability of yes. Your application supplies candidates, retrieves context, and owns the action.
“Jev-compatible” describes a useful interface. It does not establish identical predictions, confidence statistics, token accounting, input preservation, or serving behaviour. A model switch is a new deployment to evaluate.
Can you swap one model for another?
You can often reuse the decision frame. Keep the evidence, intended labels, and action logic, then adapt the request transport and evaluate the replacement. API compatibility reduces integration work; it does not preserve measured behaviour.
| Layer | Usually reusable | Check or change |
|---|---|---|
| Decision design | What the question means; the Choice, Score, or Noul answer shape | Supported description types, candidate coverage, rubric interpretation |
| Client integration | The state/questions/answers pattern | Endpoint, credentials, model selector, outer response envelope, validation |
| Application policy | The consequences, review path, and allowed operations | Acceptance signal, threshold, accuracy, and coverage on the replacement |
| Deployment | The workflow your application exposes | Weights, runtime, input preservation, language routing, concurrency, latency, cost |
A successful request proves that the interfaces connect. It does not prove that the replacement makes equivalent decisions. Use the migration checks before letting it drive the same actions.
Meet the model families
Jev
Text judgments through TypeSafe’s API and SDKs. The current stable version is jev-1.13.0; the latest and preview aliases currently resolve to it. Our frame-engineering studies used this version.
Clef and Clef-flash
27B and 9B multimodal decision models. Workers AI serves them as @cf/cloudflare/clef and @cf/cloudflare/clef-flash. The published weights use Apache 2.0.
Kev
A Qwen-based family with 0.8B, 4B, 9B, and 27B sizes. Kev 1.0 fixes a set of checkpoints. It provides a Jev-compatible server, supports local inference and fine-tuning, and documents question isolation.
Laya
English and typed-decisions checkpoints use a 421M ModernBERT model; the 322M multilingual checkpoint uses mmBERT for 100+ languages. The family is Apache 2.0. Use the Python library or its self-hosted HTTP server.
Match the deployment to the work
A repair-desk feature can begin as text routing. If the feature must inspect a photograph of a damaged wheel, the evidence requirement changes. If customer records must stay inside an existing deployment, the serving arrangement matters as much as model size.
| Requirement | Candidate to investigate | Measure before choosing |
|---|---|---|
| A managed text service | Jev | Quality on your frame, network latency, rate limits, review coverage |
| Image evidence in Workers | Hosted Clef / Clef-flash | Whether image details survive preprocessing and affect the right judgment |
| Control of weights and serving | Kev or local Clef | Hardware fit, concurrency, cold starts, operational effort |
| Small local multilingual judgments | Laya | Language routing, short evidence budgets, label descriptions, domain calibration |
These are starting points for an experiment. Avoid ranking the families by parameter count or mixing local model time with hosted end-to-end response time.
Compatibility stops at important boundaries
| Boundary | Documented difference | What to test |
|---|---|---|
| Question isolation | Jev and Kev isolate questions. Clef’s joint schema head scores options across the schema jointly. | Ask alone, then add unrelated questions. Compare probabilities and application actions. |
| Question IDs | Jev and Kev treat IDs as code identifiers. Clef’s local encoder uses an ID as instructions when instructions are omitted. | Write complete instructions. Check serialization and rename-ID behaviour on the deployed surface. |
| Confidence | TypeSafe’s public Jev reference leaves the exact formula unspecified. Laya documents 1 − normalized entropy and an additional answer_confidence field. | Select a signal and tune a fresh threshold for every model, option count, and action. |
| Model identity | A hosted model selector, a loaded Kev checkpoint, and a routed Laya checkpoint name identify different layers. | Record service model, weights revision, runtime, tokenizer, and calibration settings where available. |
Sources: Jev primitives, Clef model card, Kev API, and Laya serving contract. The isolation contrast follows the documented architectures; we have not measured its effect on our fixture.
Read the budget for the serving surface
| Surface | Input budget or behaviour | Listed input price |
|---|---|---|
| Jev 1.13 | 64k total request; state plus longest question ≤32k | $0.042 / million tokens; output free |
| Workers AI Clef | 65,536 context; 1–64 questions; long text state can be truncated | $0.24 / million input tokens |
| Workers AI Clef-flash | 65,536 context; same hosted question limit | $0.09 / million input tokens |
| Kev server | Accepts state up to 65,536 plus 8,192 per question; rejects oversized state. Validated length differs by checkpoint. | Your compute and hosting costs |
| Laya defaults | 512 tokens per English question; 1,024 for multilingual / typed-decisions. State and option descriptions compete for space. | Your compute and hosting costs |
These budgets are not interchangeable. A request accepted by a server can still lose useful accuracy at long lengths. For Laya, test whether shortened descriptions make options indistinguishable; its HTTP server also caps Choice at 100 options. For local Clef, the published encoder defaults to 16,384 tokens, independently of the hosted limit.
The Workers AI schema accepts up to four embedded PNG, JPEG, or WebP images, with byte and pixel limits; remote image URLs are not accepted. The local Clef card also documents video frame arrays. Do not infer a hosted video request field from the local interface. Consult the linked model pages for the current full schema.
Move a frame, then prove the application
- Preserve the decision. Keep the intended labels, source cohort, and action consequences fixed. Translate the schema only where the destination requires it.
- Inspect the received evidence. Check truncation, option budgets, image preprocessing, language routing, and hidden defaults.
- Keep failures visible. Distinguish transport or validation failures from a real no-match answer. Normalize outer response envelopes before validating answers.
- Re-evaluate uncertainty. Compare errors and coverage at the same consequence budget. A copied confidence cutoff is not a migration test.
- Measure the service you will use. Test warm and cold requests, realistic concurrency, timeouts, retries, cost, and downstream duplicate protection.
Use the provider evaluation protocol and request builder. Our Jev findings suggest experiments; they remain Jev measurements.