Part 4 of 12 · Models and providers

Choose what runs the judgment

The same decision can run in a hosted service, on your own GPU, or through a small encoder. Choose the evidence, deployment, and evaluation contract before choosing a model name.

Documentation snapshot: October 1, 2026. This guide has not run a cross-provider benchmark.

A shared interface is a starting point

Jev, Clef, Kev, and Laya take evidence plus typed questions and return bounded judgments. Choice selects an option, Score reports a position on described levels, and Noul reports the probability of yes. Your application supplies candidates, retrieves context, and owns the action.

“Jev-compatible” describes a useful interface. It does not establish identical predictions, confidence statistics, token accounting, input preservation, or serving behaviour. A model switch is a new deployment to evaluate.

Can you swap one model for another?

You can often reuse the decision frame. Keep the evidence, intended labels, and action logic, then adapt the request transport and evaluate the replacement. API compatibility reduces integration work; it does not preserve measured behaviour.

LayerUsually reusableCheck or change
Decision designWhat the question means; the Choice, Score, or Noul answer shapeSupported description types, candidate coverage, rubric interpretation
Client integrationThe state/questions/answers patternEndpoint, credentials, model selector, outer response envelope, validation
Application policyThe consequences, review path, and allowed operationsAcceptance signal, threshold, accuracy, and coverage on the replacement
DeploymentThe workflow your application exposesWeights, runtime, input preservation, language routing, concurrency, latency, cost

A successful request proves that the interfaces connect. It does not prove that the replacement makes equivalent decisions. Use the migration checks before letting it drive the same actions.

Meet the model families

TypeSafe · Hosted

Jev

Text judgments through TypeSafe’s API and SDKs. The current stable version is jev-1.13.0; the latest and preview aliases currently resolve to it. Our frame-engineering studies used this version.

Model reference ↗ · Draft a Jev request →

Cloudflare · Hosted or open weights

Clef and Clef-flash

27B and 9B multimodal decision models. Workers AI serves them as @cf/cloudflare/clef and @cf/cloudflare/clef-flash. The published weights use Apache 2.0.

Launch and architecture ↗ · Draft a Clef-flash request →

Jared Palmer · Self-hosted

Kev

A Qwen-based family with 0.8B, 4B, 9B, and 27B sizes. Kev 1.0 fixes a set of checkpoints. It provides a Jev-compatible server, supports local inference and fine-tuning, and documents question isolation.

Project and releases ↗ · Draft a Kev request →

ConvAI Innovations · Local encoders

Laya

English and typed-decisions checkpoints use a 421M ModernBERT model; the 322M multilingual checkpoint uses mmBERT for 100+ languages. The family is Apache 2.0. Use the Python library or its self-hosted HTTP server.

Model card ↗ · Draft a Laya request →

Match the deployment to the work

A repair-desk feature can begin as text routing. If the feature must inspect a photograph of a damaged wheel, the evidence requirement changes. If customer records must stay inside an existing deployment, the serving arrangement matters as much as model size.

RequirementCandidate to investigateMeasure before choosing
A managed text serviceJevQuality on your frame, network latency, rate limits, review coverage
Image evidence in WorkersHosted Clef / Clef-flashWhether image details survive preprocessing and affect the right judgment
Control of weights and servingKev or local ClefHardware fit, concurrency, cold starts, operational effort
Small local multilingual judgmentsLayaLanguage routing, short evidence budgets, label descriptions, domain calibration

These are starting points for an experiment. Avoid ranking the families by parameter count or mixing local model time with hosted end-to-end response time.

Compatibility stops at important boundaries

BoundaryDocumented differenceWhat to test
Question isolationJev and Kev isolate questions. Clef’s joint schema head scores options across the schema jointly.Ask alone, then add unrelated questions. Compare probabilities and application actions.
Question IDsJev and Kev treat IDs as code identifiers. Clef’s local encoder uses an ID as instructions when instructions are omitted.Write complete instructions. Check serialization and rename-ID behaviour on the deployed surface.
ConfidenceTypeSafe’s public Jev reference leaves the exact formula unspecified. Laya documents 1 − normalized entropy and an additional answer_confidence field.Select a signal and tune a fresh threshold for every model, option count, and action.
Model identityA hosted model selector, a loaded Kev checkpoint, and a routed Laya checkpoint name identify different layers.Record service model, weights revision, runtime, tokenizer, and calibration settings where available.

Sources: Jev primitives, Clef model card, Kev API, and Laya serving contract. The isolation contrast follows the documented architectures; we have not measured its effect on our fixture.

Read the budget for the serving surface

SurfaceInput budget or behaviourListed input price
Jev 1.1364k total request; state plus longest question ≤32k$0.042 / million tokens; output free
Workers AI Clef65,536 context; 1–64 questions; long text state can be truncated$0.24 / million input tokens
Workers AI Clef-flash65,536 context; same hosted question limit$0.09 / million input tokens
Kev serverAccepts state up to 65,536 plus 8,192 per question; rejects oversized state. Validated length differs by checkpoint.Your compute and hosting costs
Laya defaults512 tokens per English question; 1,024 for multilingual / typed-decisions. State and option descriptions compete for space.Your compute and hosting costs

These budgets are not interchangeable. A request accepted by a server can still lose useful accuracy at long lengths. For Laya, test whether shortened descriptions make options indistinguishable; its HTTP server also caps Choice at 100 options. For local Clef, the published encoder defaults to 16,384 tokens, independently of the hosted limit.

The Workers AI schema accepts up to four embedded PNG, JPEG, or WebP images, with byte and pixel limits; remote image URLs are not accepted. The local Clef card also documents video frame arrays. Do not infer a hosted video request field from the local interface. Consult the linked model pages for the current full schema.

Move a frame, then prove the application

  1. Preserve the decision. Keep the intended labels, source cohort, and action consequences fixed. Translate the schema only where the destination requires it.
  2. Inspect the received evidence. Check truncation, option budgets, image preprocessing, language routing, and hidden defaults.
  3. Keep failures visible. Distinguish transport or validation failures from a real no-match answer. Normalize outer response envelopes before validating answers.
  4. Re-evaluate uncertainty. Compare errors and coverage at the same consequence budget. A copied confidence cutoff is not a migration test.
  5. Measure the service you will use. Test warm and cold requests, realistic concurrency, timeouts, retries, cost, and downstream duplicate protection.

Use the provider evaluation protocol and request builder. Our Jev findings suggest experiments; they remain Jev measurements.