Find out what actually improves
A confident answer, a cheaper request, and a reliable workflow are different outcomes. Measure the decision your application takes, then investigate the request that produced it.
Our Jev experiments are documented below. The metric lab uses synthetic teaching data.
Freeze the baseline before changing it
Define success in application terms: correctly suggested queues, missed fitting requests, messages needing review, or time to a useful result. Preserve the exact request, raw answer, failure status, resolved model, dataset version, and policy version for each item.
Use labelled cases from the work the application will see, including mixed intent, missing evidence, out-of-scope content, and rare but consequential cases. Have people resolve disagreements against the same task definition. A maintainer’s historical label may reflect information absent from the request; check that the frame contains the evidence needed to reproduce it.
If validation misses guide a rewrite, those cases have become tuning data. Keep related messages and near duplicates in the same split. Criteria examples need their own disjoint pool; a paraphrase of a test item can leak its meaning just as an exact copy can.
Choose sample size for the error bound, class coverage, and operating region you need to assess. Three repeats or a few hundred items are not universal evidence of reliability.
Measure the action and keep its denominator
For a queue suggestion, the useful question is how often accepted suggestions are right, and how many incoming messages receive one. A threshold can improve the first number while leaving a rare queue almost entirely for review.
| Measure | Count it as | What it reveals |
|---|---|---|
| Accepted accuracy | Correct accepted answers ÷ all accepted answers | Mistakes among automatic suggestions |
| End-to-end coverage | Accepted answers ÷ all incoming items | How much work the workflow actually handles |
| Per-class recall at policy | Correct accepted answers for a class ÷ all labelled items in that class | The class the workflow fails to serve |
| Per-class precision at policy | Correct accepted answers for a class ÷ all accepted predictions of that class | The trustworthiness of that suggested class |
| Failure rate | Items without usable answers ÷ all incoming items | Service and integration losses that accuracy can hide |
For Noul, report precision and recall at the action’s cutoff, the positive prevalence, and false-positive counts. For Score, compare error on the ordered rubric and the downstream ranking or threshold behaviour. With rare positives, inspect precision-recall performance alongside ROC-AUC. Confidence averages alone establish none of these.
The threshold changes who gets served
Twelve invented repair-desk records: ten answers and two failed calls. Move the cutoff or inspect a queue. Counts include failures in coverage and recall, while accuracy uses accepted answers only.
| Case | Label | Answer | Confidence | Policy outcome |
|---|
Change one axis, then inspect paired cases
Keep the same items in each comparison. When a condition loses answers, compare accuracy on the commonly answered items and report the failures separately. Otherwise a variant can appear better merely by failing on difficult cases.
| Experiment | Keep fixed | Inspect |
|---|---|---|
| Remove an exclusion or example | State, model, labels, other wording | Which boundary errors change; tokens and action coverage |
| Reverse Choice option order | Option meanings and input | Label flips and cutoff crossings; keep production order fixed for calibration |
| Repeat identical requests | The complete request | Label agreement, probability shifts, action flips |
| Grow an item batch | Question template and reference form | Per-position and per-class errors, failures, usage, response time |
| Move definitions into shared state | Intended judgment and evaluation items | Probability shifts and token savings; retune the changed question |
| Insert adversarial content | Trusted policy and ordinary items | Effects on the item’s own answer and on neighbouring items |
For batch experiments, first compare the original request with the batch template at size one. This catches wording or shared-definition changes before you attribute their effects to batch size. Rotate position and batch composition when testing the form you intend to ship.
Do not reorder Score levels as an option-order test: their order defines the scale. Do not interpret a repeated mistake as independent evidence that majority voting will repair it.
What we learned from Jev experiments
We evaluated jev-1.13.0 on 50 invented support tickets with hand-written labels. Most comparisons used three repeats on September 20, 2026; we ran shared-definition follow-ups on September 24. These experiments underpin this guide and our agent skill, v0.9.4. The methods, fixture, and tools live in our research repository.
We used 30 development and 20 validation items, with labels written by one researcher. This small fixture demonstrates the workflow; it is not an audited holdout. Some probe cases and raw run bodies are outside the distributed fixture, so the published materials do not provide a complete reproduction record for every figure. Use these results to design experiments on your own workload.
| What we measured | What to test in your application |
|---|---|
Missing evidence can look confident. Nine of ten deficient inputs had an arbitrary winner at confidence ≥ 0.95 without an escape option; all ten chose other at ≥ 0.97 when it was offered. | Test empty, irrelevant, wrong-field, and incomplete state. Check the answer space before relying on confidence. |
| References mattered in long arrays. With 48 items, numeric references produced 23–24 correct queue labels; keyed references produced 47; quoted references produced 46. This comparison scores 48 items, not 50. | Compare reference forms at intended batch sizes, with position and neighbour analysis. A path is a model-interpreted reference. |
| More questions differed from more items. Over the same 16-item keyed state, requests with 16, 48, or 240 questions scored 47–48 of 48 queue answers in each repeat. | Separate question-count sweeps from state-growth sweeps. Neither study establishes unlimited free fan-out. |
| Shared definitions changed the signal. At eight items, pointer definitions saved 12% of input tokens for the Choice and a further 16% for the Nouls. Pointer wording raised the urgency Noul on 40–46 of 50 items; effects on false positives differed at cutoffs 0.5 and 0.8. | Compare inline and shared definitions at size one, then at scale. Select thresholds for the resulting question, not for its predecessor. |
| Labels stayed still while numbers moved. Five identical runs over 50 items kept every label, but Choice probabilities moved by up to 0.08 and Noul outputs by up to 0.09. | Measure flips of the application action as well as the winning label. Repeatability is separate from correctness. |
| A winner did not establish exclusive applicability. Choice assigned ≥ 0.99 to one option on 66% of items; separately worded Nouls found two applicable labels on 36%. | Use Choice for competing alternatives and per-label Nouls for several applicable conditions. Evaluate them as distinct judgments. |
| Two injection probes spared keyed and quoted neighbours. Their queue labels remained at 28–29 correct of 29; numeric references worsened in those runs. | Test broader attacks on target and neighbour answers. Naming and quoting are not an injection defense or authorization boundary. |
Read the fixture notes for inputs, denominators, and reproduction commands, and our agent skill for probe details. The useful result is a set of experiments to repeat on your data. None of these counts establishes a universal batch limit, a security guarantee, or a production threshold.
Select a cutoff; validate the accepted region
Use tuning data to propose a rule. Check cumulative precision, coverage, per-class recall, and sample support at that rule. A perfect bucket containing two items does not establish a low error rate. Report appropriate uncertainty estimates and inspect the action-specific costs.
The cost calculation uses correctness probability under explicit assumptions. Returned confidence is a candidate routing signal, not that probability. A coarse fallback needs its own evaluation; a wrong fine label does not automatically make its parent right.
Once the frame and policy are fixed, evaluate them on untouched data. Recheck when the model, questions, candidate set, state construction, batch form, language, traffic prevalence, or action costs change. Keep results from different resolved models separate.
Use our evaluation tools with their limits in view
Our agent skill includes standard-library Python helpers for labelled evaluation, description ablation, batch sweeps, and token probes. Its wrapper fields such as reference, shared_state, and array_field configure those helpers; they are not fields in the TypeSafe HTTP request.
Read the evaluation protocol before running them. They call the live API and incur charges. The fixture’s reproduction commands do not create a production benchmark. Batch warning flags are diagnostic heuristics, not significance tests; the scorer does not fit a calibrator or implement your application’s review policy.
Real cost comes from reported usage across the whole workflow. Include retries and failures when usage is available, and record gaps when it is not. Request latency divided by batch size is amortized time, not the time each item waits. Raw evaluation output can contain source text; keep credentials and retained records protected.
Next, use the build guide to connect a reviewed frame to an application with an explicit failure path.
Compare providers on the decisions you will ship
Keep one source cohort and one label definition. Translate the frame into each supported interface, record the translation, and verify that every model received the same relevant evidence. A silently shortened state or option description makes this a different experiment.
- Fix the cohort. Include language, long inputs, missing facts, rare classes, and the action boundaries that matter. Keep tuning and final reporting separate for every model.
- Test the contract. Add an unrelated question, permute options, cross the input budget, simulate a failed response, and inspect the resolved model or checkpoint. Use the migration checklist.
- Tune each policy separately. Select confidence or relevant probabilities using that model’s tuning outcomes. Compare coverage at a common error budget, not performance at an inherited numerical cutoff.
- Measure the complete service. Include transport, cold starts, model loading, retries, concurrency, and review work. Self-hosted compute has a cost even when no token invoice arrives.
- Report the deployment you tested. Preserve weights, tokenizer, calibration, runtime, question schema, and truncation settings. Our Jev fixture findings are hypotheses to test elsewhere, not measurements of Clef, Kev, or Laya.
Provider benchmark tables can identify workloads to investigate. They do not establish a universal winner or a safe threshold for this application.