Part 9 of 12 · Evaluate the frame

Find out what actually improves

A confident answer, a cheaper request, and a reliable workflow are different outcomes. Measure the decision your application takes, then investigate the request that produced it.

Our Jev experiments are documented below. The metric lab uses synthetic teaching data.

Freeze the baseline before changing it

Define success in application terms: correctly suggested queues, missed fitting requests, messages needing review, or time to a useful result. Preserve the exact request, raw answer, failure status, resolved model, dataset version, and policy version for each item.

Use labelled cases from the work the application will see, including mixed intent, missing evidence, out-of-scope content, and rare but consequential cases. Have people resolve disagreements against the same task definition. A maintainer’s historical label may reflect information absent from the request; check that the frame contains the evidence needed to reproduce it.

Design and tuneUse these cases to inspect errors and change wording.
Select the policyUse separate validation cases to choose thresholds and batch size.
Report onceKeep a final set untouched until the design and policy are fixed.

If validation misses guide a rewrite, those cases have become tuning data. Keep related messages and near duplicates in the same split. Criteria examples need their own disjoint pool; a paraphrase of a test item can leak its meaning just as an exact copy can.

Choose sample size for the error bound, class coverage, and operating region you need to assess. Three repeats or a few hundred items are not universal evidence of reliability.

Measure the action and keep its denominator

For a queue suggestion, the useful question is how often accepted suggestions are right, and how many incoming messages receive one. A threshold can improve the first number while leaving a rare queue almost entirely for review.

MeasureCount it asWhat it reveals
Accepted accuracyCorrect accepted answers ÷ all accepted answersMistakes among automatic suggestions
End-to-end coverageAccepted answers ÷ all incoming itemsHow much work the workflow actually handles
Per-class recall at policyCorrect accepted answers for a class ÷ all labelled items in that classThe class the workflow fails to serve
Per-class precision at policyCorrect accepted answers for a class ÷ all accepted predictions of that classThe trustworthiness of that suggested class
Failure rateItems without usable answers ÷ all incoming itemsService and integration losses that accuracy can hide

For Noul, report precision and recall at the action’s cutoff, the positive prevalence, and false-positive counts. For Score, compare error on the ordered rubric and the downstream ranking or threshold behaviour. With rare positives, inspect precision-recall performance alongside ROC-AUC. Confidence averages alone establish none of these.

Lab · synthetic records

The threshold changes who gets served

Twelve invented repair-desk records: ten answers and two failed calls. Move the cutoff or inspect a queue. Counts include failures in coverage and recall, while accuracy uses accepted answers only.

CaseLabelAnswerConfidencePolicy outcome

“Not estimable” means a zero denominator, not zero error. A class with no labelled examples supplies no recall evidence. “Unclear” is an answer; “no judgment” is a failure. This tiny synthetic set cannot establish a production cutoff.

Change one axis, then inspect paired cases

Keep the same items in each comparison. When a condition loses answers, compare accuracy on the commonly answered items and report the failures separately. Otherwise a variant can appear better merely by failing on difficult cases.

ExperimentKeep fixedInspect
Remove an exclusion or exampleState, model, labels, other wordingWhich boundary errors change; tokens and action coverage
Reverse Choice option orderOption meanings and inputLabel flips and cutoff crossings; keep production order fixed for calibration
Repeat identical requestsThe complete requestLabel agreement, probability shifts, action flips
Grow an item batchQuestion template and reference formPer-position and per-class errors, failures, usage, response time
Move definitions into shared stateIntended judgment and evaluation itemsProbability shifts and token savings; retune the changed question
Insert adversarial contentTrusted policy and ordinary itemsEffects on the item’s own answer and on neighbouring items

For batch experiments, first compare the original request with the batch template at size one. This catches wording or shared-definition changes before you attribute their effects to batch size. Rotate position and batch composition when testing the form you intend to ship.

Do not reorder Score levels as an option-order test: their order defines the scale. Do not interpret a repeated mistake as independent evidence that majority voting will repair it.

What we learned from Jev experiments

We evaluated jev-1.13.0 on 50 invented support tickets with hand-written labels. Most comparisons used three repeats on September 20, 2026; we ran shared-definition follow-ups on September 24. These experiments underpin this guide and our agent skill, v0.9.4. The methods, fixture, and tools live in our research repository.

We used 30 development and 20 validation items, with labels written by one researcher. This small fixture demonstrates the workflow; it is not an audited holdout. Some probe cases and raw run bodies are outside the distributed fixture, so the published materials do not provide a complete reproduction record for every figure. Use these results to design experiments on your own workload.

What we measuredWhat to test in your application
Missing evidence can look confident. Nine of ten deficient inputs had an arbitrary winner at confidence ≥ 0.95 without an escape option; all ten chose other at ≥ 0.97 when it was offered.Test empty, irrelevant, wrong-field, and incomplete state. Check the answer space before relying on confidence.
References mattered in long arrays. With 48 items, numeric references produced 23–24 correct queue labels; keyed references produced 47; quoted references produced 46. This comparison scores 48 items, not 50.Compare reference forms at intended batch sizes, with position and neighbour analysis. A path is a model-interpreted reference.
More questions differed from more items. Over the same 16-item keyed state, requests with 16, 48, or 240 questions scored 47–48 of 48 queue answers in each repeat.Separate question-count sweeps from state-growth sweeps. Neither study establishes unlimited free fan-out.
Shared definitions changed the signal. At eight items, pointer definitions saved 12% of input tokens for the Choice and a further 16% for the Nouls. Pointer wording raised the urgency Noul on 40–46 of 50 items; effects on false positives differed at cutoffs 0.5 and 0.8.Compare inline and shared definitions at size one, then at scale. Select thresholds for the resulting question, not for its predecessor.
Labels stayed still while numbers moved. Five identical runs over 50 items kept every label, but Choice probabilities moved by up to 0.08 and Noul outputs by up to 0.09.Measure flips of the application action as well as the winning label. Repeatability is separate from correctness.
A winner did not establish exclusive applicability. Choice assigned ≥ 0.99 to one option on 66% of items; separately worded Nouls found two applicable labels on 36%.Use Choice for competing alternatives and per-label Nouls for several applicable conditions. Evaluate them as distinct judgments.
Two injection probes spared keyed and quoted neighbours. Their queue labels remained at 28–29 correct of 29; numeric references worsened in those runs.Test broader attacks on target and neighbour answers. Naming and quoting are not an injection defense or authorization boundary.

Read the fixture notes for inputs, denominators, and reproduction commands, and our agent skill for probe details. The useful result is a set of experiments to repeat on your data. None of these counts establishes a universal batch limit, a security guarantee, or a production threshold.

Select a cutoff; validate the accepted region

Use tuning data to propose a rule. Check cumulative precision, coverage, per-class recall, and sample support at that rule. A perfect bucket containing two items does not establish a low error rate. Report appropriate uncertainty estimates and inspect the action-specific costs.

The cost calculation uses correctness probability under explicit assumptions. Returned confidence is a candidate routing signal, not that probability. A coarse fallback needs its own evaluation; a wrong fine label does not automatically make its parent right.

Once the frame and policy are fixed, evaluate them on untouched data. Recheck when the model, questions, candidate set, state construction, batch form, language, traffic prevalence, or action costs change. Keep results from different resolved models separate.

Use our evaluation tools with their limits in view

Our agent skill includes standard-library Python helpers for labelled evaluation, description ablation, batch sweeps, and token probes. Its wrapper fields such as reference, shared_state, and array_field configure those helpers; they are not fields in the TypeSafe HTTP request.

Read the evaluation protocol before running them. They call the live API and incur charges. The fixture’s reproduction commands do not create a production benchmark. Batch warning flags are diagnostic heuristics, not significance tests; the scorer does not fit a calibrator or implement your application’s review policy.

Real cost comes from reported usage across the whole workflow. Include retries and failures when usage is available, and record gaps when it is not. Request latency divided by batch size is amortized time, not the time each item waits. Raw evaluation output can contain source text; keep credentials and retained records protected.

Next, use the build guide to connect a reviewed frame to an application with an explicit failure path.

Compare providers on the decisions you will ship

Keep one source cohort and one label definition. Translate the frame into each supported interface, record the translation, and verify that every model received the same relevant evidence. A silently shortened state or option description makes this a different experiment.

  1. Fix the cohort. Include language, long inputs, missing facts, rare classes, and the action boundaries that matter. Keep tuning and final reporting separate for every model.
  2. Test the contract. Add an unrelated question, permute options, cross the input budget, simulate a failed response, and inspect the resolved model or checkpoint. Use the migration checklist.
  3. Tune each policy separately. Select confidence or relevant probabilities using that model’s tuning outcomes. Compare coverage at a common error budget, not performance at an inherited numerical cutoff.
  4. Measure the complete service. Include transport, cold starts, model loading, retries, concurrency, and review work. Self-hosted compute has a cost even when no token invoice arrives.
  5. Report the deployment you tested. Preserve weights, tokenizer, calibration, runtime, question schema, and truncation settings. Our Jev fixture findings are hypotheses to test elsewhere, not measurements of Clef, Kev, or Laya.

Provider benchmark tables can identify workloads to investigate. They do not establish a universal winner or a safe threshold for this application.