Part 5 of 12 · Frame engineering

Design the decision frame

The model reads the request you give it. Frame engineering makes that request answerable, tests the distinction it draws, and checks what your application does with the answer.

A repair-desk example connects the concepts. The lab uses invented answers, including deliberately confident errors.

Three parts of a frame; one policy around it

A decision frame brings together the evidence in state, the judgment in instructions, and the answer boundaries in criteria where applicable. Frame engineering is the design, testing, and refinement of those parts. The application’s action rule sits around the frame and needs evaluation too.

01 · Evidence

What can it read?

The customer’s message and a relevant record of the last visit.

02 · Judgment

What distinction matters?

Which workshop queue best fits this request?

03 · Boundaries

Where can the answer land?

Repair, fitting, accessories, or insufficient information.

04 · Application policy

What happens next?

Suggest a queue, ask for context, or send the case for review.

Version the frame and the policy separately. A new threshold changes behaviour without changing inference. A new criterion changes what the question means and can invalidate the old threshold.

A better question cannot supply a missing fact

“It still feels wrong after last time” refers to evidence outside the message. The last visit might have been a wheel repair or a fitting. A more emphatic routing question cannot establish which one happened.

Retrieve the relevant record when you can. Otherwise, keep an answer for insufficient information. Confidence summarizes the distribution over the options offered; it cannot guarantee that the right option was offered or that the needed evidence exists.

Our measurement · Jev 1.13.0

We tested ten empty, irrelevant, or wrong-field inputs. Without a no-match option, nine produced an arbitrary winner with confidence at least 0.95. With other available, all ten selected it with confidence at least 0.97. This small probe motivates an input-coverage test; it does not promise that adding other always fixes missing evidence. See the study notes.

Lab · synthetic answers

Change evidence, boundaries, or policy

Try an incomplete message with and without history or an escape option. A service failure has no answer to threshold. Every response here is hand-authored.

Evidence sent
Question and boundaries
Invented response

The threshold changes only the rule. Changing history or the escape option selects a different illustrative response. High confidence can still route a message using missing evidence. These values are not copied from the study.

Split decisions when you need to use them separately

The desk can route a message, estimate the disruption to a ride, and check whether the customer requested a loan bike. Each answer has its own meaning and can be tested alone. One broad “what should we do?” question hides those distinctions.

Keep a relational judgment intact when the relationship is the point. “Does this message address one of this customer’s open bookings?” needs both the message and the bookings. Atomic means one coherent decision, not context-free fact extraction.

Independent questions can share one request. If a first answer determines which booking record to fetch, fetch it in code and build a later request. A multi-question request does not supply a returned answer as new state. Joint scoring and a sequential workflow are different contracts.

Choose a primitive by the decision: Choice for one competing option, Noul for whether each condition applies, Score for a described degree. Multiple Nouls can be positive together; a Choice still selects one winner.

Keep four outcomes distinguishable

OutcomeWhat you haveUseful next step
A usable answerA judgment that passes the tested action ruleSuggest the queue; check permission before any consequential operation
No matchA real answer such as unclearAsk for the missing context or handle an out-of-scope request
Below the acceptance thresholdA real answer the policy does not acceptReview, clarify, or use another validated stage
No judgmentA failed call or an unusable responseRecord the failure; retry within a deadline or use the fallback

A failure has no model probability. Replacing it with false or other corrupts the outcome record and makes service problems look like successful judgments.

Improve the frame with a controlled experiment

Start with a plain question and a sufficient answer space. Inspect mistakes before adding more instructions. Missing evidence, unclear labels, a model mistake, and an application bug call for different changes.

  1. Freeze a baseline: evidence construction, questions, model, reference form, and policy.
  2. Choose one change that addresses an observed failure, such as adding the previous visit.
  3. Compare the same labelled items, retaining failed calls and high-confidence mistakes.
  4. Check whether relevant facts change the answer and harmless paraphrases preserve the decision.
  5. Accept the change for its measured effect on errors, coverage, cost, and latency.

The next chapter develops instructions and boundaries. Evaluation shows how to separate tuning from final evidence and interpret our findings.

Keep the method; verify the implementation

Evidence coverage, clear boundaries, an insufficient-information outcome, and a tested action rule remain useful across decision models. How a runtime serializes that evidence, isolates questions, shortens inputs, and reports confidence can differ. Record those choices as part of the frame. The model chapter maps the contracts; the comparison protocol tests the application on each one.