Try one useful integration
Build a small issue-triage assistant: read a report, propose labels, and defer unclear cases. Edit a request, inspect an example response, and practise evaluating the result before connecting it to real issues.
The project and its issues are fictional; the builder runs locally and sends nothing.
Choose the serving surface
The shared frame is portable; transport and operational settings need their own configuration. Pick a model in the builder below. It generates Jev SDK examples, a Workers AI binding for Clef, a Kev SDK client for a local server, or HTTP examples for Laya.
Models and providers covers the contracts. The response reader and detailed retry example later on this page remain explicitly Jev examples.
Start with Jev’s hosted API
Keep the TypeSafe key in your server environment. Choose an SDK or send HTTP requests directly. The snippets below show the starting point; follow the linked SDK reference for version-specific configuration.
pip install typesafe-sdk # or: uv add typesafe-sdk
export TYPESAFE_API_KEY=sk-... # from the TypeSafe consolenpm install @typesafe-ai/sdk # needs Node 20+
export TYPESAFE_API_KEY=sk-...export TYPESAFE_API_KEY=sk-...
# POST https://api.typesafe.ai/v1/systemone with Authorization: Bearer $TYPESAFE_API_KEYThe optional TypeSafe agent skill supplies API guidance to a coding assistant. You can use this tutorial without installing it.
Prepare a local decision service
The Kev and Laya builder examples assume a separately running service on the loopback address. These commands are for your environment; the guide does not run them. Model loading downloads weights and needs suitable hardware. Follow the project’s setup reference for runtime versions and serving configuration.
git clone https://github.com/jaredpalmer/kev.git
cd kev
uv sync --extra serve
uv run --extra serve python -m kev.serve --run jaredpalmer/kev-4b@v1.0 --port 8009The Hub revision pins the loaded weights. Record the serving-code revision and calibration settings too. See Kev setup.
pip install "laya[serve]"
LAYA_HOST=127.0.0.1 LAYA_PORT=8000 laya-serveThe builder selects english or multilingual explicitly. For an authenticated server, configure LAYA_API_KEY and give the client that value as DECISION_API_KEY. Record the installed library and model revisions. See Laya serving.
For hosted Clef, configure a Workers AI binding for the TypeScript example, or supply CLOUDFLARE_ACCOUNT_ID and CLOUDFLARE_AUTH_TOKEN for REST. See the model’s usage examples.
A Jev request, annotated
A fictional queue library receives a report that its consumer stops after a broker restart. We want to propose an issue type, identify missing information, and assess the described impact if it is a defect. Keep the first version modest: suggest labels for a maintainer to review. Sending comments or paging an engineer adds consequences that need their own acceptance rules and duplicate-action protection.
{
"state": {
"repo": "meridian/relay-queue",
"issue": {
"title": "Consumer stops pulling after broker restart",
"body": "After we restart the broker, the consumer logs reconnected once and then idles forever. Jobs pile up until the process is restarted. Happens on 3.2.0, worked on 3.1.x. No stack trace, CPU flat. Config attached.",
"author_association": "CONTRIBUTOR",
"comments": 0
}
},
"model": "jev-latest",
"questions": {
"kind": {
"type": "choice",
"instructions": "What kind of issue is `issue`, judged from its title and body?",
"criteria": {
"bug": "Behaviour that used to work or is documented does not behave as described",
"feature": "A request for behaviour the project does not claim to have",
"question": "The author is asking how to do something, not reporting a fault",
"docs": "The code is fine but the documentation is wrong or missing",
"none_of_these": "Spam, empty, or off-topic for this repository"
}
},
"needs_info": {
"type": "noul",
"instructions": "Would a maintainer have to ask the author for more before they could start work on `issue`?",
"criteria": {
"true": "Something needed to act is missing: a version, a reproduction, logs, or configuration",
"false": "The report contains enough to reproduce the problem or to decide what to do"
}
},
"severity": {
"type": "score",
"instructions": "If `issue` describes a defect, how bad is its impact on a user of the library?",
"criteria": [
"Cosmetic or a minor annoyance",
"A feature is degraded but a workaround exists",
"A feature is unusable and no workaround is described",
"Data loss, corruption, or a security exposure"
]
}
}
}
- The state is a JSON object, not a blob. Named fields let the instructions point at parts of it with backticked paths such as
issue. Facts code already knows, like the author's association, ride along as fields rather than being asked. - Jev IDs exist for your code. The docs are explicit that the id never reaches the model, so the meaning has to be complete inside
instructionsandcriteria. A key likekindtells the model nothing. - Each type has its own criteria shape. Choice takes a map of option to description, with
nullallowed when an option needs none. Score takes an ordered array of at least two level descriptions. Noul's criteria are optional and describe what yes and no mean. - There is always somewhere for the mass to go.
none_of_thesegives the Choice an exit. Without it, an off-topic issue is forced into the closest wrong label with an honest-looking probability. - The severity question is speculative. It only matters when
kindcomes back as a bug, but it is asked in the same request because the state is already there. See fan-out.
Request builder
Choose a model, then edit the state and questions; the JSON body, curl command, Python, and TypeScript update as you type. Prefilled with the triage example. Nothing is sent anywhere.Response reader
An illustrative Jev response to the triage request. Click a field to see what it means and the line of code that reads it.type. Choice and Score answers carry confidence, derived from the distribution; Noul answers carry only the probability. For Jev, output tokens are billed at zero under the documented pricing, so the number that matters for cost is input_tokens.Failures and service operations
Normalize the chosen service’s outer response, validate every expected answer, and retain missing or malformed answers as failures. Cloudflare REST wraps the model payload in result; Workers bindings return the model payload. Local servers have their own configured authentication, deadlines, and concurrency. The status table and SDK retry code below describe Jev.
Treat a failed request separately from an uncertain answer. A timeout gives you no judgment at all. Keep the issue in a review queue so the failure remains visible.
| Status | Documented meaning | What your code should do |
|---|---|---|
| 401 | Missing or invalid API key | Do not retry. Fix the Authorization header or the environment variable, and alert: this is a deploy problem, not a traffic problem. |
| 422 | Request failed validation; the body names the offending field | Do not retry. Inspect a sanitized diagnostic and fix the question shape. Common causes: a Score with one level, a Choice with no criteria map, a missing type. |
| 429 | Rate limit exceeded | Retry with exponential backoff and honor retry-after when present. Both SDKs do this by default. |
| 529 | Service temporarily overloaded | Same as 429. Cap the retries and fall through to your cascade's next stage rather than blocking the request path. |
Use bounded retries and a deadline covering attempts and waits. The current Python RetryPolicy reference defines its timeout as a total SDK-call retry budget. Connection failures, exhausted retries, and invalid successful responses all need a no-judgment path. Prefer the SDK for application integrations; our evaluation helpers have their own retry configuration.
Retries can repeat evaluation. Protect downstream actions with the issue ID and a processing version so a retry cannot post the same comment or apply the same operation twice.
Watch both request rate and token rate. Batching reduces repeated input and request count, but larger batches still consume tokens. Consult current limits when sizing traffic.
from typesafe_sdk import (
RetryPolicy, TypeSafeClient, TypeSafeAPIError,
TypeSafeAPIConnectionError, TypeSafeAPIResponseValidationError,
)
# QUESTIONS and build_state come from your reviewed request.
MODEL = "jev-1.13.0" # pin the version evaluated
ACCEPT_CONFIDENCE = 0.85 # illustrative; select on your labelled data
client = TypeSafeClient(retry=RetryPolicy(max_retries=2, timeout=3.0))
def propose_label(issue):
try:
result = client.system_one(
state=build_state(issue), questions=QUESTIONS, model=MODEL
)
except (TypeSafeAPIConnectionError, TypeSafeAPIResponseValidationError):
return {"status": "no_judgment", "next": "manual_triage"}
except TypeSafeAPIError as error:
if error.status not in (408, 429) and error.status < 500:
raise # request or configuration fault; surface it
return {"status": "no_judgment", "next": "manual_triage"}
answer = result.choices["kind"]
if answer.choice == "none_of_these":
return {"status": "no_match", "next": "manual_triage"}
if answer.confidence < ACCEPT_CONFIDENCE:
return {"status": "review", "answer": answer.choice}
return {"status": "suggestion", "answer": answer.choice}
# A suggestion is data. Authorize and deduplicate any later write in code.
Evaluate the suggested action before enabling writes
Replay historical issues with the reviewed frame and action policy. Report accepted accuracy, coverage over all issues, class recall, no-match cases, and failures. Make sure historical labels do not depend on evidence your request omits.
Use the evaluation protocol to separate design, threshold selection, and final reporting. The small sheet below explains cumulative counts; it is a visual aid for proposing a cutoff, not evidence that the cutoff is safe.
Eval sheet
Forty synthetic labelled issues with the model'skind answer and confidence. Change the band width and the target accuracy to see where an "act" threshold would land. The sample is generated in the page and is illustrative only.
| Confidence band | n | Accuracy | Acc. above | Coverage above |
|---|
| Issue | Label | Answer | Conf. |
|---|
Keep the production contract explicit
- Keep credentials server-side. Sanitize diagnostics. Raw responses and evaluation records can contain source text; key redaction does not remove all sensitive content.
- Version the inputs. Record the question version, state-building version, policy version, and service’s reported
model, plus the loaded weights revision, tokenizer, calibration, and runtime for local models. An echoed alias is not a weights pin. Pin the tested deployment after selecting thresholds; evaluate replacements before switching. - Cache the complete judgment input. Include a digest of the actual evidence and relevant policy, roster, or related records, together with question ID/version and resolved model version. Include all jointly scored questions for Clef, not only the answer being read. An unchanged issue ID does not mean unchanged content.
- Handle aliases deliberately. A cache keyed only by
jev-latestcannot distinguish model releases. Resolve/version entries before reuse; do not combine multiple versions into one evaluation claim. - Make effects idempotent. Use a separate operation key for applying a label, posting a comment, or sending an alert. Retrying inference must not duplicate downstream work.
- Check freshness and authorization. Before writing, confirm the issue and relevant policy still match the evaluated state. A typed suggestion or high confidence does not authorize an operation.
- Retain a visible fallback. Distinguish no match, low acceptance, and no judgment. Collect corrections with suitable access and retention so changes can be evaluated.
The production contract in our agent skill explains these boundaries. Use the official SDK reference for installed-version behaviour.