Jev Decision Model: Branching Design and Limitations
Jev is TypeSafe AI's decision model, which returns predefined choices and probabilities. This overview explains how to connect each question type to code, how type inference works, how to interpret confidence, and what to verify when using Jev with Korean.
Jev returns choices and probabilities that code can use instead of free-form text.
Choice is for choosing a category, Score is for evaluating on a scale, and Noul asks for the probability that something is true.
The TypeScript SDK infers the answer type from the question definition.
confidence summarizes the probability distribution and is not itself the accuracy rate.
Korean-language services must use their own data to evaluate both the automated processing rate and the rate of incorrect decisions.
Jev is TypeSafe AI's decision model that returns predefined choices and probabilities. It is suited to connecting natural-language judgments to branching conditions in code. It restricts the output format but does not guarantee the accuracy of its judgments.
Pricing is based on the announcement of September 15, 2026, and limitations are based on documents reviewed on September 17.
What does Jev do?
Jev reads input and returns structured answers to predefined questions. Put the content to evaluate in state. Define the criteria for judgment in questions. The returned answers can be used for classification or prioritization.
TypeSafe AI calls this approach a System One model. It describes the training method as RLCD, a form of reinforcement learning focused on calibrating judgment probabilities. Accuracy for each service still needs to be verified separately.
The official Introduction document states the design principle this way:
Atomic questions, composed in code TypeSafe AI, Introduction
This means assigning one judgment to each question. Handle combinations of judgments in code. Questions in a single request evaluate the same state independently. See the TypeSafe AI Introduction for details of the structure.
Comparing Choice, Score, and Noul
Choose a question type based on how you will use the returned answer in code. Identifying the responsible department is a category-selection problem. Assessing severity is an ordered evaluation. Whether a particular request is present can be divided into true and false.
Type | What it asks | Main return values | Code use
Choice | Which predefined category applies? | choice, probabilities, confidence | switch branch
Score | Which of the described levels applies? | score, probabilities, confidence | Compare score with a threshold
Noul | Is a particular condition true? | noul | Compare the probability, then use an if branch
How should you define the choices for Choice?
Choice accepts up to 255 choices per question. It evaluates both the name and description of each choice. If inputs outside the list may occur, consider an option such as other. This helps avoid forcing inputs into unsuitable categories.
Choice selects one of the given candidates. To judge separately whether all candidates are unsuitable, you need to add a question. The returned probabilities sum to 1. The definitions and limits are in the Choice documentation.
How are Score and Noul different?
Score evaluates where an input falls among described levels. Level numbers start at 0. The result is a probability-weighted average of the level numbers, so it can return a non-integer score.
Noul returns the probability that a condition is true, from 0 to 1. A value near 0.5 means true and false have similar probabilities. It should not be read as a moderate degree of dissatisfaction. Use Score to measure degree and Noul to ask whether something is present.
Comparing its role with conventional code and generative models
Leave exact calculations to code and assign only semantic judgments to the model. Writing new sentences is a job for a generative model. Jev is suited to judgments with a defined range of answers. This distinction can guide which parts of an existing feature to replace.
Task condition | Suitable approach | Reason
Calculate the interval between dates or count items | Conventional code | Can be calculated exactly according to rules
Classify a customer inquiry by responsible department | Jev Choice | Semantic judgment with predefined choices
Assess the severity of an incident report | Jev Score | Judgment among described levels
Check for a request to speak with an agent | Jev Noul | Judgment about whether a specific intent is present
Write a response or an open-ended summary | Generative model | Requires generating new text
Code to validate the format may be needed when processing text output. Jev instead returns answers directly in a defined format. That does not eliminate all validation and retries. The official SDK also provides a retry policy for communication errors and other issues.
Continue to design external input validation and business-rule checks separately. A model answer can match the required type and still be wrong for the business task. This distinction is a design interpretation based on the official Client SDKs and model limitations document.
TypeScript question definitions and answer types
The TypeScript SDK infers return types from question definitions. Choice option keys become the allowed values for the answer. This lets you catch comparisons against disallowed strings at compile time. The SDK requires Node.js 20 or later.
The following illustrative code adapts the official SDK call format. It is not an example that measures results or accuracy. The TYPESAFE_API_KEY environment variable must be set.
import { TypeSafeClient, choice } from "@typesafe-ai/sdk";
const api = new TypeSafeClient();
async function classifyMessage(message: string) {
const result = await api.systemOne({
model: "jev-1.13.0",
state: { message },
questions: {
topic: choice("Classify the subject of message.", {
account: "Account access or password problems",
delivery: "Shipment tracking or delivery problems",
other: "Any subject outside those categories",
}),
},
});
return result.answers.topic;
}
The choice type in this example is account | delivery | other. In actual TypeScript, these strings are literal types shown in quotation marks. Type checking does not, however, determine whether the inquiry was classified correctly. Check installation and call specifications against the JavaScript SDK documentation.
A confidence calculation example and common mistakes
confidence is a summary of the returned probabilities. It is not the same number as the probability of the selected answer. It should not be read directly as the actual accuracy rate, either. This distinction is especially important when setting criteria for automatic execution.
The official formula for Choice is:
confidence = (maximum probability - 1 / number of choices) / (1 - 1 / number of choices)
The probability distribution in the official documentation is 0.6, 0.3, 0.1. There are 3 choices. Substituting the maximum probability of 0.6 gives a confidence of 0.4. A selection probability of 60% and a confidence of 0.4 are different measures.
Common interpretation | Correct interpretation
A confidence of 0.4 means 40% accuracy | A value summarizing the probability distribution according to the formula
Noul 0.5 means a moderate level | True and false have been assigned similar probabilities
Score's decimal places are precise measurements | A probability-weighted average of the defined level numbers
High confidence means permission checks can be skipped | Check business permissions and execution conditions separately in code
The meanings of these numbers are based on the official Confidence documentation. Separating permission checks is a design recommendation for applying that information in production.
How should you compare pricing and response times?
The announced price is $0.042 per 1 million input tokens. Output tokens were described as free. The announced response-time range is 70 to 500 milliseconds. These figures are based on the company's September 15, 2026 announcement.
The company's multiplier comparisons came from an evaluation of a specific workflow. They used the average predictions of other large models as the reference instead of actual correct answers. Speed testing was conducted mainly in the western United States. The results therefore do not represent accuracy across all tasks or response times in Korea.
Comparison item | What to check
API cost | Actual input-token usage and applicable rate
Response time | Round-trip time in the region where the service runs
Accuracy | Results from samples labeled with correct answers for the actual task
Operating cost | Cost including retries and human review
Comparing prices alone can miss an increase in review work. An adoption decision needs to account for costs through to the final processing result. The conditions behind the announced figures are stated in the Jev public announcement.
Nine limitations of Jev 1.13
The official documentation groups Jev 1.13's failure modes into nine types. This list is based on documents reviewed on September 17, 2026. Success on some examples does not guarantee reliability for that task.
Failure mode | Design response
Interpreting wording literally | State hidden conditions explicitly in the instructions
Calculation and counting | Handle arithmetic in code
Comparing dates and times | Construct dates, then compare them in code
Double negatives and multistep reasoning | Split them into direct questions
Inputs with much irrelevant content | Pass only the necessary fields
Inputs that steer the judgment | Include leading wording in advance evaluations
Conflicts between instructions and selection criteria | Align the meanings of questions and criteria
Lack of mathematical consistency between questions | Manage logical relationships in code
Open-ended text generation | Use a generative model
Noul and Choice can assign different probabilities to the same meaning. Do not apply a threshold tuned for one type to the other unchanged. The source is Jev 1.13 jaggedness.
Conditions for use in Korean-language services
Jev processes Korean input, but performance equivalent to English is not guaranteed. The official documentation says its primary training language is English. It also notes performance differences for CJK scripts, including Korean. It provides no general accuracy figure for Korean.
Use your own samples to decide whether to use it for Korean. It is worth including inquiries with indirect wording and omissions. Evaluating only translated text may miss differences in how actual customers write. This is an evaluation recommendation based on language-specific performance differences.
· Evaluate clear and ambiguous requests separately
· Report department classification and sentiment assessment separately
· Recheck in Korean any thresholds set using English
· Compare the same samples before and after a model change
jev-latest points to a different model when a new version is released. If you tuned thresholds to a specific version, consider pinning that version. Check the official Models documentation for language and version policies.
The automation rate that is easy to miss when evaluating a model replacement
When replacing a model, consider both overall accuracy and the automation rate. Sending every ambiguous case to a person reduces the amount handled automatically. Expanding the scope of automatic processing can increase incorrect executions. The following evaluation method extends the confidence-based branching described in the official documentation.
Metric | How to calculate or record it | What it checks
Automation rate | Automatically processed cases / all evaluated cases | Manual work actually reduced
Error rate in automatic processing | Incorrect automatically processed cases / automatically processed cases | Quality of automated results
Human review rate | Cases sent for review / all evaluated cases | Remaining review workload
Final processing time | Measure from input to final completion | Waiting time including review
If no cases were processed automatically, the error rate in automatic processing cannot be calculated. Record it as not calculable, rather than 0%. Keep the denominator for each metric so models can be compared.
You can evaluate in this order:
· Have people label real work samples with the correct answers.
· Fix the question wording and choice definitions.
· Record the model version, probabilities, and final branch.
· Compare automation rates and error rates at different thresholds.
· Set operating criteria according to the level of error you can accept.
There is no universally correct threshold. The criteria must reflect the impact of an incorrect execution. The branching principles that provide a starting point are in the Confidence documentation.