{"content_id":"klwm55imio","slug":"jev-decision-model-branching-design-and-limitations","locale":"en","schema_type":"TechArticle","category":"knowledge_base","category_name":"Knowledge Base","title":"Jev Decision Model: Branching Design and Limitations","summary":"Jev is TypeSafe AI's decision model, which returns predefined choices and probabilities. This overview explains how to connect each question type to code, how type inference works, how to interpret confidence, and what to verify when using Jev with Korean.","sponsorship_disclosure":null,"affiliate_disclosure":null,"commerce_disclosure":null,"author":{"name":"Injoys Editorial Team","url":"https://injoys.com/ko/about"},"key_points":["Jev returns choices and probabilities that code can use instead of free-form text.","Choice is for choosing a category, Score is for evaluating on a scale, and Noul asks for the probability that something is true.","The TypeScript SDK infers the answer type from the question definition.","confidence summarizes the probability distribution and is not itself the accuracy rate.","Korean-language services must use their own data to evaluate both the automated processing rate and the rate of incorrect decisions."],"content_markdown":"Jev is TypeSafe AI's decision model that returns predefined choices and probabilities. It is suited to connecting natural-language judgments to branching conditions in code. It restricts the output format but does not guarantee the accuracy of its judgments.\n\nPricing is based on the announcement of September 15, 2026, and limitations are based on documents reviewed on September 17.\n\n## What does Jev do?\n\nJev reads input and returns structured answers to predefined questions. Put the content to evaluate in `state`. Define the criteria for judgment in `questions`. The returned answers can be used for classification or prioritization.\n\nTypeSafe AI calls this approach a System One model. It describes the training method as RLCD, a form of reinforcement learning focused on calibrating judgment probabilities. Accuracy for each service still needs to be verified separately.\n\nThe official Introduction document states the design principle this way:\n\n\u003e Atomic questions, composed in code\n\u003e\n\u003e TypeSafe AI, Introduction\n\nThis means assigning one judgment to each question. Handle combinations of judgments in code. Questions in a single request evaluate the same state independently. See the [TypeSafe AI Introduction](https://docs.typesafe.ai/introduction) for details of the structure.\n\n## Comparing Choice, Score, and Noul\n\nChoose a question type based on how you will use the returned answer in code. Identifying the responsible department is a category-selection problem. Assessing severity is an ordered evaluation. Whether a particular request is present can be divided into true and false.\n\n| Type | What it asks | Main return values | Code use |\n| --- | --- | --- | --- |\n| Choice | Which predefined category applies? | `choice`, `probabilities`, `confidence` | `switch` branch |\n| Score | Which of the described levels applies? | `score`, `probabilities`, `confidence` | Compare score with a threshold |\n| Noul | Is a particular condition true? | `noul` | Compare the probability, then use an `if` branch |\n\n### How should you define the choices for Choice?\n\nChoice accepts up to 255 choices per question. It evaluates both the name and description of each choice. If inputs outside the list may occur, consider an option such as `other`. This helps avoid forcing inputs into unsuitable categories.\n\nChoice selects one of the given candidates. To judge separately whether all candidates are unsuitable, you need to add a question. The returned probabilities sum to 1. The definitions and limits are in the [Choice documentation](https://docs.typesafe.ai/primitives/choice).\n\n### How are Score and Noul different?\n\nScore evaluates where an input falls among described levels. Level numbers start at 0. The result is a probability-weighted average of the level numbers, so it can return a non-integer score.\n\nNoul returns the probability that a condition is true, from 0 to 1. A value near 0.5 means true and false have similar probabilities. It should not be read as a moderate degree of dissatisfaction. Use [Score](https://docs.typesafe.ai/primitives/score) to measure degree and [Noul](https://docs.typesafe.ai/primitives/noul) to ask whether something is present.\n\n## Comparing its role with conventional code and generative models\n\nLeave exact calculations to code and assign only semantic judgments to the model. Writing new sentences is a job for a generative model. Jev is suited to judgments with a defined range of answers. This distinction can guide which parts of an existing feature to replace.\n\n| Task condition | Suitable approach | Reason |\n| --- | --- | --- |\n| Calculate the interval between dates or count items | Conventional code | Can be calculated exactly according to rules |\n| Classify a customer inquiry by responsible department | Jev Choice | Semantic judgment with predefined choices |\n| Assess the severity of an incident report | Jev Score | Judgment among described levels |\n| Check for a request to speak with an agent | Jev Noul | Judgment about whether a specific intent is present |\n| Write a response or an open-ended summary | Generative model | Requires generating new text |\n\nCode to validate the format may be needed when processing text output. Jev instead returns answers directly in a defined format. That does not eliminate all validation and retries. The official SDK also provides a retry policy for communication errors and other issues.\n\nContinue to design external input validation and business-rule checks separately. A model answer can match the required type and still be wrong for the business task. This distinction is a design interpretation based on the official [Client SDKs](https://docs.typesafe.ai/sdk) and [model limitations document](https://docs.typesafe.ai/model-jaggedness/jev-1.13).\n\n## TypeScript question definitions and answer types\n\nThe TypeScript SDK infers return types from question definitions. Choice option keys become the allowed values for the answer. This lets you catch comparisons against disallowed strings at compile time. The SDK requires Node.js 20 or later.\n\nThe following illustrative code adapts the official SDK call format. It is not an example that measures results or accuracy. The `TYPESAFE_API_KEY` environment variable must be set.\n\n```typescript\nimport { TypeSafeClient, choice } from \"@typesafe-ai/sdk\";\n\nconst api = new TypeSafeClient();\n\nasync function classifyMessage(message: string) {\n  const result = await api.systemOne({\n    model: \"jev-1.13.0\",\n    state: { message },\n    questions: {\n      topic: choice(\"Classify the subject of message.\", {\n        account: \"Account access or password problems\",\n        delivery: \"Shipment tracking or delivery problems\",\n        other: \"Any subject outside those categories\",\n      }),\n    },\n  });\n\n  return result.answers.topic;\n}\n```\n\nThe `choice` type in this example is `account | delivery | other`. In actual TypeScript, these strings are literal types shown in quotation marks. Type checking does not, however, determine whether the inquiry was classified correctly. Check installation and call specifications against the [JavaScript SDK documentation](https://docs.typesafe.ai/sdk/javascript).\n\n## A confidence calculation example and common mistakes\n\n`confidence` is a summary of the returned probabilities. It is not the same number as the probability of the selected answer. It should not be read directly as the actual accuracy rate, either. This distinction is especially important when setting criteria for automatic execution.\n\nThe official formula for Choice is:\n\n`confidence = (maximum probability - 1 / number of choices) / (1 - 1 / number of choices)`\n\nThe probability distribution in the official documentation is `0.6, 0.3, 0.1`. There are 3 choices. Substituting the maximum probability of 0.6 gives a confidence of 0.4. A selection probability of 60% and a confidence of 0.4 are different measures.\n\n| Common interpretation | Correct interpretation |\n| --- | --- |\n| A confidence of 0.4 means 40% accuracy | A value summarizing the probability distribution according to the formula |\n| Noul 0.5 means a moderate level | True and false have been assigned similar probabilities |\n| Score's decimal places are precise measurements | A probability-weighted average of the defined level numbers |\n| High confidence means permission checks can be skipped | Check business permissions and execution conditions separately in code |\n\nThe meanings of these numbers are based on the official [Confidence documentation](https://docs.typesafe.ai/confidence). Separating permission checks is a design recommendation for applying that information in production.\n\n## How should you compare pricing and response times?\n\nThe announced price is $0.042 per 1 million input tokens. Output tokens were described as free. The announced response-time range is 70 to 500 milliseconds. These figures are based on the company's September 15, 2026 announcement.\n\nThe company's multiplier comparisons came from an evaluation of a specific workflow. They used the average predictions of other large models as the reference instead of actual correct answers. Speed testing was conducted mainly in the western United States. The results therefore do not represent accuracy across all tasks or response times in Korea.\n\n| Comparison item | What to check |\n| --- | --- |\n| API cost | Actual input-token usage and applicable rate |\n| Response time | Round-trip time in the region where the service runs |\n| Accuracy | Results from samples labeled with correct answers for the actual task |\n| Operating cost | Cost including retries and human review |\n\nComparing prices alone can miss an increase in review work. An adoption decision needs to account for costs through to the final processing result. The conditions behind the announced figures are stated in the [Jev public announcement](https://typesafe.ai/blog/introducing-system-one-models-and-jev).\n\n## Nine limitations of Jev 1.13\n\nThe official documentation groups Jev 1.13's failure modes into nine types. This list is based on documents reviewed on September 17, 2026. Success on some examples does not guarantee reliability for that task.\n\n| Failure mode | Design response |\n| --- | --- |\n| Interpreting wording literally | State hidden conditions explicitly in the instructions |\n| Calculation and counting | Handle arithmetic in code |\n| Comparing dates and times | Construct dates, then compare them in code |\n| Double negatives and multistep reasoning | Split them into direct questions |\n| Inputs with much irrelevant content | Pass only the necessary fields |\n| Inputs that steer the judgment | Include leading wording in advance evaluations |\n| Conflicts between instructions and selection criteria | Align the meanings of questions and criteria |\n| Lack of mathematical consistency between questions | Manage logical relationships in code |\n| Open-ended text generation | Use a generative model |\n\nNoul and Choice can assign different probabilities to the same meaning. Do not apply a threshold tuned for one type to the other unchanged. The source is [Jev 1.13 jaggedness](https://docs.typesafe.ai/model-jaggedness/jev-1.13).\n\n## Conditions for use in Korean-language services\n\nJev processes Korean input, but performance equivalent to English is not guaranteed. The official documentation says its primary training language is English. It also notes performance differences for CJK scripts, including Korean. It provides no general accuracy figure for Korean.\n\nUse your own samples to decide whether to use it for Korean. It is worth including inquiries with indirect wording and omissions. Evaluating only translated text may miss differences in how actual customers write. This is an evaluation recommendation based on language-specific performance differences.\n\n- Evaluate clear and ambiguous requests separately\n- Report department classification and sentiment assessment separately\n- Recheck in Korean any thresholds set using English\n- Compare the same samples before and after a model change\n\n`jev-latest` points to a different model when a new version is released. If you tuned thresholds to a specific version, consider pinning that version. Check the official [Models documentation](https://docs.typesafe.ai/models) for language and version policies.\n\n## The automation rate that is easy to miss when evaluating a model replacement\n\nWhen replacing a model, consider both overall accuracy and the automation rate. Sending every ambiguous case to a person reduces the amount handled automatically. Expanding the scope of automatic processing can increase incorrect executions. The following evaluation method extends the confidence-based branching described in the official documentation.\n\n| Metric | How to calculate or record it | What it checks |\n| --- | --- | --- |\n| Automation rate | Automatically processed cases / all evaluated cases | Manual work actually reduced |\n| Error rate in automatic processing | Incorrect automatically processed cases / automatically processed cases | Quality of automated results |\n| Human review rate | Cases sent for review / all evaluated cases | Remaining review workload |\n| Final processing time | Measure from input to final completion | Waiting time including review |\n\nIf no cases were processed automatically, the error rate in automatic processing cannot be calculated. Record it as not calculable, rather than 0%. Keep the denominator for each metric so models can be compared.\n\nYou can evaluate in this order:\n\n1. Have people label real work samples with the correct answers.\n2. Fix the question wording and choice definitions.\n3. Record the model version, probabilities, and final branch.\n4. Compare automation rates and error rates at different thresholds.\n5. Set operating criteria according to the level of error you can accept.\n\nThere is no universally correct threshold. The criteria must reflect the impact of an incorrect execution. The branching principles that provide a starting point are in the [Confidence documentation](https://docs.typesafe.ai/confidence).","content_html":"\u003cp\u003eJev is TypeSafe AI's decision model that returns predefined choices and probabilities. It is suited to connecting natural-language judgments to branching conditions in code. It restricts the output format but does not guarantee the accuracy of its judgments.\u003c/p\u003e\n\u003cp\u003ePricing is based on the announcement of September 15, 2026, and limitations are based on documents reviewed on September 17.\u003c/p\u003e\n\u003ch2\u003e\n\u003ca href=\"#what-does-jev-do\" class=\"anchor\" id=\"what-does-jev-do\"\u003e\u003c/a\u003eWhat does Jev do?\u003c/h2\u003e\n\u003cp\u003eJev reads input and returns structured answers to predefined questions. Put the content to evaluate in \u003ccode\u003estate\u003c/code\u003e. Define the criteria for judgment in \u003ccode\u003equestions\u003c/code\u003e. The returned answers can be used for classification or prioritization.\u003c/p\u003e\n\u003cp\u003eTypeSafe AI calls this approach a System One model. It describes the training method as RLCD, a form of reinforcement learning focused on calibrating judgment probabilities. Accuracy for each service still needs to be verified separately.\u003c/p\u003e\n\u003cp\u003eThe official Introduction document states the design principle this way:\u003c/p\u003e\n\u003cblockquote\u003e\n\u003cp\u003eAtomic questions, composed in code\u003c/p\u003e\n\u003cp\u003eTypeSafe AI, Introduction\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003cp\u003eThis means assigning one judgment to each question. Handle combinations of judgments in code. Questions in a single request evaluate the same state independently. See the \u003ca href=\"https://docs.typesafe.ai/introduction\"\u003eTypeSafe AI Introduction\u003c/a\u003e for details of the structure.\u003c/p\u003e\n\u003ch2\u003e\n\u003ca href=\"#comparing-choice-score-and-noul\" class=\"anchor\" id=\"comparing-choice-score-and-noul\"\u003e\u003c/a\u003eComparing Choice, Score, and Noul\u003c/h2\u003e\n\u003cp\u003eChoose a question type based on how you will use the returned answer in code. Identifying the responsible department is a category-selection problem. Assessing severity is an ordered evaluation. Whether a particular request is present can be divided into true and false.\u003c/p\u003e\n\u003cdiv class=\"overflow-x-auto\"\u003e\u003ctable\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth\u003eType\u003c/th\u003e\n\u003cth\u003eWhat it asks\u003c/th\u003e\n\u003cth\u003eMain return values\u003c/th\u003e\n\u003cth\u003eCode use\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Type\"\u003eChoice\u003c/td\u003e\n\u003ctd data-label=\"What it asks\"\u003eWhich predefined category applies?\u003c/td\u003e\n\u003ctd data-label=\"Main return values\"\u003e\n\u003ccode\u003echoice\u003c/code\u003e, \u003ccode\u003eprobabilities\u003c/code\u003e, \u003ccode\u003econfidence\u003c/code\u003e\n\u003c/td\u003e\n\u003ctd data-label=\"Code use\"\u003e\n\u003ccode\u003eswitch\u003c/code\u003e branch\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Type\"\u003eScore\u003c/td\u003e\n\u003ctd data-label=\"What it asks\"\u003eWhich of the described levels applies?\u003c/td\u003e\n\u003ctd data-label=\"Main return values\"\u003e\n\u003ccode\u003escore\u003c/code\u003e, \u003ccode\u003eprobabilities\u003c/code\u003e, \u003ccode\u003econfidence\u003c/code\u003e\n\u003c/td\u003e\n\u003ctd data-label=\"Code use\"\u003eCompare score with a threshold\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Type\"\u003eNoul\u003c/td\u003e\n\u003ctd data-label=\"What it asks\"\u003eIs a particular condition true?\u003c/td\u003e\n\u003ctd data-label=\"Main return values\"\u003e\u003ccode\u003enoul\u003c/code\u003e\u003c/td\u003e\n\u003ctd data-label=\"Code use\"\u003eCompare the probability, then use an \u003ccode\u003eif\u003c/code\u003e branch\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\u003c/div\u003e\n\u003ch3\u003e\n\u003ca href=\"#how-should-you-define-the-choices-for-choice\" class=\"anchor\" id=\"how-should-you-define-the-choices-for-choice\"\u003e\u003c/a\u003eHow should you define the choices for Choice?\u003c/h3\u003e\n\u003cp\u003eChoice accepts up to 255 choices per question. It evaluates both the name and description of each choice. If inputs outside the list may occur, consider an option such as \u003ccode\u003eother\u003c/code\u003e. This helps avoid forcing inputs into unsuitable categories.\u003c/p\u003e\n\u003cp\u003eChoice selects one of the given candidates. To judge separately whether all candidates are unsuitable, you need to add a question. The returned probabilities sum to 1. The definitions and limits are in the \u003ca href=\"https://docs.typesafe.ai/primitives/choice\"\u003eChoice documentation\u003c/a\u003e.\u003c/p\u003e\n\u003ch3\u003e\n\u003ca href=\"#how-are-score-and-noul-different\" class=\"anchor\" id=\"how-are-score-and-noul-different\"\u003e\u003c/a\u003eHow are Score and Noul different?\u003c/h3\u003e\n\u003cp\u003eScore evaluates where an input falls among described levels. Level numbers start at 0. The result is a probability-weighted average of the level numbers, so it can return a non-integer score.\u003c/p\u003e\n\u003cp\u003eNoul returns the probability that a condition is true, from 0 to 1. A value near 0.5 means true and false have similar probabilities. It should not be read as a moderate degree of dissatisfaction. Use \u003ca href=\"https://docs.typesafe.ai/primitives/score\"\u003eScore\u003c/a\u003e to measure degree and \u003ca href=\"https://docs.typesafe.ai/primitives/noul\"\u003eNoul\u003c/a\u003e to ask whether something is present.\u003c/p\u003e\n\u003ch2\u003e\n\u003ca href=\"#comparing-its-role-with-conventional-code-and-generative-models\" class=\"anchor\" id=\"comparing-its-role-with-conventional-code-and-generative-models\"\u003e\u003c/a\u003eComparing its role with conventional code and generative models\u003c/h2\u003e\n\u003cp\u003eLeave exact calculations to code and assign only semantic judgments to the model. Writing new sentences is a job for a generative model. Jev is suited to judgments with a defined range of answers. This distinction can guide which parts of an existing feature to replace.\u003c/p\u003e\n\u003cdiv class=\"overflow-x-auto\"\u003e\u003ctable\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth\u003eTask condition\u003c/th\u003e\n\u003cth\u003eSuitable approach\u003c/th\u003e\n\u003cth\u003eReason\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Task condition\"\u003eCalculate the interval between dates or count items\u003c/td\u003e\n\u003ctd data-label=\"Suitable approach\"\u003eConventional code\u003c/td\u003e\n\u003ctd data-label=\"Reason\"\u003eCan be calculated exactly according to rules\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Task condition\"\u003eClassify a customer inquiry by responsible department\u003c/td\u003e\n\u003ctd data-label=\"Suitable approach\"\u003eJev Choice\u003c/td\u003e\n\u003ctd data-label=\"Reason\"\u003eSemantic judgment with predefined choices\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Task condition\"\u003eAssess the severity of an incident report\u003c/td\u003e\n\u003ctd data-label=\"Suitable approach\"\u003eJev Score\u003c/td\u003e\n\u003ctd data-label=\"Reason\"\u003eJudgment among described levels\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Task condition\"\u003eCheck for a request to speak with an agent\u003c/td\u003e\n\u003ctd data-label=\"Suitable approach\"\u003eJev Noul\u003c/td\u003e\n\u003ctd data-label=\"Reason\"\u003eJudgment about whether a specific intent is present\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Task condition\"\u003eWrite a response or an open-ended summary\u003c/td\u003e\n\u003ctd data-label=\"Suitable approach\"\u003eGenerative model\u003c/td\u003e\n\u003ctd data-label=\"Reason\"\u003eRequires generating new text\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\u003c/div\u003e\n\u003cp\u003eCode to validate the format may be needed when processing text output. Jev instead returns answers directly in a defined format. That does not eliminate all validation and retries. The official SDK also provides a retry policy for communication errors and other issues.\u003c/p\u003e\n\u003cp\u003eContinue to design external input validation and business-rule checks separately. A model answer can match the required type and still be wrong for the business task. This distinction is a design interpretation based on the official \u003ca href=\"https://docs.typesafe.ai/sdk\"\u003eClient SDKs\u003c/a\u003e and \u003ca href=\"https://docs.typesafe.ai/model-jaggedness/jev-1.13\"\u003emodel limitations document\u003c/a\u003e.\u003c/p\u003e\n\u003ch2\u003e\n\u003ca href=\"#typescript-question-definitions-and-answer-types\" class=\"anchor\" id=\"typescript-question-definitions-and-answer-types\"\u003e\u003c/a\u003eTypeScript question definitions and answer types\u003c/h2\u003e\n\u003cp\u003eThe TypeScript SDK infers return types from question definitions. Choice option keys become the allowed values for the answer. This lets you catch comparisons against disallowed strings at compile time. The SDK requires Node.js 20 or later.\u003c/p\u003e\n\u003cp\u003eThe following illustrative code adapts the official SDK call format. It is not an example that measures results or accuracy. The \u003ccode\u003eTYPESAFE_API_KEY\u003c/code\u003e environment variable must be set.\u003c/p\u003e\n\u003cpre\u003e\u003ccode\u003e\u003cspan\u003eimport { TypeSafeClient, choice } from \"@typesafe-ai/sdk\";\n\u003c/span\u003e\u003cspan\u003e\n\u003c/span\u003e\u003cspan\u003econst api = new TypeSafeClient();\n\u003c/span\u003e\u003cspan\u003e\n\u003c/span\u003e\u003cspan\u003easync function classifyMessage(message: string) {\n\u003c/span\u003e\u003cspan\u003e  const result = await api.systemOne({\n\u003c/span\u003e\u003cspan\u003e    model: \"jev-1.13.0\",\n\u003c/span\u003e\u003cspan\u003e    state: { message },\n\u003c/span\u003e\u003cspan\u003e    questions: {\n\u003c/span\u003e\u003cspan\u003e      topic: choice(\"Classify the subject of message.\", {\n\u003c/span\u003e\u003cspan\u003e        account: \"Account access or password problems\",\n\u003c/span\u003e\u003cspan\u003e        delivery: \"Shipment tracking or delivery problems\",\n\u003c/span\u003e\u003cspan\u003e        other: \"Any subject outside those categories\",\n\u003c/span\u003e\u003cspan\u003e      }),\n\u003c/span\u003e\u003cspan\u003e    },\n\u003c/span\u003e\u003cspan\u003e  });\n\u003c/span\u003e\u003cspan\u003e\n\u003c/span\u003e\u003cspan\u003e  return result.answers.topic;\n\u003c/span\u003e\u003cspan\u003e}\n\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\n\u003cp\u003eThe \u003ccode\u003echoice\u003c/code\u003e type in this example is \u003ccode\u003eaccount | delivery | other\u003c/code\u003e. In actual TypeScript, these strings are literal types shown in quotation marks. Type checking does not, however, determine whether the inquiry was classified correctly. Check installation and call specifications against the \u003ca href=\"https://docs.typesafe.ai/sdk/javascript\"\u003eJavaScript SDK documentation\u003c/a\u003e.\u003c/p\u003e\n\u003ch2\u003e\n\u003ca href=\"#a-confidence-calculation-example-and-common-mistakes\" class=\"anchor\" id=\"a-confidence-calculation-example-and-common-mistakes\"\u003e\u003c/a\u003eA confidence calculation example and common mistakes\u003c/h2\u003e\n\u003cp\u003e\u003ccode\u003econfidence\u003c/code\u003e is a summary of the returned probabilities. It is not the same number as the probability of the selected answer. It should not be read directly as the actual accuracy rate, either. This distinction is especially important when setting criteria for automatic execution.\u003c/p\u003e\n\u003cp\u003eThe official formula for Choice is:\u003c/p\u003e\n\u003cp\u003e\u003ccode\u003econfidence = (maximum probability - 1 / number of choices) / (1 - 1 / number of choices)\u003c/code\u003e\u003c/p\u003e\n\u003cp\u003eThe probability distribution in the official documentation is \u003ccode\u003e0.6, 0.3, 0.1\u003c/code\u003e. There are 3 choices. Substituting the maximum probability of 0.6 gives a confidence of 0.4. A selection probability of 60% and a confidence of 0.4 are different measures.\u003c/p\u003e\n\u003cdiv class=\"overflow-x-auto\"\u003e\u003ctable\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth\u003eCommon interpretation\u003c/th\u003e\n\u003cth\u003eCorrect interpretation\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Common interpretation\"\u003eA confidence of 0.4 means 40% accuracy\u003c/td\u003e\n\u003ctd data-label=\"Correct interpretation\"\u003eA value summarizing the probability distribution according to the formula\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Common interpretation\"\u003eNoul 0.5 means a moderate level\u003c/td\u003e\n\u003ctd data-label=\"Correct interpretation\"\u003eTrue and false have been assigned similar probabilities\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Common interpretation\"\u003eScore's decimal places are precise measurements\u003c/td\u003e\n\u003ctd data-label=\"Correct interpretation\"\u003eA probability-weighted average of the defined level numbers\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Common interpretation\"\u003eHigh confidence means permission checks can be skipped\u003c/td\u003e\n\u003ctd data-label=\"Correct interpretation\"\u003eCheck business permissions and execution conditions separately in code\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\u003c/div\u003e\n\u003cp\u003eThe meanings of these numbers are based on the official \u003ca href=\"https://docs.typesafe.ai/confidence\"\u003eConfidence documentation\u003c/a\u003e. Separating permission checks is a design recommendation for applying that information in production.\u003c/p\u003e\n\u003ch2\u003e\n\u003ca href=\"#how-should-you-compare-pricing-and-response-times\" class=\"anchor\" id=\"how-should-you-compare-pricing-and-response-times\"\u003e\u003c/a\u003eHow should you compare pricing and response times?\u003c/h2\u003e\n\u003cp\u003eThe announced price is $0.042 per 1 million input tokens. Output tokens were described as free. The announced response-time range is 70 to 500 milliseconds. These figures are based on the company's September 15, 2026 announcement.\u003c/p\u003e\n\u003cp\u003eThe company's multiplier comparisons came from an evaluation of a specific workflow. They used the average predictions of other large models as the reference instead of actual correct answers. Speed testing was conducted mainly in the western United States. The results therefore do not represent accuracy across all tasks or response times in Korea.\u003c/p\u003e\n\u003cdiv class=\"overflow-x-auto\"\u003e\u003ctable\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth\u003eComparison item\u003c/th\u003e\n\u003cth\u003eWhat to check\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Comparison item\"\u003eAPI cost\u003c/td\u003e\n\u003ctd data-label=\"What to check\"\u003eActual input-token usage and applicable rate\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Comparison item\"\u003eResponse time\u003c/td\u003e\n\u003ctd data-label=\"What to check\"\u003eRound-trip time in the region where the service runs\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Comparison item\"\u003eAccuracy\u003c/td\u003e\n\u003ctd data-label=\"What to check\"\u003eResults from samples labeled with correct answers for the actual task\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Comparison item\"\u003eOperating cost\u003c/td\u003e\n\u003ctd data-label=\"What to check\"\u003eCost including retries and human review\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\u003c/div\u003e\n\u003cp\u003eComparing prices alone can miss an increase in review work. An adoption decision needs to account for costs through to the final processing result. The conditions behind the announced figures are stated in the \u003ca href=\"https://typesafe.ai/blog/introducing-system-one-models-and-jev\"\u003eJev public announcement\u003c/a\u003e.\u003c/p\u003e\n\u003ch2\u003e\n\u003ca href=\"#nine-limitations-of-jev-113\" class=\"anchor\" id=\"nine-limitations-of-jev-113\"\u003e\u003c/a\u003eNine limitations of Jev 1.13\u003c/h2\u003e\n\u003cp\u003eThe official documentation groups Jev 1.13's failure modes into nine types. This list is based on documents reviewed on September 17, 2026. Success on some examples does not guarantee reliability for that task.\u003c/p\u003e\n\u003cdiv class=\"overflow-x-auto\"\u003e\u003ctable\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth\u003eFailure mode\u003c/th\u003e\n\u003cth\u003eDesign response\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Failure mode\"\u003eInterpreting wording literally\u003c/td\u003e\n\u003ctd data-label=\"Design response\"\u003eState hidden conditions explicitly in the instructions\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Failure mode\"\u003eCalculation and counting\u003c/td\u003e\n\u003ctd data-label=\"Design response\"\u003eHandle arithmetic in code\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Failure mode\"\u003eComparing dates and times\u003c/td\u003e\n\u003ctd data-label=\"Design response\"\u003eConstruct dates, then compare them in code\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Failure mode\"\u003eDouble negatives and multistep reasoning\u003c/td\u003e\n\u003ctd data-label=\"Design response\"\u003eSplit them into direct questions\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Failure mode\"\u003eInputs with much irrelevant content\u003c/td\u003e\n\u003ctd data-label=\"Design response\"\u003ePass only the necessary fields\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Failure mode\"\u003eInputs that steer the judgment\u003c/td\u003e\n\u003ctd data-label=\"Design response\"\u003eInclude leading wording in advance evaluations\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Failure mode\"\u003eConflicts between instructions and selection criteria\u003c/td\u003e\n\u003ctd data-label=\"Design response\"\u003eAlign the meanings of questions and criteria\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Failure mode\"\u003eLack of mathematical consistency between questions\u003c/td\u003e\n\u003ctd data-label=\"Design response\"\u003eManage logical relationships in code\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Failure mode\"\u003eOpen-ended text generation\u003c/td\u003e\n\u003ctd data-label=\"Design response\"\u003eUse a generative model\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\u003c/div\u003e\n\u003cp\u003eNoul and Choice can assign different probabilities to the same meaning. Do not apply a threshold tuned for one type to the other unchanged. The source is \u003ca href=\"https://docs.typesafe.ai/model-jaggedness/jev-1.13\"\u003eJev 1.13 jaggedness\u003c/a\u003e.\u003c/p\u003e\n\u003ch2\u003e\n\u003ca href=\"#conditions-for-use-in-korean-language-services\" class=\"anchor\" id=\"conditions-for-use-in-korean-language-services\"\u003e\u003c/a\u003eConditions for use in Korean-language services\u003c/h2\u003e\n\u003cp\u003eJev processes Korean input, but performance equivalent to English is not guaranteed. The official documentation says its primary training language is English. It also notes performance differences for CJK scripts, including Korean. It provides no general accuracy figure for Korean.\u003c/p\u003e\n\u003cp\u003eUse your own samples to decide whether to use it for Korean. It is worth including inquiries with indirect wording and omissions. Evaluating only translated text may miss differences in how actual customers write. This is an evaluation recommendation based on language-specific performance differences.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eEvaluate clear and ambiguous requests separately\u003c/li\u003e\n\u003cli\u003eReport department classification and sentiment assessment separately\u003c/li\u003e\n\u003cli\u003eRecheck in Korean any thresholds set using English\u003c/li\u003e\n\u003cli\u003eCompare the same samples before and after a model change\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003e\u003ccode\u003ejev-latest\u003c/code\u003e points to a different model when a new version is released. If you tuned thresholds to a specific version, consider pinning that version. Check the official \u003ca href=\"https://docs.typesafe.ai/models\"\u003eModels documentation\u003c/a\u003e for language and version policies.\u003c/p\u003e\n\u003ch2\u003e\n\u003ca href=\"#the-automation-rate-that-is-easy-to-miss-when-evaluating-a-model-replacement\" class=\"anchor\" id=\"the-automation-rate-that-is-easy-to-miss-when-evaluating-a-model-replacement\"\u003e\u003c/a\u003eThe automation rate that is easy to miss when evaluating a model replacement\u003c/h2\u003e\n\u003cp\u003eWhen replacing a model, consider both overall accuracy and the automation rate. Sending every ambiguous case to a person reduces the amount handled automatically. Expanding the scope of automatic processing can increase incorrect executions. The following evaluation method extends the confidence-based branching described in the official documentation.\u003c/p\u003e\n\u003cdiv class=\"overflow-x-auto\"\u003e\u003ctable\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth\u003eMetric\u003c/th\u003e\n\u003cth\u003eHow to calculate or record it\u003c/th\u003e\n\u003cth\u003eWhat it checks\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Metric\"\u003eAutomation rate\u003c/td\u003e\n\u003ctd data-label=\"How to calculate or record it\"\u003eAutomatically processed cases / all evaluated cases\u003c/td\u003e\n\u003ctd data-label=\"What it checks\"\u003eManual work actually reduced\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Metric\"\u003eError rate in automatic processing\u003c/td\u003e\n\u003ctd data-label=\"How to calculate or record it\"\u003eIncorrect automatically processed cases / automatically processed cases\u003c/td\u003e\n\u003ctd data-label=\"What it checks\"\u003eQuality of automated results\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Metric\"\u003eHuman review rate\u003c/td\u003e\n\u003ctd data-label=\"How to calculate or record it\"\u003eCases sent for review / all evaluated cases\u003c/td\u003e\n\u003ctd data-label=\"What it checks\"\u003eRemaining review workload\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd data-label=\"Metric\"\u003eFinal processing time\u003c/td\u003e\n\u003ctd data-label=\"How to calculate or record it\"\u003eMeasure from input to final completion\u003c/td\u003e\n\u003ctd data-label=\"What it checks\"\u003eWaiting time including review\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\u003c/div\u003e\n\u003cp\u003eIf no cases were processed automatically, the error rate in automatic processing cannot be calculated. Record it as not calculable, rather than 0%. Keep the denominator for each metric so models can be compared.\u003c/p\u003e\n\u003cp\u003eYou can evaluate in this order:\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003eHave people label real work samples with the correct answers.\u003c/li\u003e\n\u003cli\u003eFix the question wording and choice definitions.\u003c/li\u003e\n\u003cli\u003eRecord the model version, probabilities, and final branch.\u003c/li\u003e\n\u003cli\u003eCompare automation rates and error rates at different thresholds.\u003c/li\u003e\n\u003cli\u003eSet operating criteria according to the level of error you can accept.\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003eThere is no universally correct threshold. The criteria must reflect the impact of an incorrect execution. The branching principles that provide a starting point are in the \u003ca href=\"https://docs.typesafe.ai/confidence\"\u003eConfidence documentation\u003c/a\u003e.\u003c/p\u003e\n","tags":["AI Agents","AI Development","Decision Making","Development Tools"],"faqs":[{"question":"Does Jev replace general-purpose generative AI?","answer":"It is suited to tasks that involve choosing from predefined answers. Writing free-form responses or summaries requires a generative model."},{"question":"How many options can I include in Choice?","answer":"The official Choice documentation sets a limit of 255 options per question. If you need to handle input outside the list, consider adding an Other option."},{"question":"Does a Noul value of 0.5 mean an average level?","answer":"It means the probabilities of true and false are similar. To evaluate degree or level, Score using described levels is appropriate."},{"question":"What do the decimal places in Score mean?","answer":"The value is the sum of the defined level numbers multiplied by their probabilities. It does not represent a precise measurement or the proportion of customers in question."},{"question":"Can I treat an answer as correct if confidence is high?","answer":"confidence is a value that summarizes the probability distribution. Whether an answer is actually correct must be checked using samples from your work."},{"question":"What does the TypeScript SDK check?","answer":"It infers the answer type from the question definition. It helps find string comparisons with values outside the allowed options. It does not check whether the classification itself is accurate."},{"question":"Does using Jev mean retries are unnecessary?","answer":"Handling output formats and handling communication errors are separate. The official SDK has a default retry policy, so you should also check how to handle API errors."},{"question":"How accurate is it in Korean?","answer":"The official Models documentation does not provide a Korean accuracy figure that can be generalized. Evaluate performance for each question using real Korean input."},{"question":"How much does Jev cost?","answer":"The price announced on September 15, 2026, was $0.042 per 1 million input tokens. Output tokens were stated to be free. Check the official Models documentation again for the applicable price."},{"question":"Can I use jev-latest in a production environment?","answer":"You can, but the model it points to changes when a new version is released. To keep using the version whose thresholds you validated, specify that version's ID."}],"sources":[{"url":"https://docs.typesafe.ai/introduction","title":"TypeSafe AI Introduction","type":"source"},{"url":"https://docs.typesafe.ai/primitives/choice","title":"TypeSafe AI Choice","type":"source"},{"url":"https://docs.typesafe.ai/primitives/score","title":"TypeSafe AI Score","type":"source"},{"url":"https://docs.typesafe.ai/primitives/noul","title":"TypeSafe AI Noul","type":"source"},{"url":"https://docs.typesafe.ai/confidence","title":"TypeSafe AI Confidence","type":"source"},{"url":"https://docs.typesafe.ai/sdk","title":"TypeSafe AI Client SDKs","type":"source"},{"url":"https://docs.typesafe.ai/sdk/javascript","title":"TypeSafe AI JavaScript SDK","type":"source"},{"url":"https://typesafe.ai/blog/introducing-system-one-models-and-jev","title":"Introducing System One Models \u0026 Jev, September 15, 2026","type":"data_point"},{"url":"https://docs.typesafe.ai/model-jaggedness/jev-1.13","title":"Jev 1.13 jaggedness, reviewed September 17, 2026","type":"source"},{"url":"https://docs.typesafe.ai/models","title":"TypeSafe AI Models","type":"source"}],"images":[{"id":1601,"url":"https://injoys.com/rails/active_storage/blobs/proxy/eyJfcmFpbHMiOnsiZGF0YSI6MjQwNDYsInB1ciI6ImJsb2JfaWQifX0=--ea906281db18091febeba00926fa1dd9ebe90294/ai-cb0d0c3f.webp","is_representative":true,"generation_method":"ai_photo","license":"ai_generated","mime_type":"image/webp","width":1536,"height":1024,"translations":{"ko":{"alt":"고객지원 사무실에서 헤드셋을 쓴 두 직원이 모니터를 보며 상의하고, 책상에 청록색 표시등이 켜져 있다.","caption":"신뢰도는 정답률이 아니므로 자동 처리율과 오판율을 함께 평가해야 합니다.","description":null},"en":{"alt":"Two headset-wearing support staff confer over a monitor, with a teal light glowing at their desk and other agents behind them.","caption":"Confidence is not accuracy, so automatic handling and misclassification rates must be evaluated together.","description":null},"ja":{"alt":"カスタマーサポート室でヘッドセットを着けた二人がモニターを見ながら話し合い、机上で青緑色のランプが光っている。","caption":"信頼度は正答率ではないため、自動処理率と誤判定率を併せて評価する必要があります。","description":null},"es":{"alt":"Dos agentes con auriculares consultan un monitor; una luz verde azulada brilla en su mesa y hay más agentes al fondo.","caption":"La confianza no equivale a la tasa de aciertos, por lo que deben evaluarse tanto el procesamiento automático como los errores de clasificación.","description":null},"id":{"alt":"Dua staf dukungan berheadset berdiskusi sambil melihat monitor, dengan lampu hijau kebiruan menyala di meja dan agen lain di belakang.","caption":"Tingkat keyakinan bukan tingkat ketepatan, sehingga penanganan otomatis dan kesalahan klasifikasi perlu dievaluasi bersama.","description":null},"pt":{"alt":"Dois atendentes com fones de ouvido conversam diante de um monitor; uma luz verde-azulada brilha na mesa, com outros agentes ao fundo.","caption":"A confiança não é a taxa de acerto; o atendimento automático e os erros de classificação devem ser avaliados em conjunto.","description":null},"zh-hant":{"alt":"客服辦公室裡，兩名戴耳麥的員工看著螢幕討論，桌上亮著藍綠色指示燈，後方還有其他客服人員。","caption":"信心值不等於正確率，因此須同時評估自動處理率與誤判率。","description":null},"de":{"alt":"Zwei Supportmitarbeitende mit Headsets besprechen sich vor einem Monitor; am Arbeitsplatz leuchtet ein türkisfarbenes Licht.","caption":"Konfidenz ist keine Trefferquote; daher müssen automatisierte Bearbeitung und Fehlklassifikationen gemeinsam bewertet werden.","description":null}}},{"id":1602,"url":"https://injoys.com/rails/active_storage/blobs/proxy/eyJfcmFpbHMiOnsiZGF0YSI6MjQwNTIsInB1ciI6ImJsb2JfaWQifX0=--d07e76337161b074535590d09b1f479ecd50d738/ai-893d1efb.webp","is_representative":false,"generation_method":"ai_semi","license":"ai_generated","mime_type":"image/webp","width":1536,"height":1024,"translations":{"ko":{"alt":"여성이 분기 화살표가 그려진 테스트 보드에서 문의 카드를 사람 검토 쪽으로 옮긴다.","caption":"Jev의 결정 결과를 자동 처리와 사람 검토로 나누는 분기 설계가 중요합니다.","description":null},"en":{"alt":"A woman moves an inquiry card toward the human-review side of a branching test board.","caption":"Jev’s decisions need a clear path to either automatic processing or human review.","description":null},"ja":{"alt":"女性が分岐を示すテストボードで問い合わせカードを人による確認の側へ移している。","caption":"Jevの決定を自動処理と人による確認に振り分ける分岐設計が重要です。","description":null},"es":{"alt":"Una mujer mueve una tarjeta de consulta hacia la zona de revisión humana de un tablero con bifurcaciones.","caption":"Las decisiones de Jev requieren una ruta clara hacia el procesamiento automático o la revisión humana.","description":null},"id":{"alt":"Seorang perempuan memindahkan kartu pertanyaan ke sisi tinjauan manusia pada papan uji bercabang.","caption":"Keputusan Jev perlu diarahkan dengan jelas ke pemrosesan otomatis atau tinjauan manusia.","description":null},"pt":{"alt":"Uma mulher move um cartão de consulta para a área de revisão humana em um quadro de testes com ramificações.","caption":"As decisões do Jev precisam de um caminho claro para o processamento automático ou a revisão humana.","description":null},"zh-hant":{"alt":"一名女子在有分支箭頭的測試板上，將詢問卡移向人工審查區。","caption":"Jev 的決定需要明確分流至自動處理或人工審查。","description":null},"de":{"alt":"Eine Frau verschiebt eine Anfragekarte auf einer verzweigten Testtafel in den Bereich für menschliche Prüfung.","caption":"Für Jevs Entscheidungen braucht es eine klare Verzweigung zwischen automatischer Bearbeitung und menschlicher Prüfung.","description":null}}}],"published_at":"2026-10-02T23:30:14+09:00","updated_at":"2026-10-02T23:30:14+09:00","license":"cc_by","translation_status":"reviewed","available_locales":["ko","en","ja","es"],"data_locales":["ko","en","ja","es","id","pt","zh-hant","de"],"url":"https://injoys.com/en/articles/jev-decision-model-branching-design-and-limitations"}