AI FAQs
What is a decision model?
A broken mug connects typed AI answers, calibrated probabilities and the cost of acting. Why a likely answer does not choose the action for you.
Imagine you wrote to an online store because your mug arrived broken, and within a second the reply came back: “Please send us a photo of the damage.” Another customer, with an almost identical message, got her refund instantly. A third waited until the next day for a person to answer. Nobody read the three messages and weighed them; software did, in two steps. First, something estimated how likely each request was to be covered by the store’s policy. Then something else chose what to do with that estimate.
A decision model can mean two things: an AI model that returns typed answers with probabilities, or a formal description of a choice, including its alternatives, uncertainty, consequences and costs or preferences. In the AI products discussed here, it does the first job in our store: it reads a message or record and returns values within a predefined answer format, without composing a reply. The formal sense comes from decision analysis and describes the second step. A system like the store’s needs both, and confusing them is the easiest way to misread what these new models promise.
Let us follow the mug through both.

The example begins at the customer’s home, with a mug that arrived broken. Refunding it, requesting a photo and sending the case to a person are three possible actions. Conceptual illustration.
Where does “System One” come from?
TypeSafe AI, a company developing models for software automation, presented its first “System One model,” Jev, on September 15, 2026. The company explicitly credits psychologist and 2002 Nobel laureate in economics Daniel Kahneman and his book Thinking, Fast and Slow as the inspiration for that name (Almeida, 2026 (opens in a new tab)).
Kahneman’s System 1 works “automatically and quickly”: think of the employee who immediately recognizes a familiar damage claim. System 2 handles “effortful mental activities,” as when that employee checks an unusual policy exception. Both contribute to decisions (Kahneman, 2011, ch. 1 (opens in a new tab)). Jev borrows the image of a quick judgment; a large language model (LLM) working through intermediate reasoning steps offers an analogy to deliberate thought. This is a comparison of functions, not a shared cognitive mechanism. Jev’s probability still needs the store’s cost rule to become a refund.

Recognizing a familiar case and examining an exception: two moments in the employee analogy for Kahneman’s systems. Both can contribute to a decision.
Fast judgment also has a weakness in Kahneman’s account: it can produce biases and unwarranted confidence (Kahneman, 2011 (opens in a new tab)). TypeSafe acknowledges that System 1 is associated with errors and argues that its models can be made more reliable (Almeida, 2026 (opens in a new tab)). For our store, speed leaves another question open: how much should we trust the number?
What it returns instead of text
A request has two parts. The state is what the model must examine: the customer’s message, the order record, the refund policy. The questions are typed in advance, and each one admits a fixed kind of answer. Jev offers three: a Choice among up to 255 named options, a Score on an ordered scale, and a yes-or-no question that TypeSafe calls a Noul, answered with the probability of “yes” (Almeida, 2026 (opens in a new tab); TypeSafe AI, n.d.-a (opens in a new tab)). For our store, one Noul is enough:
{
"model": "jev-latest",
"state": {
"message": "My mug arrived broken. I'd like a refund.",
"refund_policy": "Items damaged in transit are refunded in full."
},
"questions": {
"covered": {
"type": "noul",
"instructions": "Is the situation in `message` covered by `refund_policy`?"
}
}
}
The answer for covered contains no sentence to interpret, only a value such as {"type": "noul", "noul": 0.9}.
That difference matters more than it seems. A generative language model asked the same question can write its answer token by token. Without an output constraint, the program may receive “it depends” or an explanation instead of a label. Structured-output modes can restrict the answer to a schema, including a fixed set of labels. The distinction here is also how the answer is computed: scoring the allowed options directly instead of generating a sequence. Neither a generated confidence number nor a classifier’s probability is automatically well calibrated. A decision model’s answer cannot fall outside the declared options: that is a guarantee about format, which TypeSafe presents as true by construction rather than as a measured result (Almeida, 2026 (opens in a new tab)). That narrow format guarantee is what we should retain from its “cannot hallucinate” claim. The answer can still be wrong; it just cannot be a fourth category nobody defined.
Jev currently accepts text only (TypeSafe AI, n.d.-c (opens in a new tab)). In our example, requesting a photo starts a separate check by a person or a component that can examine images.

From unstructured information to probabilities over possible answers. A conceptual illustration of an AI decision model.
Answering without writing
How does a model answer without writing? TypeSafe has not published Jev’s architecture. It describes a new architecture, a parallel sampler that produces every answer in a single query and a post-training method it calls Reinforcement Learning for Calibrated Decisions, or RLCD (Almeida, 2026 (opens in a new tab)).
Cloudflare, an internet infrastructure company, uses “decision model” for Jev and its own Clef models. It released Clef and Clef-flash with open weights on October 1, 2026, and describes its approach in more detail: a language model reads the whole input in a single pass without generating anything, and an additional scoring component evaluates valid options in parallel from that internal representation (Chen et al., 2026 (opens in a new tab)).
Turning scores into probabilities is the familiar last step of a classifier. If option receives a score , the softmax function converts the scores into probabilities that are positive and add up to 1:
Exponentiating gives each option a positive weight; dividing by the sum normalizes those weights. A difference between scores becomes a ratio between probabilities. With three options scored 2.0, 0.5 and −1.0 (invented values), the probabilities are 0.786, 0.175 and 0.039. This illustrates a scoring mechanism, not Jev’s undisclosed implementation. Avoiding sequential decoding removes work that a generative model would perform for each output token. That can reduce latency, but the total also depends on model size, input length, hardware, load and network travel. TypeSafe reports 70 to 500 ms end to end for Jev and charges only for input tokens; these are provider figures, not timings measured here (Almeida, 2026 (opens in a new tab)).
TypeSafe also derives a separate confidence value from some answers’ probabilities (note 1).
The trade is deliberate. A decision model cannot explain itself in prose, summarize, or produce an option you forgot to list. It answers the questions you wrote, over the options you wrote.
When 0.9 means 0.9
The risk of confident error brings us back to Kahneman: chapter 22 explains why feeling sure of an intuition is an unreliable guide to its validity (Kahneman, 2011, ch. 22 (opens in a new tab)). In our store, we can check the model’s probabilities against resolved claims. For this yes-or-no question, a model is calibrated when, among claims assigned a 0.9 probability of being covered, about 90% really are covered; among claims assigned 0.3, about 30% are covered. We are comparing the predicted event’s probability with its observed frequency, not calling a low probability a wrong answer. A 2017 study of image and document classifiers found that modern deep networks, although more accurate than their predecessors, had become systematically overconfident. It measured the gap with the expected calibration error: the average gap between confidence and observed accuracy across groups of predictions, weighted by the size of each group (Guo et al., 2017 (opens in a new tab)).
How do you train a model to be honest about probability? By scoring it with a rule under which honesty pays. The Brier score of a yes-or-no forecast compares the stated probability with the outcome , which is 1 if the event happened and 0 if it did not:
Lower is better. Suppose that 80% of messages like ours are truly covered. Reporting 0.8 gives an expected score of . Overstating it as 1.0 gives 0.20; understating it as 0.6 also gives 0.20. Only the truthful probability minimizes the expected penalty, which is what makes the rule strictly proper (Gneiting & Raftery, 2007 (opens in a new tab)); note 2 shows the general case. Cloudflare reports adding a Brier loss when training Clef to refine its calibration. It also uses the name RLCD for a secondary optimization objective that it describes (Chen et al., 2026 (opens in a new tab)). TypeSafe says its RLCD rewards honest probabilities (Almeida, 2026 (opens in a new tab)), but has not published enough detail to tell whether the two methods coincide. A proper scoring rule rewards accurate probabilities in expectation; finite training data and an imperfect model can still leave calibration errors.
Calibration is a property of many answers, not a promise about one. It also holds for the kind of data on which it was measured. TypeSafe notes that English is Jev’s strongest language and that other languages are handled less well, something to keep in mind for messages in Spanish or any other language. Its own list of weaknesses includes literal reading of instructions, unreliable counting and date comparison, and answers that adversarial text in the state can steer (TypeSafe AI, n.d.-b (opens in a new tab); TypeSafe AI, n.d.-c (opens in a new tab)). That last weakness matters for our store, because the message is written by the person who benefits from the decision. A probability measured on someone else’s data is a starting point; the one that counts is measured on your own cases.
From probability to action
Suppose the model reports 0.90 for our mug. Should the store refund it? The probability does not say. What answers the question is the cost of being wrong in each direction, together with the actions available.
Our store has three actions, with costs invented for the example. We count additional losses relative to correctly resolving the claim; a refund owed under the policy is not treated as an error cost. Refunding a claim the policy does not cover loses the USD 40 order. Asking for a photo costs about USD 1 in friction and, we assume, stops nine in ten uncovered claims, which leaves an expected loss of USD 4 per uncovered claim. We assume the photo check never blocks a covered claim. Sending the case to a person costs USD 3 in staff time, and we assume the person always gets it right. With the probability the model reports that the claim is covered, and the probability that it is not covered, the expected cost of each action is:
- Refund at once: .
- Ask for a photo: .
- Send to a person: .
The general rule is to choose the action with the lowest expected cost, weighting each possible state of the world by its probability given what we observed (Elkan, 2001 (opens in a new tab)):
Here is the message, ranges over “covered” and “not covered,” is what the decision model supplies, and is the cost of action when the truth is . The table applies the rule to four messages that received different probabilities. Each row shows the expected cost of the three actions in USD; the lowest one decides.
| P(covered) | Refund | Photo | Person | Action |
|---|---|---|---|---|
| 0.99 | 0.40 | 1.04 | 3.00 | Refund |
| 0.90 | 4.00 | 1.40 | 3.00 | Photo |
| 0.60 | 16.00 | 2.60 | 3.00 | Photo |
| 0.30 | 28.00 | 3.80 | 3.00 | Person |
All costs are illustrative. The exact upper boundary is , about 97.22%. Above it, refund; between 50% and that boundary, ask for a photo; below 50%, send the case to a person. At a boundary two actions tie: for this example we choose person at 50% and photo at (note 3). The mug at 0.90 gets a photo request even though the model considers it very likely covered: if that probability is well calibrated, one automatic refund in ten would go to an uncovered claim, and the photo has the lower expected cost under our assumptions.
Now remove the photo option. With only “refund” and “person,” the boundary moves to 92.5%. Adding an intermediate action did not change a single probability, yet it raised the threshold for automatic refunds by almost five percentage points. The thresholds belong to the decision, not to the model: they depend on the menu of actions and on what each mistake costs. TypeSafe’s documentation makes the same point when it advises stricter confidence thresholds for irreversible actions than for harmless ones, with the risk tolerance written into the application’s code (TypeSafe AI, n.d.-a (opens in a new tab)).

Illustrative decision regions, calculated from the example’s costs. At the same probability of 0.95, adding the photo option changes the chosen action from refund to photo. Boundary ties follow the convention in note 3.
How old is a decision model?
Ronald A. Howard named decision analysis in a 1966 paper that proposed applying decision theory systematically to real choices (Howard, 1966/1983 (opens in a new tab)). Decision theory itself is older. The discipline represents choices like the store’s with influence diagrams: a rectangle for the decision (refund, photo, person), an oval for the uncertainty (is the claim covered?), a node for the value being optimized, and arrows that represent information available for a decision or dependencies between variables (Howard & Matheson, 1984/2005 (opens in a new tab)). In that vocabulary, the decision model is the whole diagram. The new AI models fill one oval: they estimate uncertainty from text that would be difficult to cover with hand-written rules.
In business software, the older sense has a very concrete form. Decision Model and Notation, or DMN, is a standard of the Object Management Group, an industry standards consortium, for expressing decisions as diagrams and executable logic, including decision tables that list conditions and outcomes row by row. Its first formal version was adopted in September 2015. Version 1.6 followed in September 2026 and remained the latest formal release as of October 11, 2026 (Object Management Group, 2026 (opens in a new tab)). A refund rule written in DMN might say that damage reported within 30 days on an item under USD 50 is refunded. People write those rules; nothing is learned. The two traditions combine naturally: a DMN table can take the probability a decision model produces as one more input. Decision trees introduce another overlapping name; note 4 separates the two common senses.
The capability the new models package is younger than the discipline but older than the name. Choosing among labels described in words, without labeled examples of the task, has its own research history. A 2008 study called it dataless classification: label names interpreted with knowledge from Wikipedia were often enough to classify documents without labeled examples. Adding unlabeled documents improved the results to a level competitive with a supervised classifier trained on 100 labeled examples (Chang et al., 2008 (opens in a new tab)). In 2019, another study recast zero-shot text classification, classification without training examples for the task, as entailment: does the text imply a statement such as “this text is about sports”? That framing let labels of different kinds, such as topics, emotions or situations, be scored within one framework (Yin et al., 2019 (opens in a new tab)). Our Noul is an entailment question in that sense: does the message, read against the policy, imply that the claim is covered? Large language models later generalized the approach. The probability a language model assigns to the text of each option can serve as that option’s score, and Cloudflare’s first prototype started there before it trained Clef (Chen et al., 2026 (opens in a new tab)).
The recent development is the way these capabilities are packaged and offered to developers. TypeSafe launched Jev on September 15, 2026. On October 1, Cloudflare released Clef with open weights, writing that classifiers had existed for some time and that the novelty lay in working over any input without retraining for each new set of categories (Almeida, 2026 (opens in a new tab); Chen et al., 2026 (opens in a new tab)). Between those launches, OpenAI announced Decisions API at DevDay on September 29, 2026: an interface using its Luna model to answer predefined questions with a finite set of possible answers from text or images. The announcement described a limited preview (OpenAI, 2026 (opens in a new tab)). As of October 11, 2026, those product launches and announcements are 10 to 26 days old: a few weeks. Decision analysis has had its name for 60 years, and the dataless classification study is 18 years old. The answer depends on which of those three histories we mean.
Teaching it your store
Every probability so far came from a general model reading messages from a store it has never seen. Specializing a decision model means improving what it says about your cases, and there are four levels, ordered roughly by cost.
The first leaves the model untouched. The question itself is a lever: what goes into the state, how the instructions are phrased and what each option means. TypeSafe serves the same weights to every account and recommends adapting Jev to a domain through the request rather than through custom weights (TypeSafe AI, n.d.-c (opens in a new tab)). The costs are another lever, and they already live in your code. If the mug becomes a USD 120 vase, and the photo still stops nine in ten uncovered claims, the automatic-refund threshold rises from 97.22% to 99.07%, and the boundary for sending a case to a person rises from 50% to 83.33%. Nothing was retrained.
The second level recalibrates the probabilities without changing their order. Suppose you label a sample of past claims and find that, among the messages the model scored at 0.90, only three in four were actually covered. Temperature scaling can address some miscalibration with a single number, , fitted on a held-out calibration sample, then checked on separate evaluation cases. For a yes-or-no answer, it divides the log-odds of the reported probability by and converts the result back into a probability:
In this invented example, a temperature of 2 matches that one observation: it maps 0.90 to exactly 0.75, and 0.99 to 0.909. In our store, that second message would no longer be refunded automatically; it falls below 97.22% and receives a photo request. Dividing by a positive preserves the sign and order of the log-odds, so the more likely answer stays more likely; only the confidence moves. Whether calibrates the other probability ranges still needs to be checked on held-out cases. The 2017 calibration study found this one-parameter method surprisingly effective for neural classifiers (Guo et al., 2017 (opens in a new tab)). Laya, an open-weight decision model published by ConvAI Innovations, reports that fitting one temperature per question type on domain data cut its expected calibration error from 0.466 to 0.081, a figure from its own developer (ConvAI Innovations, n.d. (opens in a new tab)).
The third level trains something small on top of the model. TypeSafe suggests decomposing broad judgments into atomic questions and training a classical model downstream on Jev’s probabilities (TypeSafe AI, n.d.-c (opens in a new tab)). In our store, separate Nouls could ask whether the message describes damage in transit and whether it describes a change of mind. A logistic regression, a simple statistical model that weighs each input and returns a probability, trained on resolved claims would learn how much each answer should count. The decision model stays as it is; what you train is small, cheap and easy to inspect.
The fourth level adjusts the weights, which requires open weights or a provider that offers the service. Cloudflare published Clef under the Apache 2.0 license and announced a fine-tuning service based on reinforcement learning, in which the model is adjusted according to rewards for its answers. Its own engineers deliver the service alongside each customer for now, with a self-serve version planned, using request data captured through Cloudflare’s AI Gateway (Chen et al., 2026 (opens in a new tab)). Laya publishes a notebook that fine-tunes its model in about four hours on two NVIDIA T4 GPUs on Kaggle, a data-science platform that offers free GPU time. Its developer also notes that the strongest results on his own benchmark come after fine-tuning on that benchmark’s training data, not from the base model (ConvAI Innovations, n.d. (opens in a new tab)).
Where do the labels come from? The store already produces them: every case sent to a person ends in a human decision. Those cases, however, are not a random sample. They are the ones the model scored at 50% or below, so a model tuned only on them learns from only the low-coverage-probability region and never sees whether its confident answers were right. A useful training and evaluation set also includes a random slice of the automatic decisions, reviewed after the fact.
Costs can enter training by changing the proportion of examples from each class, tying the training setup to a particular cost ratio. Charles Elkan derives that rebalancing option, but recommends a different approach for the classifiers he examines: train on the data as given, then use the estimated probabilities to choose the action with the lowest expected cost. His equation 2 gives the optimal binary threshold from the cost matrix; he also recommends adjusting it empirically when necessary (Elkan, 2001 (opens in a new tab)). That separation is what lets us replace the mug with a vase and change the thresholds without retraining. After any fine-tuning, measure calibration again on cases the model has not seen, because the thresholds rely on those probabilities.
Where it fits and where it does not
A decision model fits when the question is a judgment a knowledgeable person could make in a few seconds and the possible answers can be listed: which team should handle a ticket, whether a command is safe to run, whether a passage supports a claim. It fits poorly where the answer must be computed, as with arithmetic, counting or comparing dates, which TypeSafe itself recommends keeping in code, and where the output must be written, as with an explanation, a summary or a reply (TypeSafe AI, n.d.-b (opens in a new tab)).
Three practical consequences follow from what we have seen. First, measure on your own cases before trusting a vendor’s figures. TypeSafe’s headline comparisons come from evaluations designed by its own team, which it acknowledges may introduce bias, and Cloudflare’s benchmark tables are likewise its own (Almeida, 2026 (opens in a new tab); Chen et al., 2026 (opens in a new tab)). For the store, a practical first step is shadow mode: keep the existing process in charge while recording what the model and cost rule would have recommended, without issuing refunds or contacting customers through that trial path. Compare those recommendations with resolved claims before enabling automatic actions.
Second, record what produced each decision: the state, the model version, the versions of the questions and policy, the probabilities, the thresholds and the action. Changing an instruction or an option description can change behavior even when the model stays the same. That record lets an auditor reconstruct the application’s rule; it does not reveal the model’s internal reasoning or prove the decision was justified. TypeSafe advises pinning a specific version once thresholds are tuned, because an alias can change the model underneath (TypeSafe AI, n.d.-c (opens in a new tab)).
Third, take special care when the decision concerns a person. A single number can carry biases learned from data without leaving a trace in any rule, so comparing errors across affected groups is one necessary part of an evaluation, alongside examining data and the policy itself.
Back to the broken mug
We can now explain the three replies. The same decision model read three messages and returned three probabilities, say 0.99, 0.90 and 0.30. The decision did not happen inside the model. It happened in a rule that compared expected costs: an immediate refund when an error was very unlikely, a photo when a cheap check was worth its friction, and a person when the case was too uncertain to automate.
So what is a decision model? In the AI products discussed here, a model that turns unstructured input into probabilities over answers you defined. In the older and broader sense, the full description of a choice, including the costs that someone, not the model, decided to accept. The first makes the second cheap enough to run on every message. It does not make it unnecessary.
Notes
- Confidence is not the decision either. For Choice and Score answers, TypeSafe also returns a
confidencevalue derived from the probabilities. For a Choice with options it is : 1 when all the probability falls on one option and 0 when it is spread evenly (TypeSafe AI, n.d.-a (opens in a new tab)). For the invented scores in the softmax example, the top probability is 0.786 and the confidence about 0.68. Both are inputs to the decision; neither is the decision. Return to the explanation - Why only the truthful probability wins. If the event occurs with probability and we report , the expected Brier score is . Its derivative, , vanishes only at , and the second derivative, 2, is positive, so that point is the unique minimum. With : and . Return to the explanation
- Where the thresholds come from. Refund beats photo when , that is, , or . Photo beats person when , that is, . Without the photo, refund beats person when , that is, , or . At equality the two costs tie. We assign the ties to person at , photo at , and person at without a photo. Real costs and different error rates give other boundaries. Return to the explanation
- A decision tree is another thing again. In machine learning, a decision tree is a predictive model that splits data with successive questions. In decision analysis, a decision tree is a diagram of choices and chance events evaluated by expected value, close to the influence diagrams above. The shared name hides two different objects. Return to the explanation
References
Almeida, D. (2026, September 15). Introducing System One models & Jev. TypeSafe AI Blog. Original source (opens in a new tab)
Chang, M.-W., Ratinov, L.-A., Roth, D., & Srikumar, V. (2008). Importance of semantic representation: Dataless classification. In Proceedings of the Twenty-Third AAAI Conference on Artificial Intelligence (pp. 830–835). AAAI Press. Proceedings record (opens in a new tab)
Chen, M., Reneau, A., & Flansburg, K. (2026, October 1). Introducing Clef: Our open-source decision models, and new RL fine-tuning platform. The Cloudflare Blog. Original source (opens in a new tab)
ConvAI Innovations. (n.d.). Laya [Project page]. Retrieved October 11, 2026, from Original source (opens in a new tab)
Elkan, C. (2001). The foundations of cost-sensitive learning. In B. Nebel (Ed.), Proceedings of the Seventeenth International Joint Conference on Artificial Intelligence (pp. 973–978). Morgan Kaufmann. Proceedings record (opens in a new tab)
Gneiting, T., & Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477), 359–378. DOI (opens in a new tab)
Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. In D. Precup & Y. W. Teh (Eds.), Proceedings of the 34th International Conference on Machine Learning (Vol. 70, pp. 1321–1330). PMLR. Proceedings (opens in a new tab)
Howard, R. A. (1983). Decision analysis: Applied decision theory. In R. A. Howard & J. E. Matheson (Eds.), Readings on the principles and applications of decision analysis (Vol. 1, pp. 95–113). Strategic Decisions Group. (Original work published 1966) Bibliographic record (opens in a new tab)
Howard, R. A., & Matheson, J. E. (2005). Influence diagrams. Decision Analysis, 2(3), 127–143. DOI (opens in a new tab) (Original work published 1984)
Kahneman, D. (2011). Thinking, fast and slow. Farrar, Straus and Giroux. Book (opens in a new tab)
Object Management Group. (2026). Decision Model and Notation (DMN) (Version 1.6). Specification (opens in a new tab)
OpenAI. (2026, September 29). DevDay 2026 recap. Announcement (opens in a new tab)
TypeSafe AI. (n.d.-a). Confidence. TypeSafe documentation. Retrieved October 11, 2026, from Original source (opens in a new tab)
TypeSafe AI. (n.d.-b). Jev 1.13 jaggedness. TypeSafe documentation. Retrieved October 11, 2026, from Original source (opens in a new tab)
TypeSafe AI. (n.d.-c). Models. TypeSafe documentation. Retrieved October 11, 2026, from Original source (opens in a new tab)
Yin, W., Hay, J., & Roth, D. (2019). Benchmarking zero-shot text classification: Datasets, evaluation and entailment approach. In K. Inui, J. Jiang, V. Ng, & X. Wan (Eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 3914–3923). Association for Computational Linguistics. DOI (opens in a new tab)
People mentioned
Daniel Kahneman
Psychologist and 2002 Nobel laureate in economics. Recognized for bringing psychological research into economics, especially the study of judgment and decision-making under uncertainty.
Sources
Ronald A. Howard
Engineer and Stanford professor who helped establish decision analysis as a discipline. His work developed systematic ways to represent uncertainty, preferences and alternatives in consequential choices.
Sources
Charles Elkan
Computer scientist whose research includes machine learning and data mining. His work on cost-sensitive learning explains how different error costs affect classification decisions.
Sources