Most software does not need another paragraph from an AI model. It needs a decision: which queue gets this ticket, whether this action is risky, or how urgent a case appears.
Developers often obtain that decision by asking a language model to produce JSON that no person will ever read. The model writes tokens, a schema constrains the tokens, and application code parses them back into an enum or boolean. Jev, TypeSafe AI’s first public “System One” model, removes the writing step. It evaluates shared state against bounded questions and returns typed answers with probability distributions.
That narrower interface is the real story. Jev does not make deterministic code obsolete, and it does not turn a valid schema into a correct judgment. It creates a purpose-built decision layer between hard rules and slower review.
| Decision layer | Best fit | Output | Main limitation |
|---|---|---|---|
| Deterministic rules | Known facts, permissions, arithmetic, limits | Exact boolean or value | Brittle when meaning depends on language |
| Jev | Repeated semantic judgments with bounded answers | Choice, Score, or Noul plus probabilities | Can return the wrong valid answer |
| Structured-output LLM | Decisions that need reasoning or generated explanation | Schema-constrained generated tokens | Slower, costlier, and still needs validation |
| Human review | High-impact or unresolved cases | Accountable judgment | Expensive and difficult to scale |
1.Why writing text is wasteful for software decisions
An LLM is optimized to continue a sequence. Even in JSON mode, it still generates one token after another. That is useful when the output is a reply, explanation, plan, or transformation. It is unnecessary overhead when the only valid outcomes are `billing`, `technical`, and `account`.
Jev changes the contract from “write an answer that matches this schema” to “score the answers declared by this schema.” TypeSafe describes the model as a frontier-intelligence function call: unstructured state goes in; typed, probabilistic decisions come out. The current API accepts text-serializable state and several questions in one request.

TypeSafe reports response times of roughly 70–500 ms and an input price of $0.042 per million tokens, with unmetered output. Its published workflow evaluations include peaks of 193.6× faster and 444.6× cheaper than the compared LLM workflows. Those are vendor-run results from TypeSafe’s own evaluation setup, not universal production benchmarks. TypeSafe itself says the largest gains are likely near the high end of real-world outcomes.
2.What Jev actually returns
Every request contains shared `state` and one or more bounded `questions`. The state might be a support ticket plus account facts. Each question selects one of three primitives.
2.1Choice selects one declared option
`Choice` is an enum with probabilities. A routing question might declare `billing`, `technical`, `account`, and `other`. Jev returns one of those keys, a probability for each key, and a confidence value. It cannot invent a fifth department, but it can choose the wrong department.
2.2Score places the state on a defined scale
`Score` maps a judgment onto levels whose meanings you write. A priority scale could define 0 as routine, 1 as time-sensitive, and 2 as service-blocking. The descriptions matter more than the numbers: Jev evaluates the semantic criteria; your code receives the selected level and its distribution.
2.3Noul evaluates a proposition
`Noul`—TypeSafe’s yes/no primitive—returns the probability that a stated criterion is true. It suits questions such as “Does this request claim an unauthorized charge?” A high probability is evidence for a policy decision, not permission to bypass the policy.
Several questions can reuse the same state in one call. That makes Jev especially interesting for triage, where one ticket may need a destination, urgency score, and escalation signal at once.
{
"state": {
"ticket": "I upgraded yesterday, the dashboard still says Free, and I was charged twice.",
"account_age_days": 420,
"recent_refunds": 0
},
"questions": {
"destination": {
"type": "choice",
"options": ["billing", "technical", "account", "other"]
},
"urgency": {
"type": "score",
"levels": {
"0": "Routine question with no current loss of access or money",
"1": "Time-sensitive problem with limited impact",
"2": "Active loss of access, money, or business operation"
}
},
"needs_specialist": {
"type": "noul",
"criterion": "Resolving this safely requires a specialist rather than a standard response"
}
}
}
The exact wire format can change while Jev is young, so production clients should follow the current TypeSafe API schema rather than copy an old example blindly.
3.Jev versus structured LLM output
Structured output solves a transport problem: it constrains generated output to a schema. Jev changes the computation itself. An LLM still generates tokens representing a choice; Jev scores the bounded choices directly.
That distinction produces three practical differences.
First, Jev exposes a distribution over the answers you declared. An enum from an LLM may be syntactically valid without revealing how close the alternatives were. Second, Jev can evaluate multiple decomposed questions against shared state in parallel. Third, Jev cannot write the customer reply or explain a novel edge case. A generative model remains the better tool when the desired output is language.
The choice is not ideological. Use rules for facts your program already knows. Use Jev for repeated semantic judgments with bounded outcomes. Use a reasoning model when the case needs synthesis, tool use, or an explanation. Use a human when impact or uncertainty makes automated error too costly.
4.Typed output is not decision correctness
“Cannot hallucinate” is easy to misread. Jev cannot emit a value outside the declared answer type. If the only destinations are `billing`, `technical`, and `account`, the response will not invent `executive_support`. That is a useful guarantee.
It is not a guarantee that `billing` is correct. A model can be wrong while remaining perfectly type-safe. It can also be confidently wrong when the production data differs from its calibration data, when question criteria overlap, or when the state contains distracting detail.
Treat the probabilities as measurements to validate, not universal truth. Build a labeled set from the actual workload, compare predicted probabilities with observed outcomes, and choose thresholds based on the cost of each error. Re-test after model upgrades or material changes in traffic.
5.A safer confidence-gated cascade
The strongest production pattern gives each layer a narrow job.
Deterministic code checks permissions, spending limits, account status, arithmetic, dates, and destructive-action constraints. Jev then handles the bounded semantic judgment. A policy function maps the returned probabilities and business impact to an action or escalation path.
For example, a support system might use the following illustrative policy:
- High confidence and low impact: route automatically.
- Ambiguous distribution: send the case to a reasoning model for a second pass.
- High impact or unresolved uncertainty: require human review.
Numbers such as 0.80 or 0.95 are not universal defaults. A threshold is defensible only after replaying labeled examples and measuring false positives, false negatives, review volume, and downstream cost.
The architecture matters because Jev should provide evidence to policy, not become the policy. Code owns the action. The model never grants itself permission to refund money, delete data, or call a destructive tool.
6.Where Jev fits—and where it does not
Jev is a strong candidate when four conditions hold: the input contains language, the output set is bounded, the decision repeats at meaningful volume, and uncertainty can be routed rather than hidden. Ticket triage, content classification, lead qualification, policy checks, agent tool gating, and priority scoring fit that shape.
Ordinary rules are better when the answer already exists in structured data. Do not ask a model whether an invoice is overdue when code can compare two dates. A conventional classifier may be better when the label set is stable, training data is plentiful, latency must be extremely low, and operating your own model is justified. An LLM is better when the output must be written, the answer space is open, or the task requires multi-step reasoning.
Jev also remains early. TypeSafe introduced it on September 15, 2026, and has not published model weights or a full independent evaluation. Vercel reported that nearly 13% of paid AI Gateway teams tried Jev within its first 24 hours—the fastest initial adoption in that gateway’s history. That proves developer interest, not durable reliability.

7.The interface shift matters more than the benchmark
Jev’s most important idea is not that one new model wins a latency chart. It is that software decisions deserve an interface designed for software: bounded answers, explicit uncertainty, and a clean handoff to deterministic policy.
That interface removes an awkward loop in which an application asks a writing model for prose-shaped data and immediately turns it back into control flow. It also makes uncertainty harder to ignore. The probability distribution can trigger review instead of being flattened into a confident-looking enum.
The right standard is therefore not “Did the response parse?” It is “Was the decision calibrated on this workload, and did the surrounding system respond safely when confidence or impact demanded review?” Jev gives developers a useful new primitive for that design. It does not remove the engineering around it.


