← Back to blog

System One Models: What They Are and When to Reach for Jev Instead of an LLM

By Javier CarraraSeptember 28, 2026

Most of the decisions an agent makes in production aren't conversations. They're classifications: which queue does this ticket belong to? Is this the right tool for this step? Is this LLM answer actually supported by the source it cites? Is this input a jailbreak attempt? For the past few years we've answered all of these with the same hammer: a generative, autoregressive LLM that first writes a prose explanation and then, hopefully, structures it into parseable JSON. It works, but it's slow, expensive, and it was never trained to do this in the first place.

In September 2026, TypeSafe AI shipped Jev, billed as the first publicly available "System One Model": a model that doesn't generate text at all — it evaluates a state and returns typed decisions with calibrated probabilities. The name is a direct nod to Daniel Kahneman's distinction between System 1 (fast, intuitive, pattern-based) and System 2 (slow, deliberate reasoning) in human thought. A conversational LLM reasoning step by step looks like System 2. Jev is built to be an agent's System 1: the part that decides fast, without narrating why.

What a System One Model is

Per TypeSafe's own documentation, a System One Model is designed to "make fast, structured decisions that software can use directly." Instead of generating free-form text, it evaluates a state and returns typed answers and probabilities. Jev, their first public model, is described as "a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out."

The core architectural difference is that Jev isn't autoregressive: it doesn't generate token by token in sequence — it produces every answer in a query in parallel, in a single pass. That's what lets it respond in milliseconds instead of seconds, and it's also why it "can't hallucinate" in the strict sense: because the space of valid outputs is defined in advance by a schema, it's mathematically impossible for it to return a value outside that schema.

Questions posed to Jev are built from three primitives:

  • Choice: pick one option from a predefined list, returning a probability for every option.
  • Score: place the state on an ordered rubric, with a probability-weighted position that can fall between two levels.
  • Noul: return a single probability that a statement is true (a probabilistic boolean).

Every question in a request is evaluated against the same state and questions can be freely mixed within a single call.

How a Jev request works: state plus typed questions (Choice, Score, Noul) evaluated in parallel, returning answers with calibrated probability

How it differs from a standard LLM

Jev / System One Standard LLM
Optimization Reinforcement Learning for Calibrated Decisions (RLCD): epistemically honest probabilities RLHF/RLVR: human preference or verifiable reward
Input Structured program state Sequential conversation messages
Output Type-safe values with calibrated confidence Free-form strings that need parsing and validation
Sampling Parallel: all outputs in a single pass Autoregressive: token by token
Hallucination risk None by design (schema constrains output) Present; requires downstream validation
Typical latency ~70–500ms Seconds to tens of seconds
Cost Input ~$0.042/MTok; output free Input $0.20–$10/MTok; output ~5x input cost

Calibration deserves a separate note: it isn't the same as accuracy. A calibrated model is one where "80% confidence" empirically means it's right 80% of the time it says that — no more, no less. General-purpose LLMs are notoriously overconfident: they say "I'm sure" with the same emotional weight on a correct answer as on a fabricated one. TypeSafe trains Jev specifically to make its confidence honest, via RLCD, which lets application code act directly on the number without an extra layer of human verification.

Latency and cost-per-decision comparison between Jev and the average frontier LLM, based on benchmarks published by TypeSafe AI

What kinds of decisions it can make

Jev doesn't replace a generative LLM — it replaces the decision step that, in most agent pipelines today, is still handled by a generative LLM or by brittle hardcoded rules. The most commonly cited use cases are:

  • Routing and triage: which queue, team, or flow an incoming request should go to.
  • Classification and tagging: categorizing documents, tickets, or events against a fixed set of labels.
  • Tool-selection validation: confirming that the tool an agent is about to call is actually the right one for the current step, before it executes.
  • Scoring against a rubric: scoring leads, résumés, or records against defined criteria.
  • Security screening: catching jailbreak attempts or policy violations in the input before it reaches the main LLM.
  • Verifying LLM output: the most interesting hybrid pattern. A generative LLM extracts a value from a document; Jev receives that extracted value along with its source and answers, as a typed Noul, whether the extraction is actually supported by the text.

The rule its own creators propose for deciding when to reach for it is simple, and all three conditions have to hold at once: you already know the set of possible answers, you make the same decision repeatedly, and you can act on a confidence score. If any one of those is missing — the answer is open-ended, it's a one-off decision, or you need a natural-language explanation of the "why" — a generative LLM is still the right tool.

Hybrid agent architecture: Jev (System 1) handles typed routing, classification, and verification decisions; the LLM (System 2) is reserved for drafting, planning, and generating content

How it does it: speculative fan-out and production architecture

A subtle but central feature of Jev's design is that questions within the same request can't see each other. This isn't a minor limitation — it's what enables the speculative fan-out pattern: asking everything you might eventually need up front, in one parallel call, instead of chaining sequential calls where each depends on the last. TypeSafe reports that batching ten questions into a single request is significantly more efficient than ten sequential calls — up to 5.9x fewer tokens in their own examples — precisely because the model doesn't have to regenerate shared context on every call.

A conceptual example of what this looks like in code (simplified — not the SDK's exact syntax):

from typesafe import Client, Choice, Score, Noul

client = Client(api_key=API_KEY)

result = client.evaluate(
    state=support_ticket_text,
    questions={
        "queue": Choice(options=["billing", "technical", "abuse", "sales"]),
        "urgency": Score(rubric="1 (low) to 5 (critical)"),
        "is_duplicate": Noul(statement="This ticket repeats an already-open complaint"),
    },
)

THRESHOLDS = {"is_duplicate": 0.85}  # the threshold lives in code, not in the prompt

if result.is_duplicate.probability > THRESHOLDS["is_duplicate"]:
    merge_with_existing_ticket()
else:
    route_to_queue(result.queue.choice)

A few production practices that recur across TypeSafe's docs and the technical guides published about Jev:

  • Decision weights and thresholds live in application code, not in the prompt. That makes them reviewable, testable, and versionable like any other piece of business logic.
  • Composite scoring is built from atomic questions. Instead of asking the model one multi-factor evaluation ("is this a good candidate?"), decompose it into separate dimensions (experience, communication, culture fit) and apply the weighting afterward, in code — which lets you re-rank without paying for another inference.
  • Jev doesn't count reliably inside a single question. If you need to count items, the recommendation is to ask one Noul per item and sum them in code, rather than asking the model to count directly.
  • The state you pass in defines exactly what the model can know. When context has several relevant parts, pass them as named fields instead of flattened, concatenated text.
  • The model's judgments are stored separately from the decision policy. This lets you change a threshold or a weight and re-evaluate without paying for a new inference.
  • In production, calls go out with bounded retries, backoff, and per-call timeouts, and in async architectures, independent decisions are parallelized rather than chained.

The caveats worth keeping in mind

No article about a product that shipped a few weeks ago should skip its limits. TypeSafe's own published benchmarks show 67.8% accuracy across their four evaluation workflows — roughly matching GPT-5.6 Terra (67.9%) but trailing GPT-5.6 Sol (74.1%) and Claude Opus 5 (73.1%). Jev's real gain isn't raw accuracy — it's speed and cost per decision, on the order of tens to hundreds of times lower, for decisions you already know how to structure.

On top of that: the evaluation workflows were designed by TypeSafe's own team, with no independent reproduction available yet; the company itself acknowledges it can't prove its pricing isn't subsidized; and — the limitation that matters most for regulated domains — Jev gives no natural-language explanation for why it landed on a number. That matters for debugging and for audits wherever an automated decision needs to be justified to a human. For now, it also only accepts text as input — no images, no other modalities.

None of these caveats invalidate the core idea. They just confirm that, like any new and specialized tool, it's worth adopting where its shape actually fits — bounded, repeated, score-actionable decisions — and not as a universal replacement for the generative LLM that, for now, remains the right piece for everything that doesn't fit that mold.

Sources:

aisystem-one-modelsjevtypesafe-aiai-agentsllmdecision-models