All posts
AI ArchitectureAI AgentsDecision ModelsRAGSovereign AI

Jev Explained: Why AI Applications Don't Need an LLM for Every Decision

11 min readThomas Stermole
Jev decision layer diagram showing context becoming a typed decision that routes work to an application workflow, an LLM or human review.

Short answer: Jev is TypeSafe AI's specialized model for returning typed, probabilistic decisions from supplied context. It is designed for bounded questions—such as choosing a route, assigning a score or estimating whether a condition holds—rather than free-form text generation. That makes Jev a possible decision layer in a system that also uses LLMs, rules and human review.

The architecture question is simple: Which parts of an AI application need language, and which need a constrained decision that software can act on?

What is Jev?

Jev is TypeSafe AI's first publicly introduced System One model. TypeSafe describes this model class as an AI interface for fast, structured decisions inside software. A request supplies state or context and typed questions; the response follows the requested answer types. The TypeSafe API documentation describes endpoints for choice, score and yes/no questions. TypeSafe founder Diogo Almeida introduced Jev in the company's System One and Jev launch post.

“System One” is TypeSafe's product and research framing, not an independently established model category. The useful architectural distinction is the output contract: a generative large language model (LLM) is built to produce flexible text, while Jev is presented as a decision model that returns a bounded, machine-readable answer.

Jev can take unstructured text as context. The decision is still constrained by the types and options defined for the request. This is different from asking an LLM to write JSON: with an LLM, the application generates or requests text and then parses and validates it; with Jev, the API is designed around typed decisions. A valid type does not guarantee a correct judgement.

What can Jev do?

The public API describes three decision primitives:

  • Choice: select from a defined set of options, for example a support queue or next workflow.
  • Score: rate content against a criterion or scale, for example how well evidence supports a requirement.
  • Noul: assess a yes/no question and return a probability, for example whether a case needs review.

These primitives can support routing, classification, relevance or suitability checks, risk scoring and approval gates. They are useful when the application can define the shape of the decision in advance.

The schema constrains the answer format; it does not prove that the chosen option or score is right. TypeSafe describes Jev's probabilities as calibrated. That is a vendor claim to validate on representative examples from the target domain, especially before using a confidence threshold to automate a consequential action.

Performance, cost and reliability claims

In its September 15, 2026 launch post, TypeSafe reported 70–500 ms latency for Jev, an input price of $0.042 per million tokens and no charge for output tokens. Those are time-specific vendor figures, not independent measurements for every production workload. The same post discusses the limits and possible bias in its own comparisons. Pricing, availability and latency should be checked against the current service and the application's network path before they become design assumptions.

“Type-safe” also does not mean “factually correct”. A constrained output can prevent a response from having the wrong schema. The selected route, score or probability can still be wrong, so it needs a task-specific evaluation set, error analysis and a fallback for uncertain cases.

Why is a decision layer architecturally interesting?

A common implementation asks an LLM to decide, produce JSON, pass schema validation and retry if the response is malformed. That pattern can be appropriate, but it couples two jobs: making the judgement and serialising it into an exact contract. Retries may repair formatting; they cannot turn a wrong judgement into a correct one.

A dedicated decision model can make this interface more direct when the question has a bounded answer. The application still owns its option lists, thresholds, business rules, permissions and side effects. Jev supplies a decision signal; ordinary code determines what that signal is allowed to trigger. This keeps the decision layer separate from the workflow that consumes it.

A decision layer between context and action

  1. 01

    Input

    Request, documents and authorised context

  2. 02

    Decision layer · Jev

    Choice, score or probability

  3. 03

    Business logic

    Thresholds, permissions and validation

  4. 04

    LLM or tool

    Generate language or run an approved action

  5. 05

    Human review

    Check uncertain or consequential cases

Jev returns a decision signal. The application controls what happens next.

Jev vs LLMs: different jobs, often one system

Generative LLMs remain a good fit when the output needs language: explaining, summarising, drafting, translating or interacting flexibly with a person. Jev is positioned for typed decisions that software can consume. A useful design can combine both rather than forcing a single model to do every job.

TaskJev as a decision modelGenerative LLM
Choose from known support queuesA natural fit for ChoicePossible, but free-form output needs validation
Assess evidence against a rubricA Choice or Score can provide a bounded signalCan also explain or summarise the evidence
Draft an employee-facing answerNot its intended focusA strong fit for natural language
Explore unfamiliar optionsA fixed choice may be too restrictiveBetter suited to open-ended exploration
Trigger a consequential actionCan provide a routing signalCan propose a plan; execution should remain controlled

One pattern is to use Jev to classify an incoming request, route high-confidence cases in application code, ask an LLM to draft a response, and send uncertain results to a person. The right split depends on the task, not on a general ranking of models.

Jev vs rules and traditional classifiers

Classification did not begin with Jev. When categories are stable and labelled examples are available, a conventional classifier may be the simpler option, especially if it can run locally. When a condition can be stated exactly, a rules engine is usually easier to test and explain.

Jev's product distinction is to expose probabilistic software decisions as a dedicated API primitive: choice, score and probability. Whether that contract is more useful than a classifier, an LLM with structured output or a rule depends on accuracy, calibration, latency, hosting, operations and the cost of errors.

Typical taskRulesJevLLM
Exact threshold or process conditionUsually the best fitUnnecessary if deterministic logic is enoughUsually unnecessary
Semantic mapping to a bounded setCan become brittle as rules growWorth evaluating as a decision modelPossible, especially when explanation is also needed
Open-ended summary or responseNot suitableNot its purposeA natural fit
Explainability and local operation are prioritiesTransparent and deterministicCheck hosting and contract requirementsDepends on the model and deployment

Jev in AI agent workflows

An agent does not need to use free-form generation for every internal branch. A bounded decision can help route an incoming request to “research”, “CRM” or “human hand-off”, or select which permitted tool is relevant to the next step.

The application must still enforce the agent's tool allow-list, user permissions, execution limits and approval policy. A model output can suggest a route; it cannot grant authority. For the broader trade-off between autonomy and predictable process paths, see AI Agent vs Workflow vs Automation.

Jev in RAG: evidence and review gates

After retrieval, a decision gate could assess whether the returned passages appear relevant or sufficient for the question. If evidence is weak, the workflow can search again, abstain or request human review. The gate should be evaluated against missing, stale, conflicting and irrelevant documents. Relevance is not the same as truth, and no decision score substitutes for retrieval metrics or source checks.

In a production RAG system, the gate must not bypass access controls or turn missing evidence into an authorised answer. It belongs alongside retrieval and answer evaluation, as described in RAG Evaluation: Metrics, Golden Datasets & Regression Tests.

A decision gate in a RAG workflow

  1. 01

    Did retrieval find evidence that matches the question?

    • No / unclear

      Search again or withhold the answer

    • Yes

      Assess evidence against the task

  2. 02

    Does the evidence meet the defined minimum?

    • Below threshold

      Request human review or abstain safely

    • Above threshold

      LLM drafts with sources; code checks policy

Retrieval, assessment and approval stay distinct. A high score is not proof of truth.

Limitations, privacy and Sovereign AI

Jev is accessed through TypeSafe's hosted API. Sending a request to an external decision service moves its context outside the application's runtime. Before using business data, establish what is sent, where it is processed or stored, how long it is retained, and which contractual terms cover processing, subprocessors and international transfers.

TypeSafe's launch post says its service was based on the US West Coast at the time of publication. That statement does not establish an EU-hosted or on-premises option, nor does a general product page settle the terms for a particular customer. Confirm the current data location, retention controls and agreement before sending sensitive information. GDPR-compliant AI covers the wider data-flow and operating-model questions.

Depending on the protection requirements, a smaller local model, a conventional classifier, deterministic rules or a hybrid may fit better. Sovereign AI is an operating and architecture decision, not a property implied by the decision-model category. See Private AI vs Public Cloud for Enterprises.

Deployment options for probabilistic decisions

  • Hosted decision service

    Typed decision through an external API; review data flows and contract terms.

    Hosted
  • Local model or classifier

    Inference in your environment; own quality, operations and updates.

    Local
  • Rules and hybrid

    Keep deterministic conditions in code and add models only for fuzzy sub-tasks.

    Hybrid
Assess hosting separately from decision logic. A hybrid stack can keep sensitive steps local.

When does Jev make sense—and when doesn't it?

Jev is worth evaluating when a recurring decision has a clear set of possible outcomes, hand-written rules are too brittle, and a generative model would produce more freedom than the task needs. It is especially relevant when the decision signal fits an existing workflow and uncertainty can safely route to review or abstention.

It is a weaker fit for open-ended generation, long explanations or creative exploration; for conditions a simple rule already handles; or when data cannot be sent to an external service. A score does not replace evaluation, monitoring or domain approval.

Start with one decision and a representative, versioned test set. Compare Jev with a rules baseline, a suitable conventional classifier and an LLM with structured output where those alternatives make sense. Measure decision errors, calibration, escalation rate, latency, cost and operating constraints—not just whether a demo looks convincing.

Frequently asked questions about Jev

What is Jev?

Jev is TypeSafe AI's first publicly introduced System One model. It accepts context and typed questions, then returns structured decisions such as a choice, score or yes/no probability.

Is Jev an LLM?

Jev is not a general-purpose text generation model. TypeSafe positions it as a decision model for structured, machine-readable answers.

What is Jev used for?

Jev may fit bounded tasks such as routing, classification, scoring and review gates when rules are too rigid and free-form generation is unnecessary. The fit needs to be evaluated for each use case.

Does Jev replace LLMs?

No. Generative LLMs remain useful for open-ended language, explanations and interaction. Jev can complement an LLM as a decision layer before or after generation.

Can Jev be used in AI agent workflows?

A typed decision model can handle bounded routing, classification or review decisions in an agent workflow. Application code should still enforce tool permissions, execution limits and approvals.

Is Jev suitable for sensitive business data?

That depends on data flows, contract terms, processing and storage locations, retention and the organisation's requirements. Those points must be checked before sending sensitive information to a hosted service.

Conclusion: a decision layer is a useful pattern to evaluate

Jev makes a practical architecture idea explicit: not every probabilistic decision needs the same interface as a chat or text-generation task. Typed answers can simplify a software path when the possible decisions are bounded and the application remains responsible for uncertainty, permissions and escalation.

Jev does not replace generative LLMs, rules or traditional classifiers. It is a distinct decision-model pattern for the AI stack, and its value depends on measured quality, controlled consequences and a deployment model that fits the data. Treat it as an option to evaluate against the alternatives in the system you are building.

Next step

Sounds relevant for your company?

In a no-obligation initial call, we clarify within 30 minutes whether and where getting started is worthwhile for you — honestly and without sales pressure.

Request an initial call