All posts
AI ArchitectureLocal LLMDecision ModelsAI AgentsSovereign AI

Local Jev Alternatives: Which Small LLMs Work as a Fast Decision Layer?

13 min readThomas Stermole
Scatter plot of local Jev alternatives comparing decision accuracy with median latency, highlighting Ministral 3 8B as the strongest speed-quality trade-off in the screening.

Short answer: Yes. Small local open-weight LLMs can act as a practical local Jev alternative for some bounded decision tasks. In my screening, Ministral 3 8B offered the strongest balance of speed and decision quality. Several smaller 1B–4B models were faster, but not accurate enough. More importantly, the experiments showed that model size alone does not determine whether a fast decision layer works. Decision contracts, state, evidence and real-world evaluation matter just as much.

TypeSafe's Jev is built around a compelling idea: not every decision inside an AI system needs a large generative language model.

Agents and automated workflows constantly make small decisions:

  • Do I need web search?
  • Which tool should I use?
  • Does this message need attention?
  • Is there enough evidence?
  • Can the workflow continue automatically?
  • Should a human review this?

Calling a large reasoning model for every such decision can add unnecessary latency and cost.

TypeSafe positions Jev as a System One Model: unstructured state goes in, typed probabilistic decisions come out. Instead of generating arbitrary text, Jev is designed specifically for fast machine-consumable decisions. I covered the concept in more detail in Jev explained: why AI applications do not need an LLM for every decision.

That led me to a different question:

Do I actually need a specialized model such as Jev, or can a small open-weight LLM perform a similar role locally?

I decided to test that architectural hypothesis.

This is not a direct Jev benchmark. I did not run Jev under the same local harness. Instead I asked:

How far can standard small local LLMs go if I constrain their decision space aggressively?

The answer turned out to be much more interesting than a simple model ranking.

What would a local Jev alternative actually mean?

Jev is not simply a smaller chat model. TypeSafe describes typed decision primitives such as Noul, Choice and Score. Possible outputs are defined in advance, and Jev returns structured decisions together with probabilities or confidence. The TypeSafe introduction to System One and Jev explains the model concept, while the API documentation describes the interface.

My local models work differently. They remain generative LLMs.

I constrain them using:

  • a small decision space,
  • an explicit decision contract,
  • structured output,
  • JSON Schema,
  • very short completions,
  • temperature 0.

The goal is not to reproduce Jev internally. The goal is to see whether a standard local LLM can perform a similar architectural role.

Which local Jev alternatives did I test?

My initial screening used 24 deliberately small and unambiguous cases:

  • 8 binary decisions,
  • 8 choice decisions,
  • 8 scores from 0 to 4.

The models were served locally through LM Studio. I measured exact-match decision quality, schema validity, p50/p95 latency and serial decisions per second.

This is a small screening dataset, not a statistically significant model leaderboard. The full experiment structure is available in the public Fast Decision Layer directory of my agents repository.

Local Jev alternatives compared by speed and accuracyThirteen local models compared by decision accuracy and median latency.

ModelExactAccuracyp50p95
Ministral 3 8B Instruct, 4-bit21/2487.5%166 ms189 ms
Gemma 4 31B IT19/2479.2%676 ms803 ms
Qwen 3.5 9B MLX, 4-bit19/2479.2%309 ms332 ms
Qwen 3.5 9B MLX, 8-bit18/2475.0%348 ms364 ms
Ministral 3 3B Instruct17/2470.8%90 ms105 ms
Nemotron 3 Nano 4B15/2462.5%256 ms276 ms
Qwen 3.8 9B MLX15/2462.5%494 ms1,163 ms
Twil LM3 MLX13/2454.2%248 ms636 ms
Liquid LFM 2.5 1.2B12/2450.0%103 ms279 ms
Granite 4.2 3B MLX11/2445.8%213 ms488 ms
MiniCPM 5 2B11/2445.8%221 ms549 ms
Qwen 3.5 4B MLX11/2445.8%129 ms143 ms
Granite 4.2 8B MLX, 4-bit9/2437.5%298 ms540 ms

Accuracy for all 13 local modelsAccuracy for all 13 local models.

The smallest models were usually not good enough

The bottom half of the table is arguably more important than the winner.

Liquid LFM 2.5 1.2B reached a median latency of just 103 milliseconds. That sounds ideal for a fast decision layer. Its accuracy was only 50 percent.

Qwen 3.5 4B responded in 129 milliseconds but reached only 45.8 percent. Granite 3B, MiniCPM 2B and Granite 8B were also not convincing.

My first conclusion was therefore:

A local Jev alternative should not be the smallest model available. It should be the smallest model that reliably satisfies the decision contract.

A 100-ms decision is not useful if it is wrong almost half of the time.

Bigger was not automatically better either. Granite 8B performed worse than Ministral 3B. Gemma 31B achieved respectable quality, but its 676-ms median latency was more than four times slower than Ministral 8B.

Parameter count alone therefore does not explain decision quality. Model family, instruction following, quantization, structured-output behavior and the exact semantic task also matter.

In this screening, Ministral 3 8B Instruct at 4-bit was the strongest speed-quality trade-off: 87.5% exact match at 166 ms median latency.

Runtime interoperability is part of model quality

Some models failed before decision quality even became the main problem.

A Qwen 3.8 27B reference run produced 0/24 comparable valid outputs. K2 Horizon could not be loaded successfully in my LM Studio setup at the time.

Several reasoning models also returned schema-valid JSON through reasoning_content while leaving the normal content field empty. A narrowly scoped normalization step made some of those models usable.

This may sound like implementation trivia. In production systems, it is not.

Practical model quality includes predictable runtime behavior and reliable interoperability.

Structured output does not equal a correct decision

After the initial screening, I moved to a more realistic use case: a tool intent router.

It had to select exactly one action:

  • ANSWER_DIRECTLY
  • WEB_SEARCH
  • READ_REPO
  • RUN_CODE

The dataset contained 60 cases, balanced across the four classes.

Ministral 8B produced 60 out of 60 schema-valid responses. Its accuracy was nevertheless only 25 percent because it selected ANSWER_DIRECTLY for almost everything.

Negative evidence from the local Jev alternative testsThree important failure modes: tiny models, tool routing and phishing.

That produced one of the most important lessons from the experiments:

Valid JSON is not the same thing as a valid decision.

JSON Schema solves the syntax problem. It does not solve the semantic problem.

The decision contract can matter more than model size

I then investigated why the router collapsed.

Using Qwen 3.5 9B, I tested more explicit router framing. On the full 60-case dataset, the model achieved:

  • 60/60 schema-valid responses,
  • 51/60 correct decisions,
  • 85% accuracy,
  • 694 ms p50,
  • 739 ms p95.

The takeaway:

A decision contract is not merely an output schema. It defines the semantic boundaries between the allowed decisions.

A specialized decision model can encode part of that behavior during training. With a general-purpose local LLM, much more of it has to come from the surrounding architecture.

The most dangerous benchmark was the perfect one

The clearest example came from an email attention gate.

The task was deliberately narrow:

Does this email create a concrete action that is still open for the recipient?

Only two outputs existed: NEEDS_ACTION and NO_ACTION.

On 48 synthetic cases, Ministral 8B achieved:

  • 100% accuracy
  • 100% recall
  • 100% precision
  • 48/48 schema-valid outputs
  • 286 ms p50

On paper, the problem looked solved.

It was not.

100 percent synthetic accuracy compared with 62.9 percent on real shadow dataThe real shadow evaluation sharply reduced the apparent quality of the perfect synthetic benchmark.

100% offline became 62.9% on real data

I then ran the approach in shadow mode against manually reviewed real email.

Across 35 deduplicated completed reviews, the result was:

MetricResult
Accuracy62.86%
NEEDS_ACTION recall88.24%
NEEDS_ACTION precision57.69%
True positives15
False negatives2
False positives11
True negatives7

Synthetic versus real shadow metricsSynthetic versus real shadow metrics.

A seemingly perfect classifier had turned into a system with just 62.9 percent accuracy.

If I had stopped after the synthetic benchmark, I would have significantly overestimated production quality.

For me, this is one of the most important results of the entire experiment. It also matches the broader production lesson in AI Agent Evaluation: offline evals are necessary, but real shadow data determines whether a gate survives outside the lab.

Some classification problems are actually state problems

The dominant error category involved optional actions: invitations, opportunities and things the recipient could do but did not necessarily have to do.

In some cases the required information simply was not present in the email. The system might also need to know:

  • Has the user already responded?
  • Is this invitation relevant?
  • Has the task already been completed elsewhere?
  • Is there already a calendar event?
  • Has another workflow closed the issue?

That changes the nature of the problem.

Some apparent classification problems are actually state problems.

No amount of prompt engineering can reconstruct information the model has never received.

Phishing exposed another limit of local Jev alternatives

I also tested small local models as a phishing gate.

ExperimentRecallPrecision
Ministral 8B, three classes44.4%80.0%
Binary phishing gate55.6%71.4%
+ Structural signals72.2%72.2%
Challenge set, LLM only33.3%100%
Challenge set + signals41.7%71.4%
Naive hard gate50.0%46.2%

Recall and precision for the six phishing experimentsRecall and precision for the six phishing experiments.

Adding deterministic structural signals initially helped.

Then I deliberately made the benchmark harder. The challenge set included legitimate newsletters, CRM messages, support messages and transactional email with different Reply-To domains.

Performance dropped sharply. A larger 14B model did not solve the fundamental issue either.

The right conclusion was not: “I need a larger LLM.”

It was:

I need better evidence.

A serious phishing system needs signals such as SPF, DKIM, DMARC, domain and URL reputation, redirect analysis, threat intelligence and attachment analysis.

A model can interpret those signals. It cannot replace them.

More model is not a substitute for missing evidence.

The best local Jev alternative was not a single model

The experiments led me to a different architecture: a Fast Decision Layer.

Fast Decision Layer using rules, a small local LLM and escalationUse deterministic logic first, a small local LLM for bounded semantic ambiguity and escalation for complex cases.

1. Deterministic rules first

If a decision can be made reliably in code, do not call a model. Use rules for known states, hard policies, thresholds, structured flags and deterministic security conditions.

2. Small local LLM for semantic ambiguity

Use a small model where the problem is semantic but tightly constrained: intent routing, relevance, prioritization, simple classification or selecting among a few known actions.

3. Escalation for difficult cases

When the problem requires more evidence or deeper reasoning, escalate to a larger LLM, web search, retrieval, specialist APIs, code execution or human review.

The principle is simple:

Deterministic where possible. Small models where useful. Expensive intelligence only where necessary.

Is Ministral 8B a local Jev alternative?

For some use cases: yes.

Ministral 3 8B can perform a similar architectural role for narrow binary or choice decisions, particularly when:

  • the decision space is small,
  • class boundaries are explicit,
  • all necessary information is available,
  • errors have limited consequences,
  • outputs are validated strictly,
  • fallback behavior exists,
  • real shadow evaluations are performed.

It is not a feature-for-feature replacement for Jev.

TypeSafe designed Jev specifically for machine-consumable decisions and describes typed decisions, probabilities, confidence and an architecture optimized for that purpose. TypeSafe's published speed/cost/accuracy comparisons are first-party evaluations; TypeSafe itself notes possible bias because the workflows were created by members of its own model-capabilities team. I have not independently reproduced those claims.

Why local Jev alternatives are still compelling

Local models offer a different set of advantages:

  • data stays inside your infrastructure,
  • no external API call for every decision,
  • predictable local inference costs,
  • offline operation,
  • less provider dependency,
  • custom decision contracts,
  • full control over model and runtime,
  • easier integration into private or on-premise systems.

For organizations with sovereignty, privacy or infrastructure requirements, those properties can matter as much as absolute benchmark latency. This is why I also evaluate the operating model, not just model quality, in Private AI vs Public Cloud for Enterprises.

What I learned

  1. Small local LLMs can make fast decisions. Sub-200-ms inference was achievable locally.
  2. Very small models were often not accurate enough. 1B–4B is not automatically sufficient.
  3. Bigger is not automatically better.
  4. Structured output does not guarantee decision quality.
  5. The decision contract is part of the system.
  6. Synthetic benchmarks can be dangerously optimistic. My clearest example: 100% offline → 62.9% in real shadow evaluation.
  7. Missing state cannot be prompted away.
  8. Missing evidence cannot be prompted away.
  9. Rules, small LLMs and large models belong at different layers of the same system.

Conclusion: What is the best local Jev alternative?

I would not look for one universal Jev replacement.

I would build a local Fast Decision Layer.

For deterministic decisions: code and rules.

For narrow semantic decisions: a capable small local LLM with a strict decision contract.

For complex or uncertain decisions: larger models, external evidence, tools or human review.

Among the models I tested, Ministral 3 8B Instruct was the strongest starting point for fast typed decisions. Several smaller models were faster, but not accurate enough. Qwen 3.5 9B also demonstrated that more complex routing tasks can benefit substantially from a stronger model and a better decision contract.

The most important result was therefore not which model won the benchmark.

The right question is not which model should make every decision. The right question is which decision belongs at which layer.

That is where I see the real potential of local Jev alternatives.

Frequently asked questions about local Jev alternatives

What is a local alternative to Jev?

For bounded decisions, a small local open-weight LLM with structured output and a precise decision contract can perform a similar role inside an architecture. Jev remains a specialized decision model and is not functionally identical.

Which local model performed best in my test?

Ministral 3 8B Instruct at 4-bit reached 87.5 percent exact match at 166 ms p50 latency in the initial 24-case screening.

Is a 1B or 4B model enough?

Usually not for the decision gates I tested. The smallest models were fast but significantly less accurate.

Is structured output enough?

No. A valid schema constrains output shape, not the correctness of the decision.

Why does shadow mode matter?

Because synthetic gold sets can underrepresent real distribution, state and ambiguity. My email gate fell from 100 percent synthetic accuracy to 62.9 percent on real shadow data.

What does a practical local fast decision layer look like?

Deterministic rules first, a small local model for bounded semantic decisions, and an escalation layer for complex or evidence-dependent cases.

Next step

Sounds relevant for your company?

In a no-obligation initial call, we clarify within 30 minutes whether and where getting started is worthwhile for you — honestly and without sales pressure.

Request an initial call