Jev Alternatives You Can Run Locally: Laya, Zero-Shot Encoders, Small LLMs

Jev is TypeSafe's closed decision model, reachable only through an API. These open models do the same classification, yes/no and scoring work on your own machine, listed from closest to Jev to furthest.

What Jev is, and why it has no local version

Jev is TypeSafe's "System One" decision model: you send a state (text or JSON, up to 32,000 tokens) and typed questions, and it returns a Choice, a yes/no probability (Noul) or a Score with calibrated probabilities instead of generated text. It is served only through APIs (typesafe/jev-1.13 on OpenRouter, TypeSafe's own API and Cloudflare), billed per input token; there are no weights to download. The models below run on your own GPU or CPU.

Encoders: the closest local replacements

Small classification encoders do what Jev does, in one forward pass and under 2 GB. Laya is the only one built to Jev's interface: typed Choice, Noul and Score questions with calibrated probabilities.

ModelAnswersCalibratedWeightsMemoryLicense
Laya (English, typed-decisions, multilingual) Choice, Noul, Score Yes 249 MB–850 MB
24 files, see the Laya page
about 1–2 GB on a GPU or CPU Apache-2.0
ModernBERT-large zeroshot v2.0 Choice, Noul No 792 MB about 1–2 GB on a GPU or CPU Apache-2.0
DeBERTa-v3-large zeroshot v2.0 Choice, Noul No 870 MB about 1–2 GB on a GPU or CPU MIT
GLiClass ModernBERT large v3.0 Choice, Noul No 1,596 MB about 1–2 GB on a GPU or CPU Apache-2.0

The zero-shot encoders turn each label into a hypothesis ("This text is about billing.") and score it with natural language inference, so they answer Choice and yes/no questions but have no ordered Score and were not trained for calibrated probabilities. They run with Hugging Face Transformers' zero-shot-classification pipeline (GLiClass with its own gliclass package).

Small LLMs with JSON-schema output

Any chat model can answer a Jev question if you force its reply into a schema, for example llama-server --json-schema or Ollama's format field, and read the probability of each option from the token log-probabilities. It works on the same GPU you already use for chat, but each answer is generated token by token, so it is slower than an encoder and the probabilities are not calibrated. VRAM is the site's estimate at Q4_K_M (MXFP4 for gpt-oss) with an 8K-token context; each model page has every quantization and context.

Which one to pick

  • You call Jev today and want the same questions offline: Laya, which takes the same typed questions.
  • Only English labels or yes/no checks, and you already use Transformers: a zero-shot encoder.
  • Many labels per text (tags): GLiClass, which scores all labels in one pass.
  • The decision needs world knowledge or reasoning more than speed, and you have 3–16 GB of VRAM: a small LLM with JSON-schema output.

Jev facts from OpenRouter's Jev guide and the TypeSafe docs; encoder sizes from the Hugging Face API; all read on .

Frequently asked questions

Can Jev run locally?

No. Jev is served only through APIs (OpenRouter, TypeSafe, Cloudflare) and its weights are not published. The closest local model is the open Laya, which takes the same Choice, Noul and Score questions.

What is the closest open-source model to Jev?

Laya, an Apache-2.0 decision encoder that reproduces Jev's interface and returns calibrated probabilities. Its files are under 1 GB, and it runs on a 4 GB GPU or a CPU.

Can a regular LLM replace Jev?

Yes, with trade-offs. With output forced into a JSON schema, small models such as Gemma 4 and Qwen3 answer choice and yes/no questions, but they are slower, need more VRAM (about 2.4–15 GB at Q4_K_M, MXFP4 for gpt-oss) and their probabilities are not calibrated.

More calculators

Updated