Jev Alternatives You Can Run Locally: Laya, Zero-Shot Encoders, Small LLMs
Jev is TypeSafe's closed decision model, reachable only through an API. These open models do the same classification, yes/no and scoring work on your own machine, listed from closest to Jev to furthest.
What Jev is, and why it has no local version
Jev is TypeSafe's "System One" decision model: you send a state (text or JSON, up to 32,000 tokens) and typed questions, and it returns a Choice, a yes/no probability (Noul) or a Score with calibrated probabilities instead of generated text. It is served only through APIs (typesafe/jev-1.13 on OpenRouter, TypeSafe's own API and Cloudflare), billed per input token; there are no weights to download. The models below run on your own GPU or CPU.
Encoders: the closest local replacements
Small classification encoders do what Jev does, in one forward pass and under 2 GB. Laya is the only one built to Jev's interface: typed Choice, Noul and Score questions with calibrated probabilities.
| Model | Answers | Calibrated | Weights | Memory | License |
|---|---|---|---|---|---|
| Laya (English, typed-decisions, multilingual) | Choice, Noul, Score | Yes | 249 MB–850 MB 24 files, see the Laya page | about 1–2 GB on a GPU or CPU | Apache-2.0 |
| ModernBERT-large zeroshot v2.0 | Choice, Noul | No | 792 MB | about 1–2 GB on a GPU or CPU | Apache-2.0 |
| DeBERTa-v3-large zeroshot v2.0 | Choice, Noul | No | 870 MB | about 1–2 GB on a GPU or CPU | MIT |
| GLiClass ModernBERT large v3.0 | Choice, Noul | No | 1,596 MB | about 1–2 GB on a GPU or CPU | Apache-2.0 |
The zero-shot encoders turn each label into a hypothesis ("This text is about billing.") and score it with natural language inference, so they answer Choice and yes/no questions but have no ordered Score and were not trained for calibrated probabilities. They run with Hugging Face Transformers' zero-shot-classification pipeline (GLiClass with its own gliclass package).
Small LLMs with JSON-schema output
Any chat model can answer a Jev question if you force its reply into a schema, for example llama-server --json-schema or Ollama's format field, and read the probability of each option from the token log-probabilities. It works on the same GPU you already use for chat, but each answer is generated token by token, so it is slower than an encoder and the probabilities are not calibrated. VRAM is the site's estimate at Q4_K_M (MXFP4 for gpt-oss) with an 8K-token context; each model page has every quantization and context.
| Model | GGUF file | VRAM |
|---|---|---|
| MiniCPM5 2B | 1,616 MB | 2.42 GB |
| Granite 4.2 3B | 2,317 MB | 3.46 GB |
| Nemotron 3 Nano 4B | 2,900 MB | 3.10 GB |
| Gemma 4 E4B | 4,977 MB | 5.64 GB |
| Qwen3 8B | 5,028 MB | 6.81 GB |
| Granite 4.2 8B | 5,539 MB | 7.32 GB |
| Qwen3.5 9B | 5,681 MB | 6.76 GB |
| Gemma 4 12B | 7,122 MB | 8.57 GB |
| gpt-oss-20b | 12,110 MB | 14.8 GB |
Which one to pick
- You call Jev today and want the same questions offline: Laya, which takes the same typed questions.
- Only English labels or yes/no checks, and you already use Transformers: a zero-shot encoder.
- Many labels per text (tags): GLiClass, which scores all labels in one pass.
- The decision needs world knowledge or reasoning more than speed, and you have 3–16 GB of VRAM: a small LLM with JSON-schema output.
Jev facts from OpenRouter's Jev guide and the TypeSafe docs; encoder sizes from the Hugging Face API; all read on .
Frequently asked questions
Can Jev run locally?
No. Jev is served only through APIs (OpenRouter, TypeSafe, Cloudflare) and its weights are not published. The closest local model is the open Laya, which takes the same Choice, Noul and Score questions.
What is the closest open-source model to Jev?
Laya, an Apache-2.0 decision encoder that reproduces Jev's interface and returns calibrated probabilities. Its files are under 1 GB, and it runs on a 4 GB GPU or a CPU.
Can a regular LLM replace Jev?
Yes, with trade-offs. With output forced into a JSON schema, small models such as Gemma 4 and Qwen3 answer choice and yes/no questions, but they are slower, need more VRAM (about 2.4–15 GB at Q4_K_M, MXFP4 for gpt-oss) and their probabilities are not calibrated.
More calculators
- Fine-Tuning VRAM Calculator GPU memory for full fine-tuning, LoRA and QLoRA in Transformers or Unsloth, checked against published runs. Open →
- Image & Video Model VRAM Calculator Peak VRAM and system RAM for FLUX.2, Wan 2.2, LTX-2 and Qwen-Image in ComfyUI, from exact file sizes, checked against public runs. Open →
- Laya VRAM Requirements Real file sizes and run-time memory of the three Laya decision-encoder checkpoints, and three ways to run them. Open →
- vLLM KV Cache & Concurrency Calculator The KV cache pool vLLM allocates, in tokens, and how many requests fit at once, with the vllm serve command. Open →
- What Can My PC Run? Detects your GPU in the browser and lists the local LLMs it runs, with the best quantization and speed. Open →
- MiniMax H3 VRAM Calculator
(ComfyUI) VRAM and system RAM for MiniMax H3 video in ComfyUI: pruned, INT8, NVFP4 and GGUF files on 8–96 GB GPUs. Open →
Updated