Why llama-server picked an 828,160-token context: the --fit regression (Sep 16–26, 2026)
llama-server picked an 828,160-token context because builds b10999 to b11200 (September 16 to 26, 2026) started --fit at the model's training context times the 4 automatic slots, 262,144 × 4 = 1,048,576 tokens for Gemma 4 31B, and shrank it only as far as memory required; from b11201 it starts at 262,144 again.
What people saw
In llama.cpp issue #29521 (opened 2026-09-27), Gemma 4 31B
(UD-Q8_K_XL with an F32 mmproj) on an Apple M5 Max with 128 GB crashed on its first prompt with a Metal
out-of-memory error. The build was 7fe450e19, build 11146
(2026-09-23), started without -c. Its log:
llama_context: n_ctx_seq (828160) > n_ctx_train (262144) -- possible training context overflow
srv load_model: the slot context (828160) exceeds the training context of the model (262144) - capping
srv load_model: initializing, n_slots = 4, n_ctx_slot = 262144, kv_unified = 'true'
The model loaded; the crash came on the first prompt. One request can use at most 262,144 tokens, so the
shared pool was sized for 3.16 times what any single request could use.
828,160 is 3,235 × 256: --fit rounds a context it has reduced down to a multiple of
256 tokens (common/-c 262144 -np 1 and it loaded and answered.
The change and the revert
- #28849 “Change max context length for auto-fitting with unified KV”,
merged 2026-09-16 as b04d4e567, first released in
b10999. Its description: auto-fitting “now tries up to model
context length × parallel slots for unified KV”. In
common/fit.cppthe starting context went fromhp_nct * n_streamstohp_nct * n_seq_max; with a unified cachen_streamsis 1 butn_seq_maxis the slot count. - #29437 “Revert "Change max context length for auto-fitting with unified KV"”, merged 2026-09-26 as 2145525a4, released in b11201.
Affected builds: b10999 to b11200. Build 11146 from the issue is inside the
range: on GitHub it is 147 commits after b10999 and 55 before b11201. It matters only when
-c is not set and the KV cache is unified. That is the default when -np is not set:
llama-server then uses 4 slots and kv_unified = true
(tools/
How --fit chooses context, slots and layers now
Read at master a6ea155 (2026-09-29):
- Slots. No
-np: 4 slots with one unified KV cache (server.cpp:156-160). An explicit-np Nwithout--kv-unifiedgives N separate streams (fit.cpp:195-197). - Starting context. No
-c: the training context times the KV streams, so 262,144 for a unified Gemma 4 31B, shared by the 4 slots (fit.cpp:266-279). With-cset, the context is left as it is (fit.cpp:454-456). - Keep it if it fits. If the projection leaves at least
--fit-targetfree on the device (default 1024 MiB), nothing changes (fit.cpp:352-357; defaults in common.h:480-483). - Otherwise shrink the context. A second measurement at
--fit-ctx(default 4,096) per stream, then linear interpolation between the two, rounded down to 256 × streams and never below the floor (fit.cpp:424-430). - Still short: move layers. Unless you set
-ngl, in which case it stops (fit.cpp:463-465), layers are filled back to front, with a MoE model's expert tensors in system memory first (fit.cpp:488-490), then experts moved back to the GPU where there is room (fit.cpp:735-738). - Per slot. Each slot is capped at the training context, and at
--kv-unified-per-slotif set (server-context.cpp:4008-4015).
Two things are not in that measurement. A draft model is fitted together with the main one
(common.cpp:1200-1231), but if measuring it fails, --fit logs
failed to measure the memory of the extra model, fitting without it and carries on
(fit.cpp:225-228). The mmproj (vision projector) is loaded after the fit
(server-context.cpp:1080 versus 1143), so it comes out
of the 1024 MiB margin.
How to tell whether you were hit
-
llama_context: n_ctx_seq (N) > n_ctx_train (M)with N above M. With a unified cache n_ctx_seq is the whole context (src/llama-context.cpp:291-294 ), so N larger than the training context means the pool started above it. -
the slot context (N) exceeds the training context of the model (M) - cappingandinitializing, n_slots = 4, …, kv_unified = 'true'(server-context.cpp:1213-1232). - The build number in
--versionis between 10999 and 11200, and you did not pass-c.
Flags to set
- One user:
-c 262144 -np 1, or the context you actually use. One slot, no shared pool. - Several users:
-np 4 -c 131072gives four separate caches of 32,768 tokens each; add--kv-unifiedto share one 131,072-token pool instead. Or leave out-cand set--kv-unified-per-slot N, which sizes the pool to slots × N (server.cpp:164-174; added in #24124, 2026-08-27). - More headroom:
--fit-target 4096(MiB per device) leaves room for an mmproj, a draft model or the desktop. - Or update: b11201 and later start from the training context again.
To see what --fit would pick for your model and GPU, and the explicit flags, use the --fit preview in the LLM VRAM calculator; it follows the steps above with the calculator's estimate in place of llama.cpp's measured buffers, and can show the b10999–b11200 behaviour. Where the context goes in memory: KV cache explained.
Questions
Why does llama-server say n_ctx_seq (828160) > n_ctx_train (262144)?
With no -c and no -np, llama-server runs 4 slots sharing one unified KV cache. Builds b10999–b11200 started --fit at n_ctx_train × 4 = 1,048,576 tokens for that one pool and cut it to what fit in memory, 828,160 (3,235 × 256) in llama.cpp issue #29521. A unified pool gives every sequence the whole context, so each slot was then capped at the training context, 262,144.
Which llama.cpp builds are affected?
b10999 (commit b04d4e5, PR #28849, merged 2026-09-16) through b11200. b11201 (commit 2145525, PR #29437, merged 2026-09-26) reverts the change. It only matters when you do not pass -c and the KV cache is unified, which is the default when you do not pass -np.
How do I stop llama-server from choosing a huge context?
Pass the context yourself: -c 262144 -np 1 gives one slot the model's full training context, and the reporter of #29521 confirmed it loads and answers. With -c set, --fit leaves the context alone and only moves layers if memory is short. For more headroom, raise --fit-target (default 1024 MiB).
What does --fit do on current llama.cpp?
Without -c it starts at the training context per KV stream (one stream when the cache is unified), keeps it if the projection leaves --fit-target free on each device, otherwise interpolates the context down to a multiple of 256 tokens per stream but not below --fit-ctx (default 4,096), and only then moves layers off the GPU, dense layers first and, for MoE models, experts to system memory.
Sources read on GitHub on .