<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>ModelVRAM changelog</title>
    <link>https://modelvram.com/changelog/</link>
    <atom:link href="https://modelvram.com/changelog/rss.xml" rel="self" type="application/rss+xml" />
    <description>New calculators, model pages, data and fixes on ModelVRAM.</description>
    <language>en</language>
    <lastBuildDate>Wed, 30 Sep 2026 12:00:00 GMT</lastBuildDate>
    <item>
      <title>Accuracy page: prompt speed and image/video checks</title>
      <link>https://modelvram.com/accuracy/#prompt-speed</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#accuracy-prompt-speed</guid>
      <pubDate>Wed, 30 Sep 2026 12:00:00 GMT</pubDate>
      <description>The accuracy page now also lists all 32 llama-bench runs behind the time-to-first-token estimate, with the error per hardware family, and the ComfyUI load logs behind the image and video file sizes.</description>
    </item>
    <item>
      <title>Vendor-reported benchmark scores next to VRAM</title>
      <link>https://modelvram.com/best-llm-for-vram/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#vendor-reported-benchmarks</guid>
      <pubDate>Wed, 30 Sep 2026 12:00:00 GMT</pubDate>
      <description>GPQA, MMLU-Pro and LiveCodeBench scores that makers published in their own model cards, linked to the source: on model pages, as a GPQA column in the per-size lists, and as a y-axis option on the chart.</description>
    </item>
    <item>
      <title>Bespoke Nimble 9B model page</title>
      <link>https://modelvram.com/llm-vram-calculator/bespoke-nimble-9b/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#bespoke-nimble-9b</guid>
      <pubDate>Wed, 30 Sep 2026 12:00:00 GMT</pubDate>
      <description>VRAM per quant and context for Bespoke Nimble 9B, with its GGUF files and Ollama tags.</description>
    </item>
    <item>
      <title>Article: the llama-server --fit context regression</title>
      <link>https://modelvram.com/llm-vram-calculator/fit-context-regression/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#fit-context-regression</guid>
      <pubDate>Wed, 30 Sep 2026 12:00:00 GMT</pubDate>
      <description>Why llama.cpp builds b10999–b11200 picked contexts such as 828,160 tokens, how --fit sizes context, slots and layers now, and the flags to set.</description>
    </item>
    <item>
      <title>GPU index page</title>
      <link>https://modelvram.com/llm-vram-calculator/gpu/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#gpu-index</guid>
      <pubDate>Wed, 30 Sep 2026 12:00:00 GMT</pubDate>
      <description>Every GPU, Mac, CPU-only PC and multi-GPU setup on one page, with how many models each runs.</description>
    </item>
    <item>
      <title>Shows cloud GPU options when a result doesn’t fit</title>
      <link>https://modelvram.com/llm-vram-calculator/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#cloud-gpu-card</guid>
      <pubDate>Wed, 30 Sep 2026 12:00:00 GMT</pubDate>
      <description>When a model does not fit a local card, the calculators name the memory it needs and the cloud GPU class that holds it, with plain, unsponsored links.</description>
    </item>
    <item>
      <title>Load models by their Ollama name</title>
      <link>https://modelvram.com/llm-vram-calculator/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#ollama-names</guid>
      <pubDate>Wed, 30 Sep 2026 12:00:00 GMT</pubDate>
      <description>Type an Ollama library name such as qwen3:30b-a3b into the LLM VRAM calculator to load that model.</description>
    </item>
    <item>
      <title>Apple LensVLM 9B model page</title>
      <link>https://modelvram.com/llm-vram-calculator/lensvlm-9b/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#lensvlm-9b</guid>
      <pubDate>Wed, 30 Sep 2026 12:00:00 GMT</pubDate>
      <description>VRAM for Apple’s LensVLM-9B per quant and context, its GGUF files and mmproj, and what its visual text compression saves in KV cache.</description>
    </item>
    <item>
      <title>Latency and throughput estimates for Laya and Jev alternatives</title>
      <link>https://modelvram.com/laya-vram-requirements/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#decision-latency</guid>
      <pubDate>Wed, 30 Sep 2026 12:00:00 GMT</pubDate>
      <description>How long one decision takes and how many per second a GPU or CPU handles, on the Laya and Jev alternatives pages.</description>
    </item>
    <item>
      <title>Ternary Bonsai 2 27B VRAM page</title>
      <link>https://modelvram.com/ternary-bonsai-2-27b-vram/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#bonsai-2-27b</guid>
      <pubDate>Wed, 30 Sep 2026 12:00:00 GMT</pubDate>
      <description>Real file sizes of the PTQ1_0, PQ2_0 and MLX files and the longest context on 8, 12 and 16 GB GPUs.</description>
    </item>
    <item>
      <title>llama-server --fit preview</title>
      <link>https://modelvram.com/llm-vram-calculator/#fit-preview</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#fit-preview</guid>
      <pubDate>Tue, 29 Sep 2026 12:00:00 GMT</pubDate>
      <description>The LLM VRAM calculator shows what llama-server’s automatic --fit would pick for context, slots and GPU layers, and the explicit flags instead.</description>
    </item>
    <item>
      <title>Chart: every model by the VRAM it needs</title>
      <link>https://modelvram.com/best-llm-for-vram/#chart</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#vram-scatter</guid>
      <pubDate>Tue, 29 Sep 2026 12:00:00 GMT</pubDate>
      <description>A scatter plot of all tracked models by VRAM at Q4_K_M against parameters, with 8–80 GB card lines.</description>
    </item>
    <item>
      <title>2026 Local LLM VRAM Report</title>
      <link>https://modelvram.com/reports/local-llm-vram-2026/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#report-2026</guid>
      <pubDate>Tue, 29 Sep 2026 12:00:00 GMT</pubDate>
      <description>The largest open model per memory tier and the 32K KV cache ranking, with CSV downloads under CC BY 4.0.</description>
    </item>
    <item>
      <title>Image and video model VRAM calculator</title>
      <link>https://modelvram.com/image-video-vram-calculator/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#image-video</guid>
      <pubDate>Tue, 29 Sep 2026 12:00:00 GMT</pubDate>
      <description>Peak VRAM and system RAM for FLUX.2, Wan 2.2, LTX-2 and Qwen-Image 2.1 in ComfyUI, checked against public ComfyUI runs.</description>
    </item>
    <item>
      <title>Jev alternatives you can run locally</title>
      <link>https://modelvram.com/jev-local-alternatives/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#jev-alternatives</guid>
      <pubDate>Tue, 29 Sep 2026 12:00:00 GMT</pubDate>
      <description>Open models that do the API-only Jev’s job on your own GPU or CPU, with file sizes and VRAM.</description>
    </item>
    <item>
      <title>Laya VRAM requirements</title>
      <link>https://modelvram.com/laya-vram-requirements/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#laya</guid>
      <pubDate>Tue, 29 Sep 2026 12:00:00 GMT</pubDate>
      <description>Real file sizes of Laya’s checkpoints in PyTorch, ggmlc and GGUF, and whether a 16 GB or 24 GB GPU or a CPU runs it.</description>
    </item>
    <item>
      <title>Predicted vs measured</title>
      <link>https://modelvram.com/accuracy/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#accuracy</guid>
      <pubDate>Tue, 29 Sep 2026 12:00:00 GMT</pubDate>
      <description>The estimates next to 20 public llama.cpp and vLLM measurements, with the error of each.</description>
    </item>
    <item>
      <title>KV cache counted as llama.cpp allocates it</title>
      <link>https://modelvram.com/llm-vram-calculator/kv-cache/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#kv-cache-llama-cpp</guid>
      <pubDate>Tue, 29 Sep 2026 12:00:00 GMT</pubDate>
      <description>Gemma 4’s global layers and sliding-window caches now match llama.cpp’s logs to the byte.</description>
    </item>
    <item>
      <title>Open dataset and best LLMs per VRAM size</title>
      <link>https://modelvram.com/data/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#open-data</guid>
      <pubDate>Tue, 29 Sep 2026 12:00:00 GMT</pubDate>
      <description>Every model’s VRAM estimates as CSV and JSON under CC BY 4.0, and the models that fit 8 to 96 GB.</description>
    </item>
    <item>
      <title>What can my PC run?</title>
      <link>https://modelvram.com/what-can-i-run/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#what-can-i-run</guid>
      <pubDate>Tue, 29 Sep 2026 12:00:00 GMT</pubDate>
      <description>Detects your GPU in the browser and lists the local LLMs it runs, with the best quantization and speed.</description>
    </item>
    <item>
      <title>vLLM KV cache and concurrency calculator</title>
      <link>https://modelvram.com/vllm-memory-calculator/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#vllm</guid>
      <pubDate>Tue, 29 Sep 2026 12:00:00 GMT</pubDate>
      <description>The KV cache pool vLLM allocates, in tokens, and how many requests fit at once, checked against startup logs.</description>
    </item>
    <item>
      <title>Fine-tuning VRAM calculator</title>
      <link>https://modelvram.com/fine-tuning-vram-calculator/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#fine-tuning</guid>
      <pubDate>Tue, 29 Sep 2026 12:00:00 GMT</pubDate>
      <description>GPU memory for full fine-tuning, LoRA and QLoRA in Transformers or Unsloth, checked against 15 published runs.</description>
    </item>
    <item>
      <title>MoE offload calculator</title>
      <link>https://modelvram.com/moe-offload-calculator/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#moe-offload</guid>
      <pubDate>Mon, 28 Sep 2026 12:00:00 GMT</pubDate>
      <description>The smallest llama.cpp --n-cpu-moe that fits your GPU, from real GGUF tensor sizes.</description>
    </item>
    <item>
      <title>Qwen-Image 2.1 VRAM calculator</title>
      <link>https://modelvram.com/qwen-image-2-1-vram-calculator/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#qwen-image</guid>
      <pubDate>Mon, 28 Sep 2026 12:00:00 GMT</pubDate>
      <description>Peak VRAM for every DiT, text encoder and VAE combination of Qwen-Image 2.1, with real measurements.</description>
    </item>
    <item>
      <title>LLM VRAM and speed calculators</title>
      <link>https://modelvram.com/llm-vram-calculator/</link>
      <guid isPermaLink="true">https://modelvram.com/changelog/#launch</guid>
      <pubDate>Sat, 26 Sep 2026 12:00:00 GMT</pubDate>
      <description>The first two tools: how much GPU memory a model needs from its Hugging Face files, and how fast it writes on your hardware.</description>
    </item>
  </channel>
</rss>
