EXL2 / EXL3 Model Format Explained: Supported Engines and Models

What Are EXL2 and EXL3?

EXL2 and EXL3 are quantization formats created by turboderp, the developer of the ExLlama inference libraries. EXL2 is loaded by ExLlamaV2 and its successor EXL3 by ExLlamaV3. On Hugging Face they appear as repositories named with exl2 or exl3 and an average bit width such as 4.0bpw.

Why It Matters

  • Fine-grained bit widths. EXL2 mixes 2- to 8-bit layers so the average can be any value such as 2.5 bpw or 4.65 bpw. The main appeal is sizing a model to fit your GPU memory exactly.
  • EXL3 builds on newer quantization research. The ExLlamaV3 README describes EXL3 as a streamlined variant of QTIP from Cornell RelaxML, aiming to lose less quality at the same bit width.
  • Designed for consumer NVIDIA GPUs. The goal is running large models fast on one or a few GPUs.

Tips for Running It Locally

  • It requires an NVIDIA GPU (CUDA). It does not run on CPU or Mac; choose GGUF or MLX there.
  • For a server, use TabbyAPI. The ExLlamaV3 README names TabbyAPI, which provides an OpenAI-compatible API, as the official and recommended backend. For a UI, TextGen (text-generation-webui) also supports ExLlamaV3.
  • EXL2 and EXL3 are different formats. If you are choosing now, prefer EXL3, where development has moved. Ollama and LM Studio (llama.cpp-based) load neither.
  • Repositories often keep each bit width on its own branch. Specify the branch (or separate repository) for the bit width you want when downloading.

Sources: the turboderp-org/exllamav3 README, the turboderp-org/exllamav2 README and the oobabooga/textgen README (all as of 2026-09-27).

Our Coverage and Data

Local Model Watch has published 1 article(s) on models available in EXL2 / EXL3: 1 where the repository itself is in EXL2 / EXL3, and 0 where we found a EXL2 / EXL3 build of the model. The lists below only include builds we have checked (the publisher’s organization and well-known quantizers); a model missing here may still have a EXL2 / EXL3 build elsewhere. Part of our model format index.

Main Engines That Load This Format

Engine Overview
ExLlamaV3 Inference library tuned for consumer NVIDIA GPUs. Its EXL3 quantization format fits large models into limited VRAM.
TextGen Desktop app for local LLMs (formerly text-generation-webui). Switches between backends such as llama.cpp, ExLlamaV3, Transformers and TensorRT-LLM, exposes OpenAI- and Anthropic-compatible APIs, and can also run as a browser-based web UI.

Models Available in EXL2 / EXL3

→ Scroll horizontally to see all columns

Published Model Where to get it Quantizations Smallest VRAM tier Article
2026-09-28 orcarouter/OrcaSAQ-2-27B This repository — 16GB OrcaSAQ-2-27B Text Generation Model: 16GB+ VRAM

“Smallest VRAM tier" is the smallest tier in the requirements table of each article (for other builds, of those builds; for image, video and audio models, always the article’s own table, which counts every component). Leave headroom for context length.

Last updated 2026-09-28 (JST). The explanation at the top of this page was written with the help of AI from the primary sources it cites. The tables under “Our Coverage and Data" are assembled by code from our article log.