LocalAI v4.11.0 Released with Failover and NeMo Audio Models

LocalAI v4.11.0 Released with Failover and NeMo Audio Models

At a Glance

Item Value
Repository mudler/LocalAI
Version v4.11.0
Published 2026-10-03
License MIT
Source type Primary source (the publisher itself)

Values determined by this site’s code when the information was collected. Dates are JST.

Overview

Version v4.11.0 of the open-source AI engine “mudler/LocalAI", written in the Go language and provided under the MIT license, has been released. LocalAI is software that allows local execution of models such as LLMs, vision, audio, image, and video across various hardware without depending on a specific GPU. The change with the largest impact on users in this version is the introduction of the “model failover chain," which groups multiple local and remote models under a single model name and automatically switches between them in the event of a failure. This significantly improves redundancy and availability for local LLM operations.

Key Changes

Introduction of Model Failover Chains and localai-proxy It is now possible to configure a prioritized chain of targets (local/remote) for a single model name. If failures such as communication errors, server errors, out-of-memory (OOM), or rate limits occur before a response is finalized, LocalAI automatically retries with the next target. Alongside this, management APIs such as GET /api/failover, real-time health monitoring, and notification functions via response headers (X-LocalAI-Served-Model) are provided. Additionally, a localai-proxy backend has been added to transparently relay requests to other LocalAI endpoints.

Audio Scene Recognition and Speaker Identification/Registration In the parakeet-cpp backend, audio data can now be processed comprehensively as “scenes" rather than just transcriptions. Transcription (ASR), speaker diarization, and audio event detection can be executed simultaneously with a single model. Furthermore, a UI for speaker registration has been added to the Studio screen, allowing users to preview clear segments of recorded audio and register them with names. Registered profiles are saved in a shared audio registry and are automatically identified as speaker names (speaker_name) in subsequent processing.

Support for Decision Models and the /v1/systemone API “Decision models," which perform structured selection, scoring, and zero-shot extraction, have been introduced as an official feature. They can be enabled by specifying known_usecases: [decisions] in the model configuration and are accessible via the POST /v1/systemone endpoint. Requests are subject to a 64 KiB body size limit, a maximum of 64 questions, and validation of options and levels. Decision models such as Laya and GLiNER2.5-Decide have also been added to the gallery.

Secure Distribution via Signed OCI Model Galleries Model galleries can now be distributed and fetched directly as OCI artifacts (oci://host/repository:tag). LocalAI resolves digests from tags, performs signature verification based on Sigstore policies, and securely deploys caches. During deployment, file path isolation and upper limits on layer counts are checked, enabling the construction of reliable galleries without preparing a separate web index.

Kimodo Text Animation (3D Skeletal Motion Generation) The kimodocpp backend and POST /3d/animate endpoint have been added to generate 3D skeletal animations from text. By specifying a UTF-8 text prompt, it generates 30 FPS binary glTF (.glb) animation data. It supports frame counts ranging from 60 to 150 frames and runs on CPU and Vulkan environments. Previewing on the Studio screen and local history management are also supported.

Single-Host Operations Management Dashboard (Operate → This machine) Implements the host management screen “Operate → This machine" for single-node environments that do not use distributed mode. In addition to gauges displaying usage rates for VRAM, RAM, CPU, and model storage, it displays a list of currently running models. From the Web UI, users can check the memory occupancy and PID of each model, view backend logs, and perform individual model stop operations directly.

Supported Models and Hardware

In this version, the number of registered items in the model gallery has expanded by 79 from 1,847 to 1,926, newly supporting numerous models across a wide range of fields such as speech recognition/diarization, decision processing, and text generation. Hardware detection has also been stabilized.

NeMo Speech Recognition and Diarization Models

NeMo-derived streaming audio models have been newly added for the nemo-speech-cpp backend.
– Single Speaker Diarization: nemo-speech-cpp-sortformer-diarization-v2 (nvidia/diar_streaming_sortformer_4spk-v2) enables streaming speaker diarization for up to 4 speakers via /v1/audio/diarization.
– Single Speech Recognition (ASR): Multilingual streaming-supported nemo-speech-cpp-nemotron-3.5-asr-streaming (nvidia/nemotron-3.5-asr-streaming-0.6b) and 25-language-supported nemo-speech-cpp-parakeet-tdt-0.6b-v3 (nvidia/parakeet-tdt-0.6b-v3) are available via /v1/audio/transcriptions.
– Integration of ASR and Diarization: In nemo-speech-cpp-nemotron-3.5-asr-streaming-diarized and nemo-speech-cpp-parakeet-tdt-0.6b-v3-diarized, Sortformer is integrated via the diar_model option, allowing speaker tags to be attached to transcription word units.

Piper Text-to-Speech Models (Italian)

Four high-quality Italian community audio models (ONNX format and configuration files) have been added to the gallery for the Piper backend.
– voice-it_IT-ugo-medium
– voice-it_IT-aurora-medium
– voice-it_IT-giorgio-high
– voice-it_IT-leonardo-high

Decision Models

Centering on the vllm-cpp backend, numerous models handling structured decision tasks have been added.
– Laya
– CUA-S1 forms
– GLiNER2.5 (GLiNER2.5-Decide)
– Qwen3-VL
– Tev1
– kev
– Nimble
– CLM

Expansion and Organization of LLMs and Various Models

The following GGUF variants and derivative models were registered as text generation and inference models.
– Qwen Derivatives and Community Models: NeoHorse (NeoHorse-1-9B, official NeoHorse 4B), Qwen3.8 Distill (Qwen3.8 35B Distill), Qwen3.8 Cyber, ByteShape, Flash Next GSQ-RCO, Occamy (Occamy-1.0), Hy-MT2 (Hy-MT2 7B), Maple Preview (Maple-Preview)
– GGUF Models for llama.cpp: Hemmingway-1 (Q4_K_M, Q8_0 formats. Supports 32K context), MiMo Distill Qwen 9B, Sharp-Spark, Swift 1.5 GSQ-RCO, ThinkingCap, Agention, Qwopus Flash V2, Cyber-Tiel-Coder, Cyber-Ornith

Note that qwen-image-2.1-uncensored, which was registered as a chat model, was actually diffusion model weights and does not work correctly with llama.cpp, so it has been removed from the gallery. Users are advised to use the bundle for stablediffusion-ggml for image generation.

Hardware Support and Detection Improvements

  • AMD APU: The system information acquisition function (xsysinfo) was revised so that GTT (Graphics Translation Table) memory is correctly summed up in APU VRAM detection.
  • Intel GPU: Fixes to the probing process avoid hangs that occurred at startup.
  • Kimodo (3D Animation): The added kimodocpp backend supports acceleration in Vulkan environments in addition to CPU execution.

How to Get It

Please check the official documentation and release page for specific installation commands in this release. Acquire the latest container images or update binaries according to your LocalAI environment.

Detailed release notes and downloads for binaries on each platform can be found on the GitHub release page.

Related Articles

What to Read Next

Sources