{"id":9443,"date":"2026-10-03T07:11:56","date_gmt":"2026-10-02T22:11:56","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/03\/localai-v4-11-0-released\/"},"modified":"2026-10-03T07:11:56","modified_gmt":"2026-10-02T22:11:56","slug":"localai-v4-11-0-released","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/03\/localai-v4-11-0-released\/","title":{"rendered":"LocalAI v4.11.0 Released with Failover and NeMo Audio Models"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/github.com\/mudler\/LocalAI\">mudler\/LocalAI<\/a><\/td>\n<\/tr>\n<tr>\n<td>Version<\/td>\n<td><a href=\"https:\/\/github.com\/mudler\/LocalAI\/releases\/tag\/v4.11.0\">v4.11.0<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-10-03<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>MIT<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code when the information was collected. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>Version v4.11.0 of the open-source AI engine &#8220;mudler\/LocalAI&#8221;, written in the Go language and provided under the MIT license, has been released. LocalAI is software that allows local execution of models such as LLMs, vision, audio, image, and video across various hardware without depending on a specific GPU. The change with the largest impact on users in this version is the introduction of the &#8220;model failover chain,&#8221; which groups multiple local and remote models under a single model name and automatically switches between them in the event of a failure. This significantly improves redundancy and availability for local LLM operations.<\/p>\n<h2>Key Changes<\/h2>\n<p><strong>Introduction of Model Failover Chains and localai-proxy<\/strong> It is now possible to configure a prioritized chain of targets (local\/remote) for a single model name. If failures such as communication errors, server errors, out-of-memory (OOM), or rate limits occur before a response is finalized, LocalAI automatically retries with the next target. Alongside this, management APIs such as <code>GET \/api\/failover<\/code>, real-time health monitoring, and notification functions via response headers (<code>X-LocalAI-Served-Model<\/code>) are provided. Additionally, a <code>localai-proxy<\/code> backend has been added to transparently relay requests to other LocalAI endpoints.<\/p>\n<p><strong>Audio Scene Recognition and Speaker Identification\/Registration<\/strong> In the <code>parakeet-cpp<\/code> backend, audio data can now be processed comprehensively as &#8220;scenes&#8221; rather than just transcriptions. Transcription (ASR), speaker diarization, and audio event detection can be executed simultaneously with a single model. Furthermore, a UI for speaker registration has been added to the Studio screen, allowing users to preview clear segments of recorded audio and register them with names. Registered profiles are saved in a shared audio registry and are automatically identified as speaker names (<code>speaker_name<\/code>) in subsequent processing.<\/p>\n<p><strong>Support for Decision Models and the <code>\/v1\/systemone<\/code> API<\/strong> &#8220;Decision models,&#8221; which perform structured selection, scoring, and zero-shot extraction, have been introduced as an official feature. They can be enabled by specifying <code>known_usecases: [decisions]<\/code> in the model configuration and are accessible via the <code>POST \/v1\/systemone<\/code> endpoint. Requests are subject to a 64 KiB body size limit, a maximum of 64 questions, and validation of options and levels. Decision models such as Laya and GLiNER2.5-Decide have also been added to the gallery.<\/p>\n<p><strong>Secure Distribution via Signed OCI Model Galleries<\/strong> Model galleries can now be distributed and fetched directly as OCI artifacts (<code>oci:\/\/host\/repository:tag<\/code>). LocalAI resolves digests from tags, performs signature verification based on Sigstore policies, and securely deploys caches. During deployment, file path isolation and upper limits on layer counts are checked, enabling the construction of reliable galleries without preparing a separate web index.<\/p>\n<p><strong>Kimodo Text Animation (3D Skeletal Motion Generation)<\/strong> The <code>kimodocpp<\/code> backend and <code>POST \/3d\/animate<\/code> endpoint have been added to generate 3D skeletal animations from text. By specifying a UTF-8 text prompt, it generates 30 FPS binary glTF (<code>.glb<\/code>) animation data. It supports frame counts ranging from 60 to 150 frames and runs on CPU and Vulkan environments. Previewing on the Studio screen and local history management are also supported.<\/p>\n<p><strong>Single-Host Operations Management Dashboard (Operate \u2192 This machine)<\/strong> Implements the host management screen &#8220;Operate \u2192 This machine&#8221; for single-node environments that do not use distributed mode. In addition to gauges displaying usage rates for VRAM, RAM, CPU, and model storage, it displays a list of currently running models. From the Web UI, users can check the memory occupancy and PID of each model, view backend logs, and perform individual model stop operations directly.<\/p>\n<h2>Supported Models and Hardware<\/h2>\n<p>In this version, the number of registered items in the model gallery has expanded by 79 from 1,847 to 1,926, newly supporting numerous models across a wide range of fields such as speech recognition\/diarization, decision processing, and text generation. Hardware detection has also been stabilized.<\/p>\n<h3>NeMo Speech Recognition and Diarization Models<\/h3>\n<p>NeMo-derived streaming audio models have been newly added for the <code>nemo-speech-cpp<\/code> backend.<br \/>\n&#8211; <strong>Single Speaker Diarization<\/strong>: <code>nemo-speech-cpp-sortformer-diarization-v2<\/code> (<code>nvidia\/diar_streaming_sortformer_4spk-v2<\/code>) enables streaming speaker diarization for up to 4 speakers via <code>\/v1\/audio\/diarization<\/code>.<br \/>\n&#8211; <strong>Single Speech Recognition (ASR)<\/strong>: Multilingual streaming-supported <code>nemo-speech-cpp-nemotron-3.5-asr-streaming<\/code> (<code>nvidia\/nemotron-3.5-asr-streaming-0.6b<\/code>) and 25-language-supported <code>nemo-speech-cpp-parakeet-tdt-0.6b-v3<\/code> (<code>nvidia\/parakeet-tdt-0.6b-v3<\/code>) are available via <code>\/v1\/audio\/transcriptions<\/code>.<br \/>\n&#8211; <strong>Integration of ASR and Diarization<\/strong>: In <code>nemo-speech-cpp-nemotron-3.5-asr-streaming-diarized<\/code> and <code>nemo-speech-cpp-parakeet-tdt-0.6b-v3-diarized<\/code>, Sortformer is integrated via the <code>diar_model<\/code> option, allowing speaker tags to be attached to transcription word units.<\/p>\n<h3>Piper Text-to-Speech Models (Italian)<\/h3>\n<p>Four high-quality Italian community audio models (ONNX format and configuration files) have been added to the gallery for the Piper backend.<br \/>\n&#8211; <code>voice-it_IT-ugo-medium<\/code><br \/>\n&#8211; <code>voice-it_IT-aurora-medium<\/code><br \/>\n&#8211; <code>voice-it_IT-giorgio-high<\/code><br \/>\n&#8211; <code>voice-it_IT-leonardo-high<\/code><\/p>\n<h3>Decision Models<\/h3>\n<p>Centering on the <code>vllm-cpp<\/code> backend, numerous models handling structured decision tasks have been added.<br \/>\n&#8211; Laya<br \/>\n&#8211; CUA-S1 forms<br \/>\n&#8211; GLiNER2.5 (GLiNER2.5-Decide)<br \/>\n&#8211; Qwen3-VL<br \/>\n&#8211; Tev1<br \/>\n&#8211; kev<br \/>\n&#8211; Nimble<br \/>\n&#8211; CLM<\/p>\n<h3>Expansion and Organization of LLMs and Various Models<\/h3>\n<p>The following GGUF variants and derivative models were registered as text generation and inference models.<br \/>\n&#8211; <strong>Qwen Derivatives and Community Models<\/strong>: NeoHorse (NeoHorse-1-9B, official NeoHorse 4B), Qwen3.8 Distill (Qwen3.8 35B Distill), Qwen3.8 Cyber, ByteShape, Flash Next GSQ-RCO, Occamy (Occamy-1.0), Hy-MT2 (Hy-MT2 7B), Maple Preview (Maple-Preview)<br \/>\n&#8211; <strong>GGUF Models for llama.cpp<\/strong>: Hemmingway-1 (Q4_K_M, Q8_0 formats. Supports 32K context), MiMo Distill Qwen 9B, Sharp-Spark, Swift 1.5 GSQ-RCO, ThinkingCap, Agention, Qwopus Flash V2, Cyber-Tiel-Coder, Cyber-Ornith<\/p>\n<p>Note that <code>qwen-image-2.1-uncensored<\/code>, which was registered as a chat model, was actually diffusion model weights and does not work correctly with llama.cpp, so it has been removed from the gallery. Users are advised to use the bundle for <code>stablediffusion-ggml<\/code> for image generation.<\/p>\n<h3>Hardware Support and Detection Improvements<\/h3>\n<ul>\n<li><strong>AMD APU<\/strong>: The system information acquisition function (<code>xsysinfo<\/code>) was revised so that GTT (Graphics Translation Table) memory is correctly summed up in APU VRAM detection.<\/li>\n<li><strong>Intel GPU<\/strong>: Fixes to the probing process avoid hangs that occurred at startup.<\/li>\n<li><strong>Kimodo (3D Animation)<\/strong>: The added <code>kimodocpp<\/code> backend supports acceleration in Vulkan environments in addition to CPU execution.<\/li>\n<\/ul>\n<h2>How to Get It<\/h2>\n<p>Please check the official documentation and release page for specific installation commands in this release. Acquire the latest container images or update binaries according to your LocalAI environment.<\/p>\n<p>Detailed release notes and downloads for binaries on each platform can be found on the <a href=\"https:\/\/github.com\/mudler\/LocalAI\/releases\/tag\/v4.11.0\">GitHub release page<\/a>.<\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/18\/localai-v4-10-0-released\/\">LocalAI v4.10.0 Released: Fleet Dashboard &amp; M5 Support<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Follow this tool<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-localai-en\/\">LocalAI overview and release history (136 releases tracked)<\/a><\/li>\n<li><strong>Other inference engines and runtimes<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-llama-cpp-en\/\">llama.cpp<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-ollama-en\/\">Ollama<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-vllm-en\/\">vLLM<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li>LocalAI v4.11.0 Release Notes: <a href=\"https:\/\/github.com\/mudler\/LocalAI\/releases\/tag\/v4.11.0\">https:\/\/github.com\/mudler\/LocalAI\/releases\/tag\/v4.11.0<\/a><\/li>\n<li>PR #12121 (feat(gallery): add four Italian community Piper voices): <a href=\"https:\/\/github.com\/mudler\/LocalAI\/pull\/12121\">https:\/\/github.com\/mudler\/LocalAI\/pull\/12121<\/a><\/li>\n<li>PR #12124 (feat(gallery): consolidate 25 pending gallery PRs): <a href=\"https:\/\/github.com\/mudler\/LocalAI\/pull\/12124\">https:\/\/github.com\/mudler\/LocalAI\/pull\/12124<\/a><\/li>\n<li>PR #12221 (batch(gallery): merge 14 gallery model-addition PRs): <a href=\"https:\/\/github.com\/mudler\/LocalAI\/pull\/12221\">https:\/\/github.com\/mudler\/LocalAI\/pull\/12221<\/a><\/li>\n<li>PR #12265 (feat(gallery): add nemo-speech-cpp diarization and ASR models): <a href=\"https:\/\/github.com\/mudler\/LocalAI\/pull\/12265\">https:\/\/github.com\/mudler\/LocalAI\/pull\/12265<\/a><\/li>\n<li>PR #12278 (chore(gallery): add Hemmingway and remove invalid chat entry): <a href=\"https:\/\/github.com\/mudler\/LocalAI\/pull\/12278\">https:\/\/github.com\/mudler\/LocalAI\/pull\/12278<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>LocalAI v4.11.0 introduces model failover chains, NeMo speech-to-text models, decision models, signed OCI galleries, and hardware improvements.<\/p>\n","protected":false},"author":1,"featured_media":9442,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[105],"tags":[163,1507,2918,2920,1547,2922],"class_list":["post-9443","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-engines-and-tools","tag-gguf-en","tag-localai-en","tag-nemo-en","tag-piper-en","tag-verified","tag-vllm-cpp-en"],"lang":"en","translations":{"en":9443,"ja":9441},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9443","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=9443"}],"version-history":[{"count":0,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9443\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/9442"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=9443"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=9443"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=9443"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}