{"id":7565,"date":"2026-09-30T09:08:33","date_gmt":"2026-09-30T00:08:33","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/30\/ollama-v0-35-0-released\/"},"modified":"2026-09-30T21:29:40","modified_gmt":"2026-09-30T12:29:40","slug":"ollama-v0-35-0-released","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/30\/ollama-v0-35-0-released\/","title":{"rendered":"Ollama v0.35.0 Released: New API for Decision Models"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/github.com\/ollama\/ollama\">ollama\/ollama<\/a><\/td>\n<\/tr>\n<tr>\n<td>Version<\/td>\n<td><a href=\"https:\/\/github.com\/ollama\/ollama\/releases\/tag\/v0.35.0\">v0.35.0<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-09-29<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>MIT<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code when the information was collected. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>Ollama v0.35.0 has been released. Ollama is a tool for easily running and managing various open-weight models, such as DeepSeek, Gemma, and Qwen, in local environments.<\/p>\n<p>The biggest change in this update is the addition of a new API endpoint, <code>\/v1\/systemone<\/code>, to support decision models. Unlike traditional text generation, this allows users to directly obtain structured data from models, such as choices, probabilities, and scores, making it easier to integrate them into automated workflows.<\/p>\n<h2>Breaking Changes and Deprecations<\/h2>\n<p>The behavior regarding deprecated parameters has been changed.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Parameter<\/th>\n<th>Change<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><code>typical_p<\/code><\/td>\n<td>If included in a request, a warning log will be output instead of failing with an error<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>This change affects existing users who use the <code>typical_p<\/code> parameter in their requests, but execution can continue. Migration to appropriate parameters is recommended in preparation for future removal.<\/p>\n<h2>Key Changes<\/h2>\n<h3>Introduction of New API for Decision Models &#8220;\/v1\/systemone&#8221;<\/h3>\n<p>A new <code>\/v1\/systemone<\/code> endpoint, based on TypeSafe&#8217;s Jev API, has been introduced. This API is specialized for tasks that require fast and type-defined &#8220;judgments&#8221; rather than text generation. Specific use cases include ticket triage, model routing, and content moderation.<\/p>\n<p>This API supports the following three question types:<\/p>\n<ul>\n<li><code>choice<\/code>: Selects one from the specified choices and returns the probability for each choice.<\/li>\n<li><code>noul<\/code>: Returns the probability that a specific condition is true.<\/li>\n<li><code>score<\/code>: Returns a score based on an ordered set of criteria.<\/li>\n<\/ul>\n<p>Because network latency can be avoided by running locally, extremely fast responses are possible. For example, the <code>nimble<\/code> 9B model running on an M5 Max achieves a low latency of an average of 91ms per decision. Developers can incorporate this feature into their applications using direct requests with <code>curl<\/code> or the official Python SDK provided by TypeSafe.<\/p>\n<h3>Expansion of Supported Models<\/h3>\n<p>New models specialized for decision-making are now available. These can be downloaded and used immediately with the <code>ollama pull<\/code> command.<\/p>\n<ul>\n<li><code>nimble<\/code>: A 9B parameter open-source decision model developed by Bespoke Labs.<\/li>\n<li><code>tev1<\/code>: Experimental decision models in 4B and 0.8B sizes provided by Together AI.<\/li>\n<\/ul>\n<h3>Improved User Experience and Stability<\/h3>\n<p>Several fixes have been made to enhance user convenience and system stability.<\/p>\n<p>First, the behavior of the Settings screen has been improved. Previously, it was necessary to wait for model detection, but users can now open the settings screen immediately without waiting for model detection.<\/p>\n<p>Additionally, bugs on macOS have been fixed. Specifically, an issue where available updates were not reflected in the update menu or icon upon application startup has been resolved. Furthermore, an issue where processing would stop and hang indefinitely while downloading MLX models has also been fixed, making usage in Apple Silicon environments more stable.<\/p>\n<h2>Hardware Requirements<\/h2>\n<p>This release supports a new set of models specialized for decision-making tasks. Unlike general chat models, these are designed to quickly perform structured judgments. The following models are currently available through Ollama:<\/p>\n<ul>\n<li><strong>nimble<\/strong>: A 9B parameter open-source decision model developed by Bespoke Labs. High accuracy has been confirmed in 3,880 judgment tests using 13 public datasets.<\/li>\n<li><strong>tev1<\/strong>: A 4B parameter experimental decision model by Together AI.<\/li>\n<li><strong>tev1:0.8b<\/strong>: An even lighter experimental model with 0.8B parameters by Together AI.<\/li>\n<\/ul>\n<p>On the hardware front, optimization particularly for Apple Silicon is progressing. Benchmarks using a MacBook Pro M5 Max report that the nimble 9B model can make decisions with an extremely low latency of an average of 91ms. This is fast enough to handle in-game decisions requiring real-time performance or instant processing of large volumes of content moderation.<\/p>\n<p>Additionally, while standard resources are currently used, future updates plan further performance improvements for Apple Silicon leveraging MLX (Apple&#8217;s machine learning framework). This is expected to further improve inference speeds in Mac environments.<\/p>\n<h2>How to Get It<\/h2>\n<p>First, update Ollama itself to the latest version (v0.35.0 or higher). Then, run the following command in your terminal to download the new decision model:<\/p>\n<pre><code class=\"language-sh\">ollama pull nimble\n<\/code><\/pre>\n<p>To use the new <code>\/v1\/systemone<\/code> endpoint from a Python environment, it is recommended to install and use TypeSafe&#8217;s official SDK. You can install it with the following command:<\/p>\n<pre><code class=\"language-sh\"># When using uv\nuv add typesafe-sdk\n\n# When using pip\npip install typesafe-sdk\n<\/code><\/pre>\n<p>To connect to local Ollama using the SDK, set the following environment variables:<\/p>\n<pre><code class=\"language-sh\">export TYPESAFE_BASE_URL=http:\/\/localhost:11434\nexport TYPESAFE_API_KEY=ollama\nexport TYPESAFE_DEFAULT_MODEL=nimble\n<\/code><\/pre>\n<p>A basic implementation example for calling from Python code is as follows. In this example, determining the responsible team, checking for refund requests, and scoring urgency based on the ticket content are performed in a single request.<\/p>\n<pre><code class=\"language-python\">from typesafe_sdk import Choice, Noul, Score, TypeSafeClient\n\nticket = &quot;I was charged twice. Please refund the extra payment.&quot;\nquestions = {\n    &quot;team&quot;: Choice(\n        instructions=&quot;Which team should handle this ticket?&quot;,\n        criteria={\n            &quot;billing&quot;: &quot;Payments and refunds&quot;,\n            &quot;technical&quot;: &quot;Bugs and integrations&quot;,\n            &quot;other&quot;: &quot;None of the above&quot;,\n        },\n    ),\n    &quot;refund&quot;: Noul(\n        instructions=&quot;Does the customer explicitly ask for a refund?&quot;,\n    ),\n    &quot;urgency&quot;: Score(\n        instructions=&quot;How urgent is this ticket?&quot;,\n        criteria=[&quot;Routine&quot;, &quot;Soon&quot;, &quot;Urgent&quot;],\n    ),\n}\n\nwith TypeSafeClient(timeout=120) as client:\n    result = client.system_one(\n        state={&quot;ticket&quot;: ticket},\n        questions=questions,\n    )\n    print(result.choices[&quot;team&quot;].choice) # billing\n    print(result.nouls[&quot;refund&quot;].noul)   # 0.997\n    print(result.scores[&quot;urgency&quot;].score) # 0.815\n<\/code><\/pre>\n<p><!-- lmw:releases --><\/p>\n<h2>Releases Since Our Last Article<\/h2>\n<p><em>Compiled by Local Model Watch from the project&#8217;s GitHub releases: the versions between this release and the last one we covered, which did not get separate articles.<\/em> <em>Full history: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-ollama-en\/\">release tracker<\/a>.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Version<\/th>\n<th>Released<\/th>\n<th>Release notes<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>v0.34.4<\/td>\n<td>2026-09-23<\/td>\n<td><a href=\"https:\/\/github.com\/ollama\/ollama\/releases\/tag\/v0.34.4\">GitHub<\/a><\/td>\n<\/tr>\n<tr>\n<td>v0.34.3<\/td>\n<td>2026-09-19<\/td>\n<td><a href=\"https:\/\/github.com\/ollama\/ollama\/releases\/tag\/v0.34.3\">GitHub<\/a><\/td>\n<\/tr>\n<tr>\n<td>v0.34.2<\/td>\n<td>2026-09-16<\/td>\n<td><a href=\"https:\/\/github.com\/ollama\/ollama\/releases\/tag\/v0.34.2\">GitHub<\/a><\/td>\n<\/tr>\n<tr>\n<td>v0.34.1<\/td>\n<td>2026-09-15<\/td>\n<td><a href=\"https:\/\/github.com\/ollama\/ollama\/releases\/tag\/v0.34.1\">GitHub<\/a><\/td>\n<\/tr>\n<tr>\n<td>v0.34.0<\/td>\n<td>2026-09-06<\/td>\n<td><a href=\"https:\/\/github.com\/ollama\/ollama\/releases\/tag\/v0.34.0\">GitHub<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:releases --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Follow this tool<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-ollama-en\/\">Ollama overview and release history (200 releases tracked)<\/a><\/li>\n<li><strong>Other inference engines and runtimes<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-llama-cpp-en\/\">llama.cpp<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-vllm-en\/\">vLLM<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-sglang-en\/\">SGLang<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/github.com\/ollama\/ollama\/releases\/tag\/v0.35.0\">https:\/\/github.com\/ollama\/ollama\/releases\/tag\/v0.35.0<\/a><\/li>\n<li><a href=\"https:\/\/ollama.com\/blog\/ollama-now-supports-jev-style-decision-models\">https:\/\/ollama.com\/blog\/ollama-now-supports-jev-style-decision-models<\/a><\/li>\n<\/ul>\n<p><!-- lmw:updates --><\/p>\n<h2>Update History<\/h2>\n<ul>\n<li>2026-09-30: Rewrote the article from re-collected sources.<\/li>\n<\/ul>\n<p><!-- \/lmw:updates --><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Ollama v0.35.0 is released with a new \/v1\/systemone API for decision models, supporting nimble and tev1 models for fast local inference.<\/p>\n","protected":false},"author":1,"featured_media":7564,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[105],"tags":[2072,136,2719,1093,2525,1547],"class_list":["post-7565","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-engines-and-tools","tag-api-en","tag-llm-en","tag-nimble-en","tag-ollama-en","tag-tev1-en","tag-verified"],"lang":"en","translations":{"en":7565,"ja":7563},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/7565","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=7565"}],"version-history":[{"count":4,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/7565\/revisions"}],"predecessor-version":[{"id":8290,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/7565\/revisions\/8290"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/7564"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=7565"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=7565"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=7565"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}