Ollama v0.35.0 Released: New API for Decision Models

Ollama v0.35.0 Released with New Decision Models Feature

At a Glance

Item Value
Repository ollama/ollama
Version v0.35.0
Published 2026-09-29
License MIT
Source type Primary source (the publisher itself)

Values determined by this site’s code when the information was collected. Dates are JST.

Overview

Ollama v0.35.0 has been released. Ollama is a tool for easily running and managing various open-weight models, such as DeepSeek, Gemma, and Qwen, in local environments.

The biggest change in this update is the addition of a new API endpoint, /v1/systemone, to support decision models. Unlike traditional text generation, this allows users to directly obtain structured data from models, such as choices, probabilities, and scores, making it easier to integrate them into automated workflows.

Breaking Changes and Deprecations

The behavior regarding deprecated parameters has been changed.

Parameter Change
typical_p If included in a request, a warning log will be output instead of failing with an error

This change affects existing users who use the typical_p parameter in their requests, but execution can continue. Migration to appropriate parameters is recommended in preparation for future removal.

Key Changes

Introduction of New API for Decision Models “/v1/systemone"

A new /v1/systemone endpoint, based on TypeSafe’s Jev API, has been introduced. This API is specialized for tasks that require fast and type-defined “judgments" rather than text generation. Specific use cases include ticket triage, model routing, and content moderation.

This API supports the following three question types:

  • choice: Selects one from the specified choices and returns the probability for each choice.
  • noul: Returns the probability that a specific condition is true.
  • score: Returns a score based on an ordered set of criteria.

Because network latency can be avoided by running locally, extremely fast responses are possible. For example, the nimble 9B model running on an M5 Max achieves a low latency of an average of 91ms per decision. Developers can incorporate this feature into their applications using direct requests with curl or the official Python SDK provided by TypeSafe.

Expansion of Supported Models

New models specialized for decision-making are now available. These can be downloaded and used immediately with the ollama pull command.

  • nimble: A 9B parameter open-source decision model developed by Bespoke Labs.
  • tev1: Experimental decision models in 4B and 0.8B sizes provided by Together AI.

Improved User Experience and Stability

Several fixes have been made to enhance user convenience and system stability.

First, the behavior of the Settings screen has been improved. Previously, it was necessary to wait for model detection, but users can now open the settings screen immediately without waiting for model detection.

Additionally, bugs on macOS have been fixed. Specifically, an issue where available updates were not reflected in the update menu or icon upon application startup has been resolved. Furthermore, an issue where processing would stop and hang indefinitely while downloading MLX models has also been fixed, making usage in Apple Silicon environments more stable.

Hardware Requirements

This release supports a new set of models specialized for decision-making tasks. Unlike general chat models, these are designed to quickly perform structured judgments. The following models are currently available through Ollama:

  • nimble: A 9B parameter open-source decision model developed by Bespoke Labs. High accuracy has been confirmed in 3,880 judgment tests using 13 public datasets.
  • tev1: A 4B parameter experimental decision model by Together AI.
  • tev1:0.8b: An even lighter experimental model with 0.8B parameters by Together AI.

On the hardware front, optimization particularly for Apple Silicon is progressing. Benchmarks using a MacBook Pro M5 Max report that the nimble 9B model can make decisions with an extremely low latency of an average of 91ms. This is fast enough to handle in-game decisions requiring real-time performance or instant processing of large volumes of content moderation.

Additionally, while standard resources are currently used, future updates plan further performance improvements for Apple Silicon leveraging MLX (Apple’s machine learning framework). This is expected to further improve inference speeds in Mac environments.

How to Get It

First, update Ollama itself to the latest version (v0.35.0 or higher). Then, run the following command in your terminal to download the new decision model:

ollama pull nimble

To use the new /v1/systemone endpoint from a Python environment, it is recommended to install and use TypeSafe’s official SDK. You can install it with the following command:

# When using uv
uv add typesafe-sdk

# When using pip
pip install typesafe-sdk

To connect to local Ollama using the SDK, set the following environment variables:

export TYPESAFE_BASE_URL=http://localhost:11434
export TYPESAFE_API_KEY=ollama
export TYPESAFE_DEFAULT_MODEL=nimble

A basic implementation example for calling from Python code is as follows. In this example, determining the responsible team, checking for refund requests, and scoring urgency based on the ticket content are performed in a single request.

from typesafe_sdk import Choice, Noul, Score, TypeSafeClient

ticket = "I was charged twice. Please refund the extra payment."
questions = {
    "team": Choice(
        instructions="Which team should handle this ticket?",
        criteria={
            "billing": "Payments and refunds",
            "technical": "Bugs and integrations",
            "other": "None of the above",
        },
    ),
    "refund": Noul(
        instructions="Does the customer explicitly ask for a refund?",
    ),
    "urgency": Score(
        instructions="How urgent is this ticket?",
        criteria=["Routine", "Soon", "Urgent"],
    ),
}

with TypeSafeClient(timeout=120) as client:
    result = client.system_one(
        state={"ticket": ticket},
        questions=questions,
    )
    print(result.choices["team"].choice) # billing
    print(result.nouls["refund"].noul)   # 0.997
    print(result.scores["urgency"].score) # 0.815

Releases Since Our Last Article

Compiled by Local Model Watch from the project’s GitHub releases: the versions between this release and the last one we covered, which did not get separate articles. Full history: release tracker.

Version Released Release notes
v0.34.4 2026-09-23 GitHub
v0.34.3 2026-09-19 GitHub
v0.34.2 2026-09-16 GitHub
v0.34.1 2026-09-15 GitHub
v0.34.0 2026-09-06 GitHub

What to Read Next

Sources

Update History

  • 2026-09-30: Rewrote the article from re-collected sources.