Ollama v0.35.0 Released: New API for Decision Models

At a Glance
| Item | Value |
|---|---|
| Repository | ollama/ollama |
| Version | v0.35.0 |
| Published | 2026-09-29 |
| License | MIT |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code when the information was collected. Dates are JST.
Overview
Ollama v0.35.0 has been released. Ollama is a tool for easily running and managing various open-weight models, such as DeepSeek, Gemma, and Qwen, in local environments.
The biggest change in this update is the addition of a new API endpoint, /v1/systemone, to support decision models. Unlike traditional text generation, this allows users to directly obtain structured data from models, such as choices, probabilities, and scores, making it easier to integrate them into automated workflows.
Breaking Changes and Deprecations
The behavior regarding deprecated parameters has been changed.
| Parameter | Change |
|---|---|
typical_p |
If included in a request, a warning log will be output instead of failing with an error |
This change affects existing users who use the typical_p parameter in their requests, but execution can continue. Migration to appropriate parameters is recommended in preparation for future removal.
Key Changes
Introduction of New API for Decision Models “/v1/systemone"
A new /v1/systemone endpoint, based on TypeSafe’s Jev API, has been introduced. This API is specialized for tasks that require fast and type-defined “judgments" rather than text generation. Specific use cases include ticket triage, model routing, and content moderation.
This API supports the following three question types:
choice: Selects one from the specified choices and returns the probability for each choice.noul: Returns the probability that a specific condition is true.score: Returns a score based on an ordered set of criteria.
Because network latency can be avoided by running locally, extremely fast responses are possible. For example, the nimble 9B model running on an M5 Max achieves a low latency of an average of 91ms per decision. Developers can incorporate this feature into their applications using direct requests with curl or the official Python SDK provided by TypeSafe.
Expansion of Supported Models
New models specialized for decision-making are now available. These can be downloaded and used immediately with the ollama pull command.
nimble: A 9B parameter open-source decision model developed by Bespoke Labs.tev1: Experimental decision models in 4B and 0.8B sizes provided by Together AI.
Improved User Experience and Stability
Several fixes have been made to enhance user convenience and system stability.
First, the behavior of the Settings screen has been improved. Previously, it was necessary to wait for model detection, but users can now open the settings screen immediately without waiting for model detection.
Additionally, bugs on macOS have been fixed. Specifically, an issue where available updates were not reflected in the update menu or icon upon application startup has been resolved. Furthermore, an issue where processing would stop and hang indefinitely while downloading MLX models has also been fixed, making usage in Apple Silicon environments more stable.
Hardware Requirements
This release supports a new set of models specialized for decision-making tasks. Unlike general chat models, these are designed to quickly perform structured judgments. The following models are currently available through Ollama:
- nimble: A 9B parameter open-source decision model developed by Bespoke Labs. High accuracy has been confirmed in 3,880 judgment tests using 13 public datasets.
- tev1: A 4B parameter experimental decision model by Together AI.
- tev1:0.8b: An even lighter experimental model with 0.8B parameters by Together AI.
On the hardware front, optimization particularly for Apple Silicon is progressing. Benchmarks using a MacBook Pro M5 Max report that the nimble 9B model can make decisions with an extremely low latency of an average of 91ms. This is fast enough to handle in-game decisions requiring real-time performance or instant processing of large volumes of content moderation.
Additionally, while standard resources are currently used, future updates plan further performance improvements for Apple Silicon leveraging MLX (Apple’s machine learning framework). This is expected to further improve inference speeds in Mac environments.
How to Get It
First, update Ollama itself to the latest version (v0.35.0 or higher). Then, run the following command in your terminal to download the new decision model:
ollama pull nimble
To use the new /v1/systemone endpoint from a Python environment, it is recommended to install and use TypeSafe’s official SDK. You can install it with the following command:
# When using uv
uv add typesafe-sdk
# When using pip
pip install typesafe-sdk
To connect to local Ollama using the SDK, set the following environment variables:
export TYPESAFE_BASE_URL=http://localhost:11434
export TYPESAFE_API_KEY=ollama
export TYPESAFE_DEFAULT_MODEL=nimble
A basic implementation example for calling from Python code is as follows. In this example, determining the responsible team, checking for refund requests, and scoring urgency based on the ticket content are performed in a single request.
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
ticket = "I was charged twice. Please refund the extra payment."
questions = {
"team": Choice(
instructions="Which team should handle this ticket?",
criteria={
"billing": "Payments and refunds",
"technical": "Bugs and integrations",
"other": "None of the above",
},
),
"refund": Noul(
instructions="Does the customer explicitly ask for a refund?",
),
"urgency": Score(
instructions="How urgent is this ticket?",
criteria=["Routine", "Soon", "Urgent"],
),
}
with TypeSafeClient(timeout=120) as client:
result = client.system_one(
state={"ticket": ticket},
questions=questions,
)
print(result.choices["team"].choice) # billing
print(result.nouls["refund"].noul) # 0.997
print(result.scores["urgency"].score) # 0.815
Releases Since Our Last Article
Compiled by Local Model Watch from the project’s GitHub releases: the versions between this release and the last one we covered, which did not get separate articles. Full history: release tracker.
| Version | Released | Release notes |
|---|---|---|
| v0.34.4 | 2026-09-23 | GitHub |
| v0.34.3 | 2026-09-19 | GitHub |
| v0.34.2 | 2026-09-16 | GitHub |
| v0.34.1 | 2026-09-15 | GitHub |
| v0.34.0 | 2026-09-06 | GitHub |
What to Read Next
- Follow this tool → Ollama overview and release history (200 releases tracked)
- Other inference engines and runtimes → llama.cpp / vLLM / SGLang
Sources
- https://github.com/ollama/ollama/releases/tag/v0.35.0
- https://ollama.com/blog/ollama-now-supports-jev-style-decision-models
Update History
- 2026-09-30: Rewrote the article from re-collected sources.

