OpenJev-GGUF 27.4B Decision-Making Specialized Model: 24GB+ VRAM

OpenJev-GGUF 27.4B Decision-Making Specialized Model: 24GB+ VRAM

At a Glance

Item Value
Repository ggml-org/OpenJev-GGUF
Published 2026-10-02
License cc-by-nc-4.0
Formats GGUF
Source type Primary source (the publisher itself)

Values determined by this site’s code when the information was collected. Dates are JST.

Overview

ggml-org has released “ggml-org/OpenJev-GGUF", a GGUF quantized version of “OpenJev", an open-weight model specialized in decision-making. This model is designed to accept text, web pages, and screenshots as inputs to perform choice-based decision-making (decision-model). Instead of generating free text and parsing it for user-requested questions in JSON format, it features a mechanism that directly outputs probabilities or scores for each choice in a single forward pass.

Specifications

The specifications of the original model openjev/openjev are as follows:
– Parameters: 27.4B
– Architecture: Qwen3_5ForConditionalGeneration (qwen3_5)
– Context length: Up to 16,384 tokens

Performance

Here are the benchmark results for the original model openjev/openjev published by its creators. Since this is a GGUF quantized version, the following figures are measured values on the unquantized original model.

First, here is a comparison with other models across 10,000 text-based questions (using 34 public sources):

model accuracy
Jev (hosted API) 85.4% (8,540 of 10,000)
OpenJev 84.0% (8,403 of 10,000)
the same base model before tuning, same readout, its own calibration 80.4% (8,036 of 10,000)
Nimble 9B (open) 75.7% (7,574 of 10,000)

From this table, we can see that OpenJev achieves an accuracy of 84.0%, significantly outperforming the pre-tuning base model (80.4%) and the open model Nimble 9B (75.7%). Meanwhile, while it fell slightly short of the hosted API Jev (hosted API) at 85.4%, it comes very close with a small difference of 1.4 points.

Next is the breakdown of accuracy by category:

→ Scroll horizontally to see all columns

kind of work questions OpenJev Jev (hosted API) before tuning
intent / routing / topic 1,218 92.8% 92.8% 93.2%
sentiment / stance 1,214 82.1% 81.5% 78.3%
spam / hate 695 80.4% 81.2% 80.1%
legal 1,305 78.9% 78.8% 77.8%
ethics / policy judgement 877 78.1% 75.8% 65.5%
commonsense reasoning 2,260 85.8% 88.3% 80.3%
science / facts / claims 1,558 82.5% 89.0% 81.0%
reading + language 873 89.2% 89.3% 83.3%

Looking at these category results, OpenJev shows high accuracy in tasks such as “intent / routing / topic" (92.8%) and “reading + language" (89.2%). It also scores higher than the hosted API Jev in “ethics / policy judgement" (78.1%) and “sentiment / stance" (82.1%). However, in fields like “science / facts / claims" (82.5%) and “commonsense reasoning" (85.8%), it falls slightly behind the hosted API Jev (89.0% and 88.3% respectively).

Additionally, the comparison results before and after tuning on agent-based decision-making tasks are as follows:

test steps before tuning OpenJev
desktop: next action from a screenshot 2,000 76.5% 88.0%
web: next action on unseen websites 975 68.5% 87.4%
web: next action on unseen domains 1,000 65.7% 84.5%
answer flips when the options are shuffled 2,000 18.5% 2.3%

From these results, OpenJev achieves a significant accuracy improvement of over 10 points compared to pre-tuning when determining next actions for desktop operations from screenshots or on unseen websites and domains. Furthermore, answer flips when options are shuffled decreased drastically from 18.5% to 2.3%, demonstrating extremely high stability as a decision-making model.

Comparisons regarding multilingual support and long-context reading comprehension are as follows:

test questions before tuning OpenJev
inference in German, French, Hindi, Chinese (XNLI) 240 72.5% 82.5%
intent in German, French, Hindi, Japanese (MASSIVE) 240 80.4% 85.8%
questions about 2,600 to 8,500-token articles (QuALITY) 120 91.7% 91.7%

This table indicates that accuracy has improved over pre-tuning in multilingual tasks involving German, French, Hindi, Chinese, and Japanese. On the other hand, in question tasks regarding long articles (QuALITY), the score remains unchanged at 91.7% before and after tuning, maintaining equivalent performance.

Finally, here is the number of completed tasks when executing end-to-end browser tasks (100 types of MiniWoB tasks, 1 trial each):

model tasks completed
OpenJev 39
Jev (hosted API) 39
the same base model before tuning 38

These results confirm that in end-to-end browser tasks, OpenJev completes 39 tasks, matching the hosted Jev API and showing a slight improvement from the 38 tasks completed before tuning.

Strengths and Use Cases

This model is a “decision model" that accepts text or images (screenshots) as inputs and makes optimal decisions from pre-defined choices. Unlike general chat bots that generate free text before parsing, it calculates probabilities directly from token scores at the initial output position, enabling fast and reliable processing.

Based on the performance and features of the original model openjev/openjev, it is highly suited for the following applications and tasks:

  • Routing and Triage:
    Determining the intent, topic, responsible team, priority, or language of user inquiry messages to automatically route them to the appropriate destination.

  • Moderation and Safety Judgment:
    Evaluating whether input text falls under harmful content, spam, hate speech, policy violations, or legal threats.

  • LLM as a Judge:
    Objectively determining whether a generative AI model’s output is based on reliable evidence (presence of hallucinations), follows provided evaluation criteria (rubrics), or which of multiple answers is superior.

  • Business Processes and Document Management:
    Automating routine judgment tasks such as approving, withholding, or rejecting invoices, determining the severity of system alerts, or classifying specific clause types in contracts.

  • Browser and Desktop Agents:
    Reading screen screenshots, HTML (DOM structure), or JSON data to determine the next element to interact with or action to execute (such as clicks or input). It can also judge whether a task is complete or blocked.

  • Scoring:
    Calculating ordered step-by-step evaluations accompanied by expected values or confidence levels.

Furthermore, because this model allows dynamic specification of choices (labels) per request, task-specific pre-training or fixed labeling is unnecessary. It can process up to 52 choices in a single request during a single forward pass, and can simultaneously handle multiple questions (e.g., routing, sentiment analysis, urgency) in one request. A major strength is its support for multiple languages including Japanese (English, German, French, Hindi, Chinese, and Japanese).

How It Differs from Similar Models

In a previous article introduced on our site, Decision-Making Specialized Model “OpenJev" Testing GGUF Model Released, we covered “tinyopenjev-for-testing-gguf", also released by ggml-org. However, that was an ultra-small test-only model (approx. 34M parameters) created solely to verify whether inference engines and loaders function properly, and its output content carried no meaning.

In contrast, the newly released “OpenJev-GGUF" is a practical quantized model that retains the entire structure of the original model openjev/openjev. Unlike the test model, it can be used directly for actual decision-making tasks and agent inference purposes.

Hardware Requirements

Estimated requirements (calculated by Local Model Watch) — 27.4B parameters (taken from the base model openjev/openjev)

Your VRAM Quantization File size Est. memory needed
24GB (RTX 4090 / 3090, etc.) Q4_K_M 17.7GB 21.2GB
32GB (RTX 5090, etc.) Q8_0 26.6GB 32.0GB
80GB class (A100 / H100) BF16 50.1GB 60.1GB

Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.

Can You Run It Locally?

Runs in Ollama, LM Studio and llama.cpp as-is.

It is distributed in GGUF, so no conversion is needed.

License — cc-by-nc-4.0 (Commercial use prohibited): Commercial use is prohibited (NC = NonCommercial). Internal business use can also count as commercial, so avoid this license for work use.

Compression: the Q4_K_M build measures 5.56 bits per weight — about 35% the size of the original 16-bit weights, calculated by this site from the actual file sizes.

Compiled by this site’s code from the published formats, converted builds we have found, and each engine’s own model registry. “Not found" means we have not seen such a build, not that none exists. License summaries are not legal advice — check the publisher’s original terms before relying on them.

Our Own Measurements

Measurements We Did Not Take

This model’s license (cc-by-nc-4.0) does not permit commercial use. Because this site carries advertising, we did not run the model (no CPU run, answers, quantization comparison, conversion or generation).

Recent Models in the Same Size Class

Models with 15–40B parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site’s estimates; licenses are as stated on the model cards.

→ Scroll horizontally to see all columns

Model Parameters Smallest VRAM tier License Our article
Qwen/Qwen3.8-27B 27.4B 80GB apache-2.0 clef Vision-Language Model: 80GB+ VRAM (2026-10-02)
Cloudflare/clef 27.4B 80GB apache-2.0 clef Structured Decision-Making Model: 80GB+ VRAM (2026-10-01)
Qwen/Qwen3.8-27B 27.8B 8GB apache-2.0 Qwen3.8-27B Vision-Language Model: Our Test Answers, 8GB+ VRAM (2026-09-26)
bartowski/vectionlabs_Salience-27B-R6-GGUF 27.8B 12GB apache-2.0 vectionlabs_Salience-27B-R6-GGUF: Our Test Answers, 12GB+ VRAM (2026-09-15)
agentionai/Signal-3.8-27B-GGUF 27.8B 16GB apache-2.0 Signal-3.8-27B-GGUF: Our Test Answers, 16GB+ VRAM (2026-09-12)

How to Get It

This model is distributed in GGUF format and can be run using inference engines such as llama.cpp.

When launching, you can use the llama serve command included in llama.cpp to load the model directly from the Hugging Face repository and start a server. Running the following command starts the API server in your local environment:

llama serve -hf ggml-org/OpenJev-GGUF

After startup, you can send decision-making requests via the dedicated API endpoint /v1/systemone.

Note that the weights of this model are released under the “CC BY-NC 4.0" license. While it is free for non-commercial research and use, acquiring a separate commercial license is required for commercial use.

Related Articles

What to Read Next

Sources