Testing OpenJev GGUF: A Tiny Model for Runtime Loaders

Testing OpenJev GGUF: A Tiny Model for Runtime Loaders

At a Glance

Item Value
Repository ggml-org/tinyopenjev-for-testing-gguf
Published 2026-10-02
License cc-by-nc-4.0
Formats GGUF
Source type Primary source (the publisher itself)

Values determined by this site’s code when the information was collected. Dates are JST.

Overview

ggml-org has released “tinyopenjev-for-testing-gguf", a tiny GGUF model created for the purpose of testing the open-weight decision-making model “OpenJev". This model is designed to verify whether inference engines and loaders function correctly, serving as a heavily trimmed-down subset of the original openjev/openjev model structure.

This model is strictly for testing purposes, and its outputs hold no meaning. It is designed for developers to verify loader and runtime behaviors, and is not suitable for practical inference tasks.

Claims and Evidence

The released tinyopenjev-for-testing-gguf is a GGUF format model (text model and mmproj) with Q8_0 quantization. It features a vastly reduced structure compared to the original model, with the publisher detailing the following specifics:

  • Text layers: Retains only the first 4 layers out of 64 total layers
  • Vision blocks: Retains only the first 2 blocks out of 27 total blocks
  • Hidden size: Reduced from 5120 to 64
  • Others: Reduction in attention head count and smaller FFN (Feed-Forward Network) dimensions

Due to this configuration, the parameter count is kept to approximately 34M, the majority of which is occupied by token embeddings and output matrices. The publisher explicitly states that because the outputs of this model are meaningless, it should not be used for any purpose other than testing loaders and runtimes.

Detailed data on the performance of the original model, openjev/openjev, has also been published. OpenJev is a model designed to quickly perform “typed decisions" from text, web pages, and screenshots. It evaluates up to 52 choices in a single forward pass and features a mechanism that outputs probability scores directly without waiting for free-form text generation.

Accuracy verification results conducted by the publisher based on 10,000 text questions (using 34 public sources) are as follows:

model accuracy
Jev (hosted API) 85.4% (8,540 of 10,000)
OpenJev 84.0% (8,403 of 10,000)
the same base model before tuning, same readout, its own calibration 80.4% (8,036 of 10,000)
Nimble 9B (open) 75.7% (7,574 of 10,000)

In these measurements, OpenJev is reported to have achieved an accuracy within 1.4 points of the hosted Jev API, outperforming the pre-tuned base model by 3.7 points and Nimble 9B by 8.3 points.

Decision-making accuracy by category is reported with the following figures:

→ Scroll horizontally to see all columns

kind of work questions OpenJev Jev (hosted API) before tuning
intent / routing / topic 1,218 92.8% 92.8% 93.2%
sentiment / stance 1,214 82.1% 81.5% 78.3%
spam / hate 695 80.4% 81.2% 80.1%
legal 1,305 78.9% 78.8% 77.8%
ethics / policy judgement 877 78.1% 75.8% 65.5%
commonsense reasoning 2,260 85.8% 88.3% 80.3%
science / facts / claims 1,558 82.5% 89.0% 81.0%
reading + language 873 89.2% 89.3% 83.3%

For agentic decision-making performance (such as determining the next action from a screenshot), the following improvements are shown:

test steps before tuning OpenJev
desktop: next action from a screenshot 2,000 76.5% 88.0%
web: next action on unseen websites 975 68.5% 87.4%
web: next action on unseen domains 1,000 65.7% 84.5%
answer flips when the options are shuffled 2,000 18.5% 2.3%

Notably, the stability of answers when option orders are shuffled (answer flips) showed significant improvement, dropping from 18.5% before tuning down to 2.3%.

Benchmark results for multilingual support and long-context reading comprehension are as follows:

test questions before tuning OpenJev
inference in German, French, Hindi, Chinese (XNLI) 240 72.5% 82.5%
intent in German, French, Hindi, Japanese (MASSIVE) 240 80.4% 85.8%
questions about 2,600 to 8,500-token articles (QuALITY) 120 91.7% 91.7%

It achieved an accuracy of 85.8% on the MASSIVE benchmark, which includes Japanese.

Regarding inference speed, median values measured using a single H100 GPU at FP8 precision are reported as follows:
– Short text decisions: approx. 80 ms
– Desktop operations (including screenshots): 176 ms
– Web operations (approx. 1,460 prompt tokens, 23 choices): approx. 210 ms
– Parallel requests for 1,200-token different pages (concurrency 8): 7 requests per second

These rapid decisions are achieved through a mechanism that reads answers directly from the scores of the first output position, eliminating free-form text generation and long generation wait times.

Prerequisites

The conditions for using and verifying the test model and original model presented in this material are as follows:

Target Models and Architecture

  • Test model: tinyopenjev-for-testing-gguf (a heavily sliced reduction of the original model structure)
  • Original model: openjev/openjev (27.4B parameters)
  • Architecture: Qwen3_5ForConditionalGeneration (built based on Qwen/Qwen3.8-27B)
  • Quantization format: Q8_0 (GGUF version), FP8 (during original model speed measurements)

Software Environment

  • Inference engine: vllm==0.29.0
  • Client library: openai==3.16.2, httpx==0.28.1
  • Auxiliary tool: openjev/helper/shim.py (shim for decision-making API)

Configurations and Limitations

  • Context length: Up to 16,384 tokens
  • Image input: Up to 1 per request
  • Number of choices: Up to 52 in a single pass (multiple passes used for more choices)
  • License: CC BY-NC 4.0 (restricted to non-commercial use; commercial use requires a separate license)

What Can Be Replicated Locally

Readers can use the public repositories to perform the following verifications in their own environments:

Loader Verification Using the Test Model

By using ggml-org/tinyopenjev-for-testing-gguf, you can verify with minimal resources whether GGUF format loaders and runtimes function correctly. Since this model includes mmproj alongside the text model, multimodal input pipeline testing is also possible. However, because the model structure (layer count and hidden size) has been drastically reduced, the validity of inference results cannot be verified.

Serving and API Usage of the Original Model

You can build an environment to execute actual decision-making tasks using the original model openjev/openjev. The replication steps are as follows:

  1. Prepare the Model:
    Obtain the model files locally using the command hf download openjev/openjev --local-dir openjev.

  2. Start the Inference Server:
    Launch the server using vllm with the following command: bash vllm serve./openjev --host 127.0.0.1 --served-model-name qwen --port 8000 \ --enable-prefix-caching --max-model-len 16384 --gpu-memory-utilization 0.90 \ --limit-mm-per-prompt '{"image":1}' --trust-remote-code --max-num-seqs 256 \ --max-logprobs 64 --gdn-prefill-backend triton --quantization fp8

  3. Build the Decision-Making API:
    Build an API endpoint specialized for decision-making using the included shim.py: bash VLLM=http://localhost:8000/v1 TOKENIZER=./openjev \ READOUT_T=0.85 READOUT_NOUL_T=1.829074 READOUT_NOUL_BIAS=0 \ READOUT_TARGETED=1 READOUT_INSTR_STYLE=pyrepr SHIM_STAGGER=1 \ python openjev/helper/shim.py --host 127.0.0.1 --port 3000

  4. Execute Requests:
    Send decision-making requests for specific texts or images using the curl command or similar tools. Requests can include questions, choices, and scoring criteria in JSON format.

What the Material Does Not Cover

This material does not include the following information:

  • Specific text examples of the “meaningless content" output by the test model tinyopenjev-for-testing-gguf, nor detailed explanations of its exact behaviors.
  • Benchmark results for the GGUF-quantized model using the full 10,000-question set are not shown, and measurements using a 1,789-question subset are treated merely as “current data".
  • Specific inference latency or throughput measurement data on consumer-grade GPUs other than the H100 (e.g., NVIDIA GeForce RTX series) are not provided.
  • Specific costs or support structure details for acquiring a commercial license are not mentioned.
  • The exact contents of the training datasets or training times used between the pre-tuned base model and OpenJev are not specified.

Hardware Requirements

Estimated requirements (calculated by Local Model Watch) — 27.4B parameters (taken from the base model openjev/openjev)

Your VRAM Quantization File size Est. memory needed
4GB (laptop iGPU / phone class) Q8_0 0.0GB 0.1GB

Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.

Related Articles

What to Read Next

Sources