Testing OpenJev GGUF: A Tiny Model for Runtime Loaders

At a Glance
| Item | Value |
|---|---|
| Repository | ggml-org/tinyopenjev-for-testing-gguf |
| Published | 2026-10-02 |
| License | cc-by-nc-4.0 |
| Formats | GGUF |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code when the information was collected. Dates are JST.
Overview
ggml-org has released “tinyopenjev-for-testing-gguf", a tiny GGUF model created for the purpose of testing the open-weight decision-making model “OpenJev". This model is designed to verify whether inference engines and loaders function correctly, serving as a heavily trimmed-down subset of the original openjev/openjev model structure.
This model is strictly for testing purposes, and its outputs hold no meaning. It is designed for developers to verify loader and runtime behaviors, and is not suitable for practical inference tasks.
Claims and Evidence
The released tinyopenjev-for-testing-gguf is a GGUF format model (text model and mmproj) with Q8_0 quantization. It features a vastly reduced structure compared to the original model, with the publisher detailing the following specifics:
- Text layers: Retains only the first 4 layers out of 64 total layers
- Vision blocks: Retains only the first 2 blocks out of 27 total blocks
- Hidden size: Reduced from 5120 to 64
- Others: Reduction in attention head count and smaller FFN (Feed-Forward Network) dimensions
Due to this configuration, the parameter count is kept to approximately 34M, the majority of which is occupied by token embeddings and output matrices. The publisher explicitly states that because the outputs of this model are meaningless, it should not be used for any purpose other than testing loaders and runtimes.
Detailed data on the performance of the original model, openjev/openjev, has also been published. OpenJev is a model designed to quickly perform “typed decisions" from text, web pages, and screenshots. It evaluates up to 52 choices in a single forward pass and features a mechanism that outputs probability scores directly without waiting for free-form text generation.
Accuracy verification results conducted by the publisher based on 10,000 text questions (using 34 public sources) are as follows:
| model | accuracy |
|---|---|
| Jev (hosted API) | 85.4% (8,540 of 10,000) |
| OpenJev | 84.0% (8,403 of 10,000) |
| the same base model before tuning, same readout, its own calibration | 80.4% (8,036 of 10,000) |
| Nimble 9B (open) | 75.7% (7,574 of 10,000) |
In these measurements, OpenJev is reported to have achieved an accuracy within 1.4 points of the hosted Jev API, outperforming the pre-tuned base model by 3.7 points and Nimble 9B by 8.3 points.
Decision-making accuracy by category is reported with the following figures:
→ Scroll horizontally to see all columns
| kind of work | questions | OpenJev | Jev (hosted API) | before tuning |
|---|---|---|---|---|
| intent / routing / topic | 1,218 | 92.8% | 92.8% | 93.2% |
| sentiment / stance | 1,214 | 82.1% | 81.5% | 78.3% |
| spam / hate | 695 | 80.4% | 81.2% | 80.1% |
| legal | 1,305 | 78.9% | 78.8% | 77.8% |
| ethics / policy judgement | 877 | 78.1% | 75.8% | 65.5% |
| commonsense reasoning | 2,260 | 85.8% | 88.3% | 80.3% |
| science / facts / claims | 1,558 | 82.5% | 89.0% | 81.0% |
| reading + language | 873 | 89.2% | 89.3% | 83.3% |
For agentic decision-making performance (such as determining the next action from a screenshot), the following improvements are shown:
| test | steps | before tuning | OpenJev |
|---|---|---|---|
| desktop: next action from a screenshot | 2,000 | 76.5% | 88.0% |
| web: next action on unseen websites | 975 | 68.5% | 87.4% |
| web: next action on unseen domains | 1,000 | 65.7% | 84.5% |
| answer flips when the options are shuffled | 2,000 | 18.5% | 2.3% |
Notably, the stability of answers when option orders are shuffled (answer flips) showed significant improvement, dropping from 18.5% before tuning down to 2.3%.
Benchmark results for multilingual support and long-context reading comprehension are as follows:
| test | questions | before tuning | OpenJev |
|---|---|---|---|
| inference in German, French, Hindi, Chinese (XNLI) | 240 | 72.5% | 82.5% |
| intent in German, French, Hindi, Japanese (MASSIVE) | 240 | 80.4% | 85.8% |
| questions about 2,600 to 8,500-token articles (QuALITY) | 120 | 91.7% | 91.7% |
It achieved an accuracy of 85.8% on the MASSIVE benchmark, which includes Japanese.
Regarding inference speed, median values measured using a single H100 GPU at FP8 precision are reported as follows:
– Short text decisions: approx. 80 ms
– Desktop operations (including screenshots): 176 ms
– Web operations (approx. 1,460 prompt tokens, 23 choices): approx. 210 ms
– Parallel requests for 1,200-token different pages (concurrency 8): 7 requests per second
These rapid decisions are achieved through a mechanism that reads answers directly from the scores of the first output position, eliminating free-form text generation and long generation wait times.
Prerequisites
The conditions for using and verifying the test model and original model presented in this material are as follows:
Target Models and Architecture
- Test model:
tinyopenjev-for-testing-gguf(a heavily sliced reduction of the original model structure) - Original model:
openjev/openjev(27.4B parameters) - Architecture:
Qwen3_5ForConditionalGeneration(built based onQwen/Qwen3.8-27B) - Quantization format: Q8_0 (GGUF version), FP8 (during original model speed measurements)
Software Environment
- Inference engine:
vllm==0.29.0 - Client library:
openai==3.16.2,httpx==0.28.1 - Auxiliary tool:
openjev/helper/shim.py(shim for decision-making API)
Configurations and Limitations
- Context length: Up to 16,384 tokens
- Image input: Up to 1 per request
- Number of choices: Up to 52 in a single pass (multiple passes used for more choices)
- License: CC BY-NC 4.0 (restricted to non-commercial use; commercial use requires a separate license)
What Can Be Replicated Locally
Readers can use the public repositories to perform the following verifications in their own environments:
Loader Verification Using the Test Model
By using ggml-org/tinyopenjev-for-testing-gguf, you can verify with minimal resources whether GGUF format loaders and runtimes function correctly. Since this model includes mmproj alongside the text model, multimodal input pipeline testing is also possible. However, because the model structure (layer count and hidden size) has been drastically reduced, the validity of inference results cannot be verified.
Serving and API Usage of the Original Model
You can build an environment to execute actual decision-making tasks using the original model openjev/openjev. The replication steps are as follows:
-
Prepare the Model:
Obtain the model files locally using the commandhf download openjev/openjev --local-dir openjev. -
Start the Inference Server:
Launch the server usingvllmwith the following command:bash vllm serve./openjev --host 127.0.0.1 --served-model-name qwen --port 8000 \ --enable-prefix-caching --max-model-len 16384 --gpu-memory-utilization 0.90 \ --limit-mm-per-prompt '{"image":1}' --trust-remote-code --max-num-seqs 256 \ --max-logprobs 64 --gdn-prefill-backend triton --quantization fp8 -
Build the Decision-Making API:
Build an API endpoint specialized for decision-making using the includedshim.py:bash VLLM=http://localhost:8000/v1 TOKENIZER=./openjev \ READOUT_T=0.85 READOUT_NOUL_T=1.829074 READOUT_NOUL_BIAS=0 \ READOUT_TARGETED=1 READOUT_INSTR_STYLE=pyrepr SHIM_STAGGER=1 \ python openjev/helper/shim.py --host 127.0.0.1 --port 3000 -
Execute Requests:
Send decision-making requests for specific texts or images using thecurlcommand or similar tools. Requests can include questions, choices, and scoring criteria in JSON format.
What the Material Does Not Cover
This material does not include the following information:
- Specific text examples of the “meaningless content" output by the test model
tinyopenjev-for-testing-gguf, nor detailed explanations of its exact behaviors. - Benchmark results for the GGUF-quantized model using the full 10,000-question set are not shown, and measurements using a 1,789-question subset are treated merely as “current data".
- Specific inference latency or throughput measurement data on consumer-grade GPUs other than the H100 (e.g., NVIDIA GeForce RTX series) are not provided.
- Specific costs or support structure details for acquiring a commercial license are not mentioned.
- The exact contents of the training datasets or training times used between the pre-tuned base model and OpenJev are not specified.
Hardware Requirements
Estimated requirements (calculated by Local Model Watch) — 27.4B parameters (taken from the base model openjev/openjev)
| Your VRAM | Quantization | File size | Est. memory needed |
|---|---|---|---|
| 4GB (laptop iGPU / phone class) | Q8_0 | 0.0GB | 0.1GB |
Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.
Related Articles
- OpenJev-GGUF 27.4B Decision-Making Specialized Model: 24GB+ VRAM
- MiMo-V2.6-Flash-RL-GGUF Multimodal MoE Model: ~141GB Memory
- MiMo-V2.6-Distill-Qwen-9B-GGUF: Our Test Answers, 12GB+ VRAM
- ggml v0.24.0 Released with Backend Improvements and API Updates
What to Read Next
- Find models by VRAM (This model runs from the 4GB tier) → Other models that run on a 8GB GPU
- Engines mentioned in this article → ggml / vLLM
- What Q8_0 mean → Quantization and model-format glossary

