{"id":8896,"date":"2026-10-02T07:11:57","date_gmt":"2026-10-01T22:11:57","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/02\/ggml-org-openjev-gguf-released-2\/"},"modified":"2026-10-02T09:14:11","modified_gmt":"2026-10-02T00:14:11","slug":"ggml-org-openjev-gguf-released-2","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/ggml-org-openjev-gguf-released-2\/","title":{"rendered":"OpenJev-GGUF 27.4B Decision-Making Specialized Model: 24GB+ VRAM"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/ggml-org\/OpenJev-GGUF\">ggml-org\/OpenJev-GGUF<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-10-02<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>cc-by-nc-4.0<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-gguf-en\/\">GGUF<\/a><\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code when the information was collected. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>ggml-org has released &#8220;ggml-org\/OpenJev-GGUF&#8221;, a GGUF quantized version of &#8220;OpenJev&#8221;, an open-weight model specialized in decision-making. This model is designed to accept text, web pages, and screenshots as inputs to perform choice-based decision-making (decision-model). Instead of generating free text and parsing it for user-requested questions in JSON format, it features a mechanism that directly outputs probabilities or scores for each choice in a single forward pass.<\/p>\n<h2>Specifications<\/h2>\n<p>The specifications of the original model <code>openjev\/openjev<\/code> are as follows:<br \/>\n&#8211; Parameters: 27.4B<br \/>\n&#8211; Architecture: Qwen3_5ForConditionalGeneration (qwen3_5)<br \/>\n&#8211; Context length: Up to 16,384 tokens<\/p>\n<h2>Performance<\/h2>\n<p>Here are the benchmark results for the original model <code>openjev\/openjev<\/code> published by its creators. Since this is a GGUF quantized version, the following figures are measured values on the unquantized original model.<\/p>\n<p>First, here is a comparison with other models across 10,000 text-based questions (using 34 public sources):<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>model<\/th>\n<th>accuracy<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Jev (hosted API)<\/td>\n<td>85.4% (8,540 of 10,000)<\/td>\n<\/tr>\n<tr>\n<td><strong>OpenJev<\/strong><\/td>\n<td><strong>84.0%<\/strong> (8,403 of 10,000)<\/td>\n<\/tr>\n<tr>\n<td>the same base model before tuning, same readout, its own calibration<\/td>\n<td>80.4% (8,036 of 10,000)<\/td>\n<\/tr>\n<tr>\n<td>Nimble 9B (open)<\/td>\n<td>75.7% (7,574 of 10,000)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>From this table, we can see that OpenJev achieves an accuracy of 84.0%, significantly outperforming the pre-tuning base model (80.4%) and the open model Nimble 9B (75.7%). Meanwhile, while it fell slightly short of the hosted API Jev (hosted API) at 85.4%, it comes very close with a small difference of 1.4 points.<\/p>\n<p>Next is the breakdown of accuracy by category:<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>kind of work<\/th>\n<th>questions<\/th>\n<th>OpenJev<\/th>\n<th>Jev (hosted API)<\/th>\n<th>before tuning<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>intent \/ routing \/ topic<\/td>\n<td>1,218<\/td>\n<td>92.8%<\/td>\n<td>92.8%<\/td>\n<td>93.2%<\/td>\n<\/tr>\n<tr>\n<td>sentiment \/ stance<\/td>\n<td>1,214<\/td>\n<td>82.1%<\/td>\n<td>81.5%<\/td>\n<td>78.3%<\/td>\n<\/tr>\n<tr>\n<td>spam \/ hate<\/td>\n<td>695<\/td>\n<td>80.4%<\/td>\n<td>81.2%<\/td>\n<td>80.1%<\/td>\n<\/tr>\n<tr>\n<td>legal<\/td>\n<td>1,305<\/td>\n<td>78.9%<\/td>\n<td>78.8%<\/td>\n<td>77.8%<\/td>\n<\/tr>\n<tr>\n<td>ethics \/ policy judgement<\/td>\n<td>877<\/td>\n<td>78.1%<\/td>\n<td>75.8%<\/td>\n<td>65.5%<\/td>\n<\/tr>\n<tr>\n<td>commonsense reasoning<\/td>\n<td>2,260<\/td>\n<td>85.8%<\/td>\n<td>88.3%<\/td>\n<td>80.3%<\/td>\n<\/tr>\n<tr>\n<td>science \/ facts \/ claims<\/td>\n<td>1,558<\/td>\n<td>82.5%<\/td>\n<td>89.0%<\/td>\n<td>81.0%<\/td>\n<\/tr>\n<tr>\n<td>reading + language<\/td>\n<td>873<\/td>\n<td>89.2%<\/td>\n<td>89.3%<\/td>\n<td>83.3%<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>Looking at these category results, OpenJev shows high accuracy in tasks such as &#8220;intent \/ routing \/ topic&#8221; (92.8%) and &#8220;reading + language&#8221; (89.2%). It also scores higher than the hosted API Jev in &#8220;ethics \/ policy judgement&#8221; (78.1%) and &#8220;sentiment \/ stance&#8221; (82.1%). However, in fields like &#8220;science \/ facts \/ claims&#8221; (82.5%) and &#8220;commonsense reasoning&#8221; (85.8%), it falls slightly behind the hosted API Jev (89.0% and 88.3% respectively).<\/p>\n<p>Additionally, the comparison results before and after tuning on agent-based decision-making tasks are as follows:<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>test<\/th>\n<th>steps<\/th>\n<th>before tuning<\/th>\n<th>OpenJev<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>desktop: next action from a screenshot<\/td>\n<td>2,000<\/td>\n<td>76.5%<\/td>\n<td><strong>88.0%<\/strong><\/td>\n<\/tr>\n<tr>\n<td>web: next action on unseen websites<\/td>\n<td>975<\/td>\n<td>68.5%<\/td>\n<td><strong>87.4%<\/strong><\/td>\n<\/tr>\n<tr>\n<td>web: next action on unseen domains<\/td>\n<td>1,000<\/td>\n<td>65.7%<\/td>\n<td><strong>84.5%<\/strong><\/td>\n<\/tr>\n<tr>\n<td>answer flips when the options are shuffled<\/td>\n<td>2,000<\/td>\n<td>18.5%<\/td>\n<td><strong>2.3%<\/strong><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>From these results, OpenJev achieves a significant accuracy improvement of over 10 points compared to pre-tuning when determining next actions for desktop operations from screenshots or on unseen websites and domains. Furthermore, answer flips when options are shuffled decreased drastically from 18.5% to 2.3%, demonstrating extremely high stability as a decision-making model.<\/p>\n<p>Comparisons regarding multilingual support and long-context reading comprehension are as follows:<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>test<\/th>\n<th>questions<\/th>\n<th>before tuning<\/th>\n<th>OpenJev<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>inference in German, French, Hindi, Chinese (XNLI)<\/td>\n<td>240<\/td>\n<td>72.5%<\/td>\n<td><strong>82.5%<\/strong><\/td>\n<\/tr>\n<tr>\n<td>intent in German, French, Hindi, Japanese (MASSIVE)<\/td>\n<td>240<\/td>\n<td>80.4%<\/td>\n<td><strong>85.8%<\/strong><\/td>\n<\/tr>\n<tr>\n<td>questions about 2,600 to 8,500-token articles (QuALITY)<\/td>\n<td>120<\/td>\n<td>91.7%<\/td>\n<td>91.7%<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>This table indicates that accuracy has improved over pre-tuning in multilingual tasks involving German, French, Hindi, Chinese, and Japanese. On the other hand, in question tasks regarding long articles (QuALITY), the score remains unchanged at 91.7% before and after tuning, maintaining equivalent performance.<\/p>\n<p>Finally, here is the number of completed tasks when executing end-to-end browser tasks (100 types of MiniWoB tasks, 1 trial each):<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>model<\/th>\n<th>tasks completed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>OpenJev<\/strong><\/td>\n<td><strong>39<\/strong><\/td>\n<\/tr>\n<tr>\n<td>Jev (hosted API)<\/td>\n<td>39<\/td>\n<\/tr>\n<tr>\n<td>the same base model before tuning<\/td>\n<td>38<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>These results confirm that in end-to-end browser tasks, OpenJev completes 39 tasks, matching the hosted Jev API and showing a slight improvement from the 38 tasks completed before tuning.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>This model is a &#8220;decision model&#8221; that accepts text or images (screenshots) as inputs and makes optimal decisions from pre-defined choices. Unlike general chat bots that generate free text before parsing, it calculates probabilities directly from token scores at the initial output position, enabling fast and reliable processing.<\/p>\n<p>Based on the performance and features of the original model <code>openjev\/openjev<\/code>, it is highly suited for the following applications and tasks:<\/p>\n<ul>\n<li>\n<p><strong>Routing and Triage<\/strong>:<br \/>\n  Determining the intent, topic, responsible team, priority, or language of user inquiry messages to automatically route them to the appropriate destination.<\/p>\n<\/li>\n<li>\n<p><strong>Moderation and Safety Judgment<\/strong>:<br \/>\n  Evaluating whether input text falls under harmful content, spam, hate speech, policy violations, or legal threats.<\/p>\n<\/li>\n<li>\n<p><strong>LLM as a Judge<\/strong>:<br \/>\n  Objectively determining whether a generative AI model&#8217;s output is based on reliable evidence (presence of hallucinations), follows provided evaluation criteria (rubrics), or which of multiple answers is superior.<\/p>\n<\/li>\n<li>\n<p><strong>Business Processes and Document Management<\/strong>:<br \/>\n  Automating routine judgment tasks such as approving, withholding, or rejecting invoices, determining the severity of system alerts, or classifying specific clause types in contracts.<\/p>\n<\/li>\n<li>\n<p><strong>Browser and Desktop Agents<\/strong>:<br \/>\n  Reading screen screenshots, HTML (DOM structure), or JSON data to determine the next element to interact with or action to execute (such as clicks or input). It can also judge whether a task is complete or blocked.<\/p>\n<\/li>\n<li>\n<p><strong>Scoring<\/strong>:<br \/>\n  Calculating ordered step-by-step evaluations accompanied by expected values or confidence levels.<\/p>\n<\/li>\n<\/ul>\n<p>Furthermore, because this model allows dynamic specification of choices (labels) per request, task-specific pre-training or fixed labeling is unnecessary. It can process up to 52 choices in a single request during a single forward pass, and can simultaneously handle multiple questions (e.g., routing, sentiment analysis, urgency) in one request. A major strength is its support for multiple languages including Japanese (English, German, French, Hindi, Chinese, and Japanese).<\/p>\n<h2>How It Differs from Similar Models<\/h2>\n<p>In a previous article introduced on our site, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/02\/openjev-testing-gguf-released\/\">Decision-Making Specialized Model &#8220;OpenJev&#8221; Testing GGUF Model Released<\/a>, we covered &#8220;tinyopenjev-for-testing-gguf&#8221;, also released by ggml-org. However, that was an ultra-small test-only model (approx. 34M parameters) created solely to verify whether inference engines and loaders function properly, and its output content carried no meaning.<\/p>\n<p>In contrast, the newly released &#8220;OpenJev-GGUF&#8221; is a practical quantized model that retains the entire structure of the original model <code>openjev\/openjev<\/code>. Unlike the test model, it can be used directly for actual decision-making tasks and agent inference purposes.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 27.4B parameters (taken from the base model openjev\/openjev)<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>24GB (RTX 4090 \/ 3090, etc.)<\/td>\n<td>Q4_K_M<\/td>\n<td>17.7GB<\/td>\n<td>21.2GB<\/td>\n<\/tr>\n<tr>\n<td>32GB (RTX 5090, etc.)<\/td>\n<td>Q8_0<\/td>\n<td>26.6GB<\/td>\n<td>32.0GB<\/td>\n<\/tr>\n<tr>\n<td>80GB class (A100 \/ H100)<\/td>\n<td>BF16<\/td>\n<td>50.1GB<\/td>\n<td>60.1GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<p><!-- lmw:runnability --><\/p>\n<h2>Can You Run It Locally?<\/h2>\n<p><strong>Runs in Ollama, LM Studio and llama.cpp as-is.<\/strong><\/p>\n<p>It is distributed in GGUF, so no conversion is needed.<\/p>\n<p><strong>License \u2014 <code>cc-by-nc-4.0<\/code> (Commercial use prohibited):<\/strong> <strong>Commercial use is prohibited<\/strong> (NC = NonCommercial). Internal business use can also count as commercial, so avoid this license for work use.<\/p>\n<p><strong>Compression:<\/strong> the Q4_K_M build measures 5.56 bits per weight \u2014 about 35% the size of the original 16-bit weights, calculated by this site from the actual file sizes.<\/p>\n<p><em>Compiled by this site&#8217;s code from the published formats, converted builds we have found, and each engine&#8217;s own model registry. &#8220;Not found&#8221; means we have not seen such a build, not that none exists. License summaries are not legal advice \u2014 check the publisher&#8217;s original terms before relying on them.<\/em><\/p>\n<p><!-- \/lmw:runnability --><\/p>\n<p><!-- lmw:lab --><\/p>\n<h2>Our Own Measurements<\/h2>\n<h3>Measurements We Did Not Take<\/h3>\n<p>This model&#8217;s license (<code>cc-by-nc-4.0<\/code>) does not permit commercial use. Because this site carries advertising, we did not run the model (no CPU run, answers, quantization comparison, conversion or generation).<\/p>\n<p><!-- \/lmw:lab --><\/p>\n<p><!-- lmw:peers --><\/p>\n<h2>Recent Models in the Same Size Class<\/h2>\n<p><em>Models with <\/em><em>15\u201340B<\/em><em> parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site&#8217;s estimates; licenses are as stated on the model cards.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Parameters<\/th>\n<th>Smallest VRAM tier<\/th>\n<th>License<\/th>\n<th>Our article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Qwen\/Qwen3.8-27B<\/td>\n<td>27.4B<\/td>\n<td>80GB<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/cloudflare-releases-clef-and-clef-flash\/\">clef Vision-Language Model: 80GB+ VRAM<\/a> (2026-10-02)<\/td>\n<\/tr>\n<tr>\n<td>Cloudflare\/clef<\/td>\n<td>27.4B<\/td>\n<td>80GB<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/cloudflare-clef-27b-multimodal-decision-model\/\">clef Structured Decision-Making Model: 80GB+ VRAM<\/a> (2026-10-01)<\/td>\n<\/tr>\n<tr>\n<td>Qwen\/Qwen3.8-27B<\/td>\n<td>27.8B<\/td>\n<td>8GB<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/qwen3-8-27b-overview-specs-performance\/\">Qwen3.8-27B Vision-Language Model: Our Test Answers, 8GB+ VRAM<\/a> (2026-09-26)<\/td>\n<\/tr>\n<tr>\n<td>bartowski\/vectionlabs_Salience-27B-R6-GGUF<\/td>\n<td>27.8B<\/td>\n<td>12GB<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/salience-27b-r6-gguf-2\/\">vectionlabs_Salience-27B-R6-GGUF: Our Test Answers, 12GB+ VRAM<\/a> (2026-09-15)<\/td>\n<\/tr>\n<tr>\n<td>agentionai\/Signal-3.8-27B-GGUF<\/td>\n<td>27.8B<\/td>\n<td>16GB<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/12\/signal-3-8-27b-gguf-overview\/\">Signal-3.8-27B-GGUF: Our Test Answers, 16GB+ VRAM<\/a> (2026-09-12)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:peers --><\/p>\n<h2>How to Get It<\/h2>\n<p>This model is distributed in GGUF format and can be run using inference engines such as <code>llama.cpp<\/code>.<\/p>\n<p>When launching, you can use the <code>llama serve<\/code> command included in <code>llama.cpp<\/code> to load the model directly from the Hugging Face repository and start a server. Running the following command starts the API server in your local environment:<\/p>\n<pre><code class=\"language-bash\">llama serve -hf ggml-org\/OpenJev-GGUF\n<\/code><\/pre>\n<p>After startup, you can send decision-making requests via the dedicated API endpoint <code>\/v1\/systemone<\/code>.<\/p>\n<p>Note that the weights of this model are released under the &#8220;CC BY-NC 4.0&#8221; license. While it is free for non-commercial research and use, acquiring a separate commercial license is required for commercial use.<\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/testing-openjev-gguf-tiny-model-for-runtime-loaders\/\">Testing OpenJev GGUF: A Tiny Model for Runtime Loaders<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/mimo-v2-6-flash-rl-gguf-released\/\">MiMo-V2.6-Flash-RL-GGUF Multimodal MoE Model: ~141GB Memory<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/mimo-v2-6-distill-qwen-9b-gguf\/\">MiMo-V2.6-Distill-Qwen-9B-GGUF: Our Test Answers, 12GB+ VRAM<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/14\/ggml-v0-24-0-released-2\/\">ggml v0.24.0 Released with Backend Improvements and API Updates<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Find models by VRAM<\/strong> (This model runs from the 24GB tier) \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-24gb-en\/\">Other models that run on a 24GB GPU<\/a><\/li>\n<li><strong>Engines that run this model<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-llama-cpp-en\/\">llama.cpp<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-ollama-en\/\">Ollama<\/a><\/li>\n<li><strong>What Q4_K_M, Q8_0 mean and where to get this model<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-gguf-en\/\">GGUF format guide and models<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">Quantization and model-format glossary<\/a><\/li>\n<li><strong>Other models for the same task<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/models-by-task-en\/#task-vision\">Other vision-language models<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/ggml-org\/OpenJev-GGUF\">https:\/\/huggingface.co\/ggml-org\/OpenJev-GGUF<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/openjev\/openjev\">https:\/\/huggingface.co\/openjev\/openjev<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Explore OpenJev-GGUF, a GGUF quantized open-weight decision-making model by ggml-org designed for structured choices and agent tasks.<\/p>\n","protected":false},"author":1,"featured_media":8895,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[310],"tags":[2856,161,1970,163,1073,2818,1547,1952],"class_list":["post-8896","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-new-models","tag-decision-model-en","tag-ggml-en","tag-ggml-org-en","tag-gguf-en","tag-llama-cpp-en","tag-openjev-en","tag-verified","tag-vlm-en"],"lang":"en","translations":{"en":8896,"ja":8894},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/8896","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=8896"}],"version-history":[{"count":1,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/8896\/revisions"}],"predecessor-version":[{"id":9022,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/8896\/revisions\/9022"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/8895"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=8896"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=8896"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=8896"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}