Local Model Watch

News, quick references and our own measurements for open-weight models you can run locally

  • All Articles
  • Our Measurements
  • New Models
  • Image, Video and Audio
  • Community
  • Engines and Tools
  • Weekly Roundup
  • About This Site
  • Editorial Policy
  • Guides & Data
    • Models by VRAM
    • Model Families
    • Engines & Tools
    • Benchmark Glossary
    • Quantization Glossary
  • 
  • 

    Menu

  • 

    Sidebar

  • 

    Prev

  • 

    Next

  • 

    Search

  • 日本語
  1. Home>
  2. New Models

ELYZA Releases ELYZA-Thinking-1.0 32B/33B Reasoning Models

October 2, 2026October 3, 2026

  • X
  • Bluesky
  • RSS
ELYZA Releases ELYZA-Thinking-1.0 32B/33B Reasoning Models

Contents
  • 1. At a Glance
  • 2. Overview
  • 3. Specifications
  • 4. Performance
  • 5. Strengths and Use Cases
  • 6. Hardware Requirements
  • 7. Can You Run It Locally?
  • 8. Recent Models in the Same Size Class
  • 9. How to Get It
  • 10. Other Models for the Same Task
  • 11. What to Read Next
  • 12. Sources

At a Glance

Item Value
Repository elyza/ELYZA-Thinking-1.0-llm-jp-4-32b-a3b
Publisher guide ELYZA: models and licenses
Published 2026-10-02
License apache-2.0
Formats safetensors
Source type Primary source (the publisher itself)

Values determined by this site’s code when the information was collected. Dates are JST.

Overview

ELYZA, Inc. has released ELYZA-Thinking-1.0, a reasoning model in 32B and 33B sizes based on the llm-jp-4 series of fully domestic Japanese foundational models. Among the released models, the 32B model (elyza/ELYZA-Thinking-1.0-llm-jp-4-32b-a3b) adopts an MoE architecture and has undergone mid-training and post-training to improve Japanese and English language understanding and generation capabilities. It specializes in Japanese-specific knowledge retrieval, instruction following with complex constraints, and reasoning capabilities in mathematics, coding, and STEM fields.

Specifications

  • Parameters: 32.1B (32B-A3B) / 33B (33B model)
  • Architecture: Qwen3MoeForCausalLM (32B) / LlamaForCausalLM (33B)
  • Context Length: 65,536
  • Active Parameters: 3,827,476,992 (32B)
  • Routed Experts: 128 (32B)
  • Active Experts: 8 (32B)
  • Hidden Size: 2,560 (32B) / 5,120 (33B)
  • Layers: 32 (32B) / 64 (33B)
  • Heads: 40 (32B) / 40 (33B)

Performance

The benchmark comparison table for the MoE model (ELYZA-thinking-1.0-llm-jp-4-32b-a3b) as listed in the model card is as follows. Scores for same-family models and other open-weight models are shown for comparison.

→ Scroll horizontally to see all columns

Benchmark ELYZA-thinking-1.0- llm-jp-4- 32b-a3b llm-jp-4- 32b-a3b- thinking llm-jp-4.1- 32b-a3b- thinking Qwen3- 30B-A3B gpt-oss-20b Nemotron 3.5 Lightning 30B-A3B Qwen3.5- 35B-A3B Gemma 4 26B-A4B
Knowledge & STEM
MMLU-Pro (en) 73.14 69.70 73.90 78.29 75.50 79.70 84.12 82.90
JMMLU (ja) 81.05 81.32 83.09 84.25 82.79 85.88 89.76 88.33
GPQA-Diamond (en) 63.86 51.93 60.80 63.26 68.59 75.98 84.38 78.60
GPQA (ja) 57.24 51.03 57.06 56.57 63.80 67.30 77.08 74.82
Math
MATH-500 (en) 95.20 84.20 93.80 97.20 97.40 98.00 99.00 98.80
JMATH-500 (ja) 88.20 83.20 87.40 91.60 92.20 94.00 93.20 92.00
AIME 2024+2025 (en) 62.29 34.48 57.71 76.25 88.23 89.90 92.92 90.62
PolyMath (ja, high+top) 34.60 12.80 27.40 43.45 55.10 58.20 56.40 65.10
Japanese QA
JamC-QA (ja) 59.51 53.75 53.79 45.30 41.10 55.17 60.29 66.18
JEMHopQA (ja) 71.59 62.71 65.51 55.20 54.16 55.17 58.88 61.93
Instruction Following
IFEval (en) 93.70 83.12 92.75 89.21 88.85 94.96 88.13 95.35
IFBench (en) 59.16 49.64 56.18 36.55 60.10 71.95 59.96 71.95
M-IFEval-ja (ja) 81.75 63.05 77.32 63.72 73.12 79.76 76.88 87.72
JFBench (ja) 38.15 27.14 30.87 24.23 18.38 34.39 28.83 30.30
Translation
WMT20 en→ja (ja) 23.98 23.53 23.76 22.17 23.20 20.29 24.89 27.90
WMT20 ja→en (en) 22.00 20.28 20.41 21.58 21.35 20.34 23.70 26.22
Coding
LiveCodeBench v6 (en) 53.66 26.86 32.40 56.40 59.94 75.94 74.91 77.09
JHumanEval (ja) 95.37 91.52 92.38 93.72 74.63 96.83 93.54 98.35
Function Calling
BFCL v4 (en) 36.31 13.29 27.96 42.65 48.09 62.37 67.03 66.88
Nejumi BFCL (ja) 41.89 11.58 39.19 49.03 47.10 52.12 60.62 57.34
Average
Overall 58.46 46.38 53.89 56.31 57.26 64.52 66.37 68.58
Japanese 59.61 49.16 56.65 56.73 55.04 62.02 64.24 66.68
English 55.94 41.16 49.72 56.84 61.45 68.98 69.98 71.55

According to evaluations by the publishers, ELYZA-Thinking-1.0-llm-jp-4-32b-a3b demonstrates high performance on Japanese-related tasks. In particular, it records excellent figures in IFEval and M-IFEval-ja, which measure instruction-following capability, as well as in Japanese coding via JHumanEval, showing overall score improvements compared to the base model group. On the other hand, in Function Calling and advanced English mathematics/competitive programming benchmarks (such as AIME and LiveCodeBench v6), it scores lower than some overseas large-scale models and other models, indicating room for improvement in specific English-language tasks and tool selection accuracy.

Strengths and Use Cases

The ELYZA-Thinking-1.0 series is built through a three-stage training process consisting of mid-training, SFT (supervised fine-tuning), and reinforcement learning (RLVR) to enhance Japanese and English language understanding and generation capabilities. It has been trained on localized Japanese reasoning data for mathematics, coding, and STEM fields, specializing in Japanese instruction following and knowledge retrieval. Additionally, it leverages synthetic knowledge data from Japanese Wikipedia and Wikidata to strengthen the understanding and retrieval of Japanese-specific knowledge. Furthermore, it supports tool calling and agent use cases, with agent data localized into Japanese while maintaining tool calls, function names, arguments, and JSON structures.

Hardware Requirements

Estimated requirements (calculated by Local Model Watch) — 32.1B parameters

Your VRAM Quantization File size Est. memory needed
80GB class (A100 / H100) BF16 59.9GB 71.8GB

Inference engine support (architecture name matched against each project’s own model registry in its source code, checked 2026-10-03): llama.cpp: registered, vLLM: registered, MLX (mlx-lm): registered.

Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.

Can You Run It Locally?

Not usable in Ollama, LM Studio and llama.cpp yet — we have found no GGUF build.

The publisher ships safetensors only. However, llama.cpp’s registry does list this architecture, so conversion to GGUF is possible and the model will run once someone publishes a converted build. Today it can be run with transformers or vLLM, using the memory figures in the table above.

License — apache-2.0 (Commercial use allowed): Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.

Compiled by this site’s code from the published formats, converted builds we have found, and each engine’s own model registry. “Not found" means we have not seen such a build, not that none exists. License summaries are not legal advice — check the publisher’s original terms before relying on them.

Recent Models in the Same Size Class

Models with 15–40B parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site’s estimates; licenses are as stated on the model cards.

→ Scroll horizontally to see all columns

Model Parameters Smallest VRAM tier License Our article
XingChen-AGI/Xing4.0-29B-A4B 31.2B 24GB apache-2.0 Xing4.0-29B-A4B Text Generation Model: Our Test Answers, 24GB+ VRAM (2026-09-28)
Altworld/Hemmingway-1 26.9B 12GB cc-by-nc-4.0 Hemmingway-1 Text Generation Model: 12GB+ VRAM, GGUF Builds (2026-09-28)
orcarouter/OrcaSAQ-2-27B 27.8B 16GB apache-2.0 OrcaSAQ-2-27B Text Generation Model: Our Test Answers, 16GB+ VRAM (2026-09-28)
prism-ml/Ternary-Bonsai-2-27B-gguf 27.8B 8GB apache-2.0 Ternary-Bonsai-2-27B-gguf Text Generation Model: 8GB+ VRAM (2026-09-18)
Edge0/Edge0-35B-A3B-preview 36.0B 24GB apache-2.0 Edge0-35B-A3B-preview 35B MoE Model for Phone-Class Memory: 24GB+ VRAM (2026-09-11)

How to Get It

The model is distributed in safetensors format and is available on Hugging Face under the Apache-2.0 license. There are no access restrictions or gated license agreements required on Hugging Face to use it.

It is compatible with engines such as vLLM and Transformers, and an example command for launching a server using vLLM is as follows:

vllm serve elyza/ELYZA-Thinking-1.0-llm-jp-4-32b-a3b \
  --max-model-len 65536 \
  --trust-remote-code \
  --moe-backend triton \
  --reasoning-parser-plugin vllm_plugins/multi_parser_loader.py \
  --reasoning-parser llmjp4 \
  --tool-parser-plugin vllm_plugins/llmjp_harmony_tool_parser.py \
  --tool-call-parser llmjp_harmony \
  --enable-auto-tool-choice

Other Models for the Same Task

Recent text generation models covered by Local Model Watch, newest first. Grouped by the task each publisher declares on Hugging Face (pipeline_tag); the smallest VRAM tier is this site’s estimate.

  • Xiaomi Releases MiMo-V2.6-Flash-MOPD and Pro-MOPD Models
  • MiMo-V2.6-Pro-RL Text Generation Model: ~417GB Memory, GGUF Builds (Over 80GB)
  • DeepSeek-V4-Pro-0813 Text Generation Model: ~998GB Memory, GGUF Builds (1650.5B, Over 80GB)
  • Tev1-0.8B-experimental Text Generation Model: 4GB+ VRAM (873M, 4GB)

See all text generation models →

What to Read Next

  • Find models by VRAM (This model needs at least 80GB) → VRAM quick reference
  • Engines that run this model → vLLM
  • Formats this model is available in → Safetensors format guide and models
  • Learn about the publisher → ELYZA: models, licenses and articles
  • How to read MMLU-Pro, GPQA Diamond, MATH → Benchmark glossary

Sources

  • https://huggingface.co/elyza/ELYZA-Thinking-1.0-llm-jp-4-32b-a3b
  • https://huggingface.co/elyza/ELYZA-Thinking-1.0-llm-jp-4-33b
  • https://huggingface.co/llm-jp/llm-jp-4-33b-base

New ModelsELYZA,ELYZA-Thinking-1.0,llm-jp-4,MoE,Verified,テキスト生成


  • X
  • Bluesky
  • RSS
Ai2 Releases AstaBrief 8B: Fast Open-Source Scientific Report…
Ai2 Releases AstaBrief 8B: Fast Open-Source Scientific Report…
Next
SGLang v0.5.21 Released with Dynamic PD Role Switching
SGLang v0.5.21 Released with Dynamic PD Role Switching
Prev

About Local Model Watch

Local Model Watch is a reference site, in English and Japanese, for open-weight generative models (text, image, video and audio) you can run on your own hardware: articles on new releases, hub pages, and measurements we take by running models on our own server. Memory requirements in every article are computed by this site from the actual distributed file sizes.

Start here

  • Our own measurements
  • Local models by VRAM
  • Model families
  • Publishers
  • Model formats
  • Models by task
  • Inference engines and tools
  • Trending models
  • Benchmark glossary
  • Quantization and model formats
  • All articles

About this site / Editorial policy

Latest Articles

October 3, 2026 : Engines and Tools

Magnitude 0.2.4 Released: Faster Inference and Lower Memory

Magnitude 0.2.4 is released with signifi ...

October 3, 2026 : New Models

Ai2 Releases AstaBrief 8B: Fast Open-Source Scientific Report…

Allen Institute for AI releases AstaBrie ...

October 2, 2026 : New Models

ELYZA Releases ELYZA-Thinking-1.0 32B/33B Reasoning Models

ELYZA has released ELYZA-Thinking-1.0, a ...

October 2, 2026 : Engines and Tools

SGLang v0.5.21 Released with Dynamic PD Role Switching

SGLang v0.5.21 is out, bringing dynamic ...

October 2, 2026 : New Models

clef Vision-Language Model: 80GB+ VRAM

Cloudflare has released Clef and Clef-fl ...

Categories

  • Community (9)
  • Engines and Tools (41)
  • Image, Video and Audio (25)
  • New Models (40)
  • Weekly Roundup (3)

Archives

  • October 2026
  • September 2026
  • Local Models by VRAM: Quick Reference
  • About This Site
  • Editorial Policy
  • Contact
  • Privacy Policy

Copyright © 2026 Local Model Watch All Rights Reserved.

WordPress Luxeritas Theme is provided by "Thought is free".

  • 
    Home
  • 
    Menu
  • 
    Page Top
 PAGE TOP