ELYZA Releases ELYZA-Thinking-1.0 32B/33B Reasoning Models

At a Glance
| Item | Value |
|---|---|
| Repository | elyza/ELYZA-Thinking-1.0-llm-jp-4-32b-a3b |
| Publisher guide | ELYZA: models and licenses |
| Published | 2026-10-02 |
| License | apache-2.0 |
| Formats | safetensors |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code when the information was collected. Dates are JST.
Overview
ELYZA, Inc. has released ELYZA-Thinking-1.0, a reasoning model in 32B and 33B sizes based on the llm-jp-4 series of fully domestic Japanese foundational models. Among the released models, the 32B model (elyza/ELYZA-Thinking-1.0-llm-jp-4-32b-a3b) adopts an MoE architecture and has undergone mid-training and post-training to improve Japanese and English language understanding and generation capabilities. It specializes in Japanese-specific knowledge retrieval, instruction following with complex constraints, and reasoning capabilities in mathematics, coding, and STEM fields.
Specifications
- Parameters: 32.1B (32B-A3B) / 33B (33B model)
- Architecture: Qwen3MoeForCausalLM (32B) / LlamaForCausalLM (33B)
- Context Length: 65,536
- Active Parameters: 3,827,476,992 (32B)
- Routed Experts: 128 (32B)
- Active Experts: 8 (32B)
- Hidden Size: 2,560 (32B) / 5,120 (33B)
- Layers: 32 (32B) / 64 (33B)
- Heads: 40 (32B) / 40 (33B)
Performance
The benchmark comparison table for the MoE model (ELYZA-thinking-1.0-llm-jp-4-32b-a3b) as listed in the model card is as follows. Scores for same-family models and other open-weight models are shown for comparison.
→ Scroll horizontally to see all columns
| Benchmark | ELYZA-thinking-1.0- llm-jp-4- 32b-a3b | llm-jp-4- 32b-a3b- thinking | llm-jp-4.1- 32b-a3b- thinking | Qwen3- 30B-A3B | gpt-oss-20b | Nemotron 3.5 Lightning 30B-A3B | Qwen3.5- 35B-A3B | Gemma 4 26B-A4B |
|---|---|---|---|---|---|---|---|---|
| Knowledge & STEM | ||||||||
| MMLU-Pro (en) | 73.14 | 69.70 | 73.90 | 78.29 | 75.50 | 79.70 | 84.12 | 82.90 |
| JMMLU (ja) | 81.05 | 81.32 | 83.09 | 84.25 | 82.79 | 85.88 | 89.76 | 88.33 |
| GPQA-Diamond (en) | 63.86 | 51.93 | 60.80 | 63.26 | 68.59 | 75.98 | 84.38 | 78.60 |
| GPQA (ja) | 57.24 | 51.03 | 57.06 | 56.57 | 63.80 | 67.30 | 77.08 | 74.82 |
| Math | ||||||||
| MATH-500 (en) | 95.20 | 84.20 | 93.80 | 97.20 | 97.40 | 98.00 | 99.00 | 98.80 |
| JMATH-500 (ja) | 88.20 | 83.20 | 87.40 | 91.60 | 92.20 | 94.00 | 93.20 | 92.00 |
| AIME 2024+2025 (en) | 62.29 | 34.48 | 57.71 | 76.25 | 88.23 | 89.90 | 92.92 | 90.62 |
| PolyMath (ja, high+top) | 34.60 | 12.80 | 27.40 | 43.45 | 55.10 | 58.20 | 56.40 | 65.10 |
| Japanese QA | ||||||||
| JamC-QA (ja) | 59.51 | 53.75 | 53.79 | 45.30 | 41.10 | 55.17 | 60.29 | 66.18 |
| JEMHopQA (ja) | 71.59 | 62.71 | 65.51 | 55.20 | 54.16 | 55.17 | 58.88 | 61.93 |
| Instruction Following | ||||||||
| IFEval (en) | 93.70 | 83.12 | 92.75 | 89.21 | 88.85 | 94.96 | 88.13 | 95.35 |
| IFBench (en) | 59.16 | 49.64 | 56.18 | 36.55 | 60.10 | 71.95 | 59.96 | 71.95 |
| M-IFEval-ja (ja) | 81.75 | 63.05 | 77.32 | 63.72 | 73.12 | 79.76 | 76.88 | 87.72 |
| JFBench (ja) | 38.15 | 27.14 | 30.87 | 24.23 | 18.38 | 34.39 | 28.83 | 30.30 |
| Translation | ||||||||
| WMT20 en→ja (ja) | 23.98 | 23.53 | 23.76 | 22.17 | 23.20 | 20.29 | 24.89 | 27.90 |
| WMT20 ja→en (en) | 22.00 | 20.28 | 20.41 | 21.58 | 21.35 | 20.34 | 23.70 | 26.22 |
| Coding | ||||||||
| LiveCodeBench v6 (en) | 53.66 | 26.86 | 32.40 | 56.40 | 59.94 | 75.94 | 74.91 | 77.09 |
| JHumanEval (ja) | 95.37 | 91.52 | 92.38 | 93.72 | 74.63 | 96.83 | 93.54 | 98.35 |
| Function Calling | ||||||||
| BFCL v4 (en) | 36.31 | 13.29 | 27.96 | 42.65 | 48.09 | 62.37 | 67.03 | 66.88 |
| Nejumi BFCL (ja) | 41.89 | 11.58 | 39.19 | 49.03 | 47.10 | 52.12 | 60.62 | 57.34 |
| Average | ||||||||
| Overall | 58.46 | 46.38 | 53.89 | 56.31 | 57.26 | 64.52 | 66.37 | 68.58 |
| Japanese | 59.61 | 49.16 | 56.65 | 56.73 | 55.04 | 62.02 | 64.24 | 66.68 |
| English | 55.94 | 41.16 | 49.72 | 56.84 | 61.45 | 68.98 | 69.98 | 71.55 |
According to evaluations by the publishers, ELYZA-Thinking-1.0-llm-jp-4-32b-a3b demonstrates high performance on Japanese-related tasks. In particular, it records excellent figures in IFEval and M-IFEval-ja, which measure instruction-following capability, as well as in Japanese coding via JHumanEval, showing overall score improvements compared to the base model group. On the other hand, in Function Calling and advanced English mathematics/competitive programming benchmarks (such as AIME and LiveCodeBench v6), it scores lower than some overseas large-scale models and other models, indicating room for improvement in specific English-language tasks and tool selection accuracy.
Strengths and Use Cases
The ELYZA-Thinking-1.0 series is built through a three-stage training process consisting of mid-training, SFT (supervised fine-tuning), and reinforcement learning (RLVR) to enhance Japanese and English language understanding and generation capabilities. It has been trained on localized Japanese reasoning data for mathematics, coding, and STEM fields, specializing in Japanese instruction following and knowledge retrieval. Additionally, it leverages synthetic knowledge data from Japanese Wikipedia and Wikidata to strengthen the understanding and retrieval of Japanese-specific knowledge. Furthermore, it supports tool calling and agent use cases, with agent data localized into Japanese while maintaining tool calls, function names, arguments, and JSON structures.
Hardware Requirements
Estimated requirements (calculated by Local Model Watch) — 32.1B parameters
| Your VRAM | Quantization | File size | Est. memory needed |
|---|---|---|---|
| 80GB class (A100 / H100) | BF16 | 59.9GB | 71.8GB |
Inference engine support (architecture name matched against each project’s own model registry in its source code, checked 2026-10-03): llama.cpp: registered, vLLM: registered, MLX (mlx-lm): registered.
Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.
Can You Run It Locally?
Not usable in Ollama, LM Studio and llama.cpp yet — we have found no GGUF build.
The publisher ships safetensors only. However, llama.cpp’s registry does list this architecture, so conversion to GGUF is possible and the model will run once someone publishes a converted build. Today it can be run with transformers or vLLM, using the memory figures in the table above.
License — apache-2.0 (Commercial use allowed): Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.
Compiled by this site’s code from the published formats, converted builds we have found, and each engine’s own model registry. “Not found" means we have not seen such a build, not that none exists. License summaries are not legal advice — check the publisher’s original terms before relying on them.
Recent Models in the Same Size Class
Models with 15–40B parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site’s estimates; licenses are as stated on the model cards.
→ Scroll horizontally to see all columns
| Model | Parameters | Smallest VRAM tier | License | Our article |
|---|---|---|---|---|
| XingChen-AGI/Xing4.0-29B-A4B | 31.2B | 24GB | apache-2.0 | Xing4.0-29B-A4B Text Generation Model: Our Test Answers, 24GB+ VRAM (2026-09-28) |
| Altworld/Hemmingway-1 | 26.9B | 12GB | cc-by-nc-4.0 | Hemmingway-1 Text Generation Model: 12GB+ VRAM, GGUF Builds (2026-09-28) |
| orcarouter/OrcaSAQ-2-27B | 27.8B | 16GB | apache-2.0 | OrcaSAQ-2-27B Text Generation Model: Our Test Answers, 16GB+ VRAM (2026-09-28) |
| prism-ml/Ternary-Bonsai-2-27B-gguf | 27.8B | 8GB | apache-2.0 | Ternary-Bonsai-2-27B-gguf Text Generation Model: 8GB+ VRAM (2026-09-18) |
| Edge0/Edge0-35B-A3B-preview | 36.0B | 24GB | apache-2.0 | Edge0-35B-A3B-preview 35B MoE Model for Phone-Class Memory: 24GB+ VRAM (2026-09-11) |
How to Get It
The model is distributed in safetensors format and is available on Hugging Face under the Apache-2.0 license. There are no access restrictions or gated license agreements required on Hugging Face to use it.
It is compatible with engines such as vLLM and Transformers, and an example command for launching a server using vLLM is as follows:
vllm serve elyza/ELYZA-Thinking-1.0-llm-jp-4-32b-a3b \
--max-model-len 65536 \
--trust-remote-code \
--moe-backend triton \
--reasoning-parser-plugin vllm_plugins/multi_parser_loader.py \
--reasoning-parser llmjp4 \
--tool-parser-plugin vllm_plugins/llmjp_harmony_tool_parser.py \
--tool-call-parser llmjp_harmony \
--enable-auto-tool-choice
Other Models for the Same Task
Recent text generation models covered by Local Model Watch, newest first. Grouped by the task each publisher declares on Hugging Face (pipeline_tag); the smallest VRAM tier is this site’s estimate.
- Xiaomi Releases MiMo-V2.6-Flash-MOPD and Pro-MOPD Models
- MiMo-V2.6-Pro-RL Text Generation Model: ~417GB Memory, GGUF Builds (Over 80GB)
- DeepSeek-V4-Pro-0813 Text Generation Model: ~998GB Memory, GGUF Builds (1650.5B, Over 80GB)
- Tev1-0.8B-experimental Text Generation Model: 4GB+ VRAM (873M, 4GB)
See all text generation models →
What to Read Next
- Find models by VRAM (This model needs at least 80GB) → VRAM quick reference
- Engines that run this model → vLLM
- Formats this model is available in → Safetensors format guide and models
- Learn about the publisher → ELYZA: models, licenses and articles
- How to read MMLU-Pro, GPQA Diamond, MATH → Benchmark glossary

