Hemmingway-1 Text Generation Model: 12GB+ VRAM, GGUF Builds

- 1. At a Glance
- 2. Overview
- 3. Specifications
- 4. Performance
- 5. Strengths and Use Cases
- 6. How It Differs from Similar Models
- 7. Hardware Requirements
- 8. Can You Run It Locally?
- 9. Our Own Measurements
- 10. Recent Models in the Same Size Class
- 11. How to Get It
- 12. Quantized and Converted Variants
- 13. Related Articles
- 14. What to Read Next
- 15. Sources
- 16. Update History
At a Glance
| Item | Value |
|---|---|
| Repository | Altworld/Hemmingway-1 |
| Family guide | Qwen3.8 guide (8 articles) |
| Publisher guide | Alibaba (Qwen): models and licenses |
| Published | 2026-09-20 |
| License | cc-by-nc-4.0 |
| Formats | safetensors |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code at collection time. Dates are JST.
Overview
Altworld has released Hemmingway-1, a language model specialized in writing everyday text “as if written by a human," on Hugging Face. It is a 27B model built by fine-tuning Qwen3.8-27B, with a context length of 262,144 tokens. The license is CC BY-NC 4.0, making it free for personal, research, and other non-commercial uses, while commercial use requires a separate agreement with the publishers.
The publishers state a clear rationale for the model. When asked to “write a message to my landlord," many models return three options, an introduction, and explanations of the options. Hemmingway-1 was created to return only the message itself. It targets the kinds of texts people write every day, such as messages, emails, slightly awkward communications with colleagues, and procrastination-inducing tasks.
The publishers also offer an app with the same name (for Mac and Android) as well as a web version, and this weight release corresponds to opening up the underlying model. The base Qwen3.8-27B model was covered in Qwen3.8-27B Multimodal Vision-Language Model: 8GB+ VRAM, GGUF Builds.
Specifications
From the model card table:
- Parameters: 27B
- Base Model: Qwen3.8-27B
- Context Length: 262,144 tokens
- Languages: Primarily English (the model card explicitly states “English-first")
- License: CC BY-NC 4.0 (Free for non-commercial use. Separate agreement required for commercial use)
While the base Qwen3.8-27B is a multimodal model that can also take images as input, Hemmingway-1’s configuration file is set up exclusively for text (Qwen3_5ForCausalLM), and image input is not intended. The distribution files include weights for MTP (Multi-Token Prediction) for predicting subsequent tokens as a separate file.
Performance
All evaluations in the model card are presented in graphs. The figures below are read from the values depicted in those graphs. CommunicationBench, Human-Likeness, StoryBench, and The Message, Not a Memo are benchmarks created by the publishers themselves, which the publishers explicitly state. The evaluations are blind, one-on-one comparisons of model responses, conducted twice with swapped order, and a model different from the evaluation targets was used for judging.
Everyday Writing (CommunicationBench): Scores based on head-to-head comparisons with other models across 80 real-world requests.
| Model | Score |
|---|---|
| Hemmingway-1 | 1026 |
| Fable 5.1 | 1024 |
| Fable 5 | 1009 |
| GLM-5.3 | 1007 |
| Kimi K3 | 996 |
| GPT-6 Astra | 976 |
| Grok 4.6 | 967 |
| DeepSeek V4 Pro | 963 |
| Qwen3.8 27B base | 954 |
Although it takes the top spot, the gap with the second-place Fable 5.1 is only 2 points. Given this type of head-to-head comparison score, it is reasonable to view this within the margin of error, and accurate to say it “matches top-tier commercial models." On the other hand, it is up 72 points from the base Qwen3.8-27B, clearly showing the effectiveness of the fine-tuning.
Human-Likeness: Scores determined by asking evaluators in the same comparison “which response was written by a human." Hemmingway-1 scored 1032, leading the second-place Fable 5.1 (1006) by a 26-point margin. GLM-5.3 follows at 996, Fable 5 at 992, Kimi K3 at 985, GPT-6 Astra at 964, and the base Qwen3.8 27B at 952. The gap here is clear, and it can be said that this model’s greatest strength is “human-likeness" rather than “cleverness." Broken down by request type, it performs strongly in financial/procedural matters, work communications, difficult-to-write requests that require multiple rewrites, and persuasive messaging. For difficult-to-write requests, the percentage of judgments considering it “written by a human" was 72% compared to GPT-6 Astra’s 9%. Conversely, it loses to storytelling-oriented models in adversarial stories and long-form narrative exchanges.
Excessive Introductions (The Message, Not a Memo): The percentage of responses where the core text is buried within explanations and options (lower is better).
| Model | Percentage (%) |
|---|---|
| GPT-6 Astra | 4 |
| Qwen3.8 27B base | 19 |
| Grok 4.6 | 22 |
| Hemmingway-1 | 39 |
| Fable 5.1 | 67 |
| Kimi K3 | 92 |
| Fable 5 | 93 |
| GLM-5.3 | 93 |
The main text of the model card states that “Fable 5, GLM-5.3, and Kimi K3 bury the text in 9 out of 10 cases," but looking closely at the graph, Hemmingway-1 (39%) includes more introductions than GPT-6 Astra (4%), the base Qwen3.8 27B (19%), and Grok 4.6 (22%). Contrary to the selling point of “returning only the text itself," it is not the best in this category, and moreover, fine-tuning has made it worse than the base model. If one emphasizes “returning only the message itself," this point should be discounted.
Emotion Comprehension (EQ-Bench 4): This is the only benchmark published by a third party rather than the publishers. Hemmingway-1 scored 1330, ranking third behind Fable 5 (1341) and Kimi K3 (1332), and outperforming GPT-5.5 (1316), Opus 4.7 (1312), and Opus 4.8 (1285). The fact that a 27B model falls within an 11-point difference of top commercial models in a third-party evaluation is a result that can be trusted more than the self-made benchmarks.
Creative Writing (StoryBench): Scored 1197, tying with Kimi K3. While it falls short of Fable 5 Max (1277) and GLM-5.3 (1254), it outperforms Qwen3.8-Max (1081) and DeepSeek V4 Pro (1054), and is up 504 points from the base Qwen3.8 27B (693).
Strengths and Use Cases
Use cases are clearly narrowed down: writing short, practical everyday text in English for people. This includes emails, messages, communications with colleagues or landlords, and difficult-to-write messages. For long-form storytelling, the publishers themselves acknowledge that storytelling-oriented models are superior.
As a caveat, the publishers explicitly state: “Primarily in English, and may write with confidence even when wrong. Do not use for medical, legal, or financial judgments." A model optimized to write human-like text will also write errors in a natural, human-like manner. Factual accuracy of texts must be verified independently before sending.
Evaluations for Japanese texts are not provided in the model card.
How It Differs from Similar Models
- Difference from the base Qwen3.8-27B: It adopts a text-only configuration, with significant score increases in everyday writing (+72), human-likeness (+80), and creative writing (+504). On the other hand, the percentage of adding introductions increased from 19% to 39%.
- Difference from general-purpose top-tier models: Virtually tied with Fable 5.1 in overall everyday writing scores, and clearly superior in human-likeness. As a 27B model that can be run locally, it reaches a level comparable to top commercial models specifically for this use case. However, it must be discounted that most comparisons are based on the publishers’ own evaluations.
- Difference from storytelling models: Falls short of Fable 5.1 Max and GLM-5.3 in creative StoryBench, and loses in adversarial and long-form narrative exchanges.
Hardware Requirements
Estimated requirements (calculated by Local Model Watch) — 26.9B parameters
| Your VRAM | Quantization | File size | Est. memory needed |
|---|---|---|---|
| 12GB (RTX 4070 / 3060 12GB, etc.) | IQ2_M | 9.8GB | 11.8GB |
| 16GB (RTX 5060 Ti 16GB / 4060 Ti 16GB, etc.) | Q3_K_L | 13.2GB | 15.8GB |
| 24GB (RTX 4090 / 3090, etc.) | Q5_K_M | 19.5GB | 23.4GB |
| 32GB (RTX 5090, etc.) | Q6_K_L | 23.2GB | 27.9GB |
| 48GB (RTX 6000 Ada / A6000, etc.) | Q8_0 | 27.1GB | 32.5GB |
| 80GB class (A100 / H100) | BF16 | 50.9GB | 61.1GB |
Inference engine support (architecture name matched against each project’s own model registry in its source code, checked 2026-09-28): llama.cpp: registered, vLLM: registered, MLX (mlx-lm): not registered. “Not registered" means the name is absent from that registry today, not that the model cannot run.
Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. File sizes are measured from the converted build bartowski/Altworld_Hemmingway-1-GGUF. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.
Can You Run It Locally?
Runs in Ollama, LM Studio and llama.cpp via a converted build.
The publisher ships safetensors, but bartowski/Altworld_Hemmingway-1-GGUF provides a GGUF build you can use.
License — cc-by-nc-4.0 (Commercial use prohibited):
Commercial use is prohibited (NC = NonCommercial). Internal business use can also count as commercial, so avoid this license for work use.
Compression: the IQ2_M build measures 3.13 bits per weight — about 20% the size of the original 16-bit weights, calculated by this site from the actual file sizes.
Compiled by this site’s code from the published formats, converted builds we have found, and each engine’s own model registry. “Not found" means we have not seen such a build, not that none exists. License summaries are not legal advice — check the publisher’s original terms before relying on them.
Our Own Measurements
Values we measured ourselves on our server (no GPU) by actually reading and running this model’s files — not figures copied from the model card. How we measure, and the results for every model: Our Measurements.
Japanese Token Efficiency
| Tokenizer | Tokens per 1,000 Japanese characters | Ratio to the same text in English |
|---|---|---|
| This model | 546 | 0.99× |
| Qwen3 | 688 | 1.26× |
| Llama 3.2 | 744 | 1.36× |
| Gemma 3 | 564 | 1.03× |
| gpt-oss | 795 | 1.45× |
| LLM-jp-3 | 497 | 0.85× |
It needs about 21% fewer tokens than the Qwen3 tokenizer for the same Japanese text, so about 1.26× as much Japanese fits in the same context length, and generation is faster per character.
Counted with the tokenizer.json of Altworld/Hemmingway-1 on a fixed text we wrote ourselves (876 Japanese characters across news, conversation, technical docs, a formal email, travel writing and a recipe) and its English translation. Fewer tokens mean more Japanese fits in the context window.
Inside the GGUF File
File: Altworld_Hemmingway-1-Q4_K_M.gguf (16.24GB, Q4_K_M). We read only the header (metadata) of the file, not the weights.
| Item | Value |
|---|---|
| Architecture (as named in the GGUF) | qwen35 |
| Maximum trained context length | 262,144 tokens |
| Layers | 65 |
| Vocabulary size | 248,320 |
| Chat template | Included (mentions tool calls, has a thinking switch) |
| imatrix | Used (Altworld_Hemmingway-1-calibration-v6.txt, 582 chunks) |
| Weight types (share of parameters) | Q4_K 74.5% / Q6_K 18.4% / Q8_0 4.8% / Q4_0 1.6% / other 0.7% |
| Average bits per weight | 5.10 bits |
| Embedding / output layer type | Q4_K / Q6_K |
The quant name in the file name describes the file as a whole; in practice layers mix several types. The average bits per weight is the measured file data divided by the number of weights.
Recent Models in the Same Size Class
Models with 15–40B parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site’s estimates; licenses are as stated on the model cards.
→ Scroll horizontally to see all columns
| Model | Parameters | Smallest VRAM tier | License | Our article |
|---|---|---|---|---|
| XingChen-AGI/Xing4.0-29B-A4B | 31.2B | 24GB | apache-2.0 | Xing4.0-29B-A4B 29B MoE Model Strong in Coding Agents: 24GB+ VRAM (2026-09-28) |
| Edge0/Edge0-35B-A3B-preview | 36.0B | 24GB | apache-2.0 | Edge0-35B-A3B-preview 35B MoE Model for Phone-Class Memory: 24GB+ VRAM (2026-09-11) |
| bartowski/Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF | 26.5B | 12GB | apache-2.0 | Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF: 12GB+ VRAM (2026-09-11) |
| nex-agi/Nex-N2.5-mini | 35.1B | 16GB | apache-2.0 | Nex-N2.5-mini Agent Model for Long-Horizon Tasks: 16GB+ VRAM (2026-09-09) |
How to Get It
- Distribution format: Transformers-format (safetensors) weights are available at
Altworld/Hemmingway-1on Hugging Face. - Example download command:
huggingface-cli download Altworld/Hemmingway-1 - Supported engines: The model card provides examples for launching with vLLM (
vllm serve Altworld/Hemmingway-1 --max-model-len 262144) and Transformers. The model card does not mention llama.cpp, Ollama, or LM Studio. - Running on local hardware: There is no official GGUF version provided by the publishers. When looking for quantized versions, please verify that the conversion source is
Altworld/Hemmingway-1. - License: CC BY-NC 4.0. Non-commercial use allows downloading, running, fine-tuning, and sharing, provided that Hemmingway-1 is credited. Commercial use requires a separate agreement with the publishers. Please check the original license text before use.
Quantized and Converted Variants
→ Scroll horizontally to see all columns
| Added | Publisher | Format | Repository | Smallest VRAM tier (build, est. memory) |
|---|---|---|---|---|
| 2026-09-28 | bartowski | GGUF (imatrix) | bartowski/Altworld_Hemmingway-1-GGUF | IQ2_M 11.8GB (fits in 12GB VRAM) |
| 2026-09-28 | mradermacher | GGUF | mradermacher/Hemmingway-1-GGUF | Q3_K_M 15.1GB (fits in 16GB VRAM) |
| 2026-09-28 | mlx-community | MLX | mlx-community/Hemmingway-1-OptiQ-4bit | MLX 4bit 20.8GB (fits in 24GB VRAM) |
File sizes of each build:
- Available builds in bartowski/Altworld_Hemmingway-1-GGUF: IQ2_XXS 8.3GB / IQ2_XS 8.5GB / IQ2_S 9.0GB / IQ2_M 9.8GB / Q2_K 10.1GB / IQ3_XXS 11.5GB / Q3_K_S 11.9GB / IQ3_XS 11.9GB / Q3_K_M 12.5GB / Q3_K_L 13.2GB / IQ3_M 13.8GB / IQ4_XS 14.4GB / Q4_0 15.2GB / Q4_K_S 15.2GB / IQ4_NL 16.2GB / Q4_K_M 16.2GB / Q4_1 16.6GB / Q4_K_L 17.5GB / Q5_K_S 18.2GB / Q5_K_M 19.5GB / Q6_K_S 21.3GB / Q6_K 22.2GB / Q6_K_L 23.2GB / Q8_0 27.1GB / BF16 50.9GB
- Available builds in mradermacher/Hemmingway-1-GGUF: Q2_K 10.1GB / Q3_K_S 11.4GB / Q3_K_M 12.6GB / Q3_K_L 13.6GB / IQ4_XS 14.4GB / Q4_K_S 14.7GB / Q4_K_M 15.7GB / Q5_K_S 17.7GB / Q5_K_M 18.2GB / Q6_K 20.9GB / Q8_0 27.1GB
- Available builds in mlx-community/Hemmingway-1-OptiQ-4bit: MLX 4bit 17.4GB
In addition, 27 converted build(s) from other uploaders exist on Hugging Face; this site lists only builds from the model’s publisher or established quantization maintainers.
This section is appended automatically by Local Model Watch when a converted build of this model appears after publication. Memory figures are estimated from the size of the distributed files. What the quantization names mean: glossary.
Related Articles
- OrcaSAQ-2-27B Text Generation Model: 16GB+ VRAM
- Qwen3.8-27B TWIN-TURBO Uncensored GGUF Released
- Qwen3.8-27B Multimodal Vision-Language Model: 8GB+ VRAM, GGUF Builds
- vectionlabs_Salience-27B-R6-GGUF Vision-Language Model: 12GB+ VRAM
What to Read Next
- Find models by VRAM (This model runs from the 12GB tier) → Other models that run on a 12GB GPU
- Explore the same model family → Qwen3.8 family overview (8 articles, 11 converted builds)
- Engines that run this model → llama.cpp / Ollama / vLLM
- What IQ2_M, Q3_K_L, Q5_K_M mean and where to get this model → GGUF format guide and models / MLX format guide and models / Quantization and model-format glossary
- Learn about the publisher → Alibaba (Qwen): models, licenses and articles
Sources
Update History
- 2026-09-28: Added converted builds to “Quantized and Converted Variants”: bartowski/Altworld_Hemmingway-1-GGUF, mradermacher/Hemmingway-1-GGUF, mlx-community/Hemmingway-1-OptiQ-4bit
- 2026-09-28: Updated the hardware requirements table with the actual file sizes of bartowski/Altworld_Hemmingway-1-GGUF.
- 2026-09-28: Changed the title to show what the article covers (VRAM requirements, file list, etc.).
- 2026-09-28: Added our own measurements: Japanese token efficiency, what is inside the GGUF file.

