{"id":9655,"date":"2026-10-04T02:09:14","date_gmt":"2026-10-03T17:09:14","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/04\/glm-5-3-flash-gguf-released\/"},"modified":"2026-10-05T01:38:15","modified_gmt":"2026-10-04T16:38:15","slug":"glm-5-3-flash-gguf-released","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/04\/glm-5-3-flash-gguf-released\/","title":{"rendered":"GLM-5.3-Flash-GGUF Vision-Language Model: ~150GB Memory"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/ggml-org\/GLM-5.3-Flash-GGUF\">ggml-org\/GLM-5.3-Flash-GGUF<\/a><\/td>\n<\/tr>\n<tr>\n<td>Publisher guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-zai-en\/\">Z.ai (Zhipu AI): models and licenses<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-10-04<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>other<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-gguf-en\/\">GGUF<\/a><\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code when the information was collected. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<p><!-- lmw:lab-summary --><\/p>\n<p><strong>What we checked ourselves<\/strong><\/p>\n<ul>\n<li>It needs about 7% more tokens than the Qwen3 tokenizer for the same Japanese text, so only about 0.94\u00d7 as much Japanese fits in the same context length.<\/li>\n<\/ul>\n<p>Details and conditions are in \u201cOur Own Measurements\u201d below.<\/p>\n<p><!-- \/lmw:lab-summary --><\/p>\n<h2>Overview<\/h2>\n<p>ggml-org has released &#8220;ggml-org\/GLM-5.3-Flash-GGUF&#8221;, which is a GGUF quantized version of the open-weight multimodal model &#8220;GLM-5.3-Flash-BF16&#8221; published by zai-org.<\/p>\n<p>The original model, GLM-5.3-Flash, is the first native multimodal model in the GLM-5 series. It adopts a Mixture of Experts (MoE) configuration with a total parameter count of 320B and 18B active parameters. It is designed to balance efficient long-context processing and high performance by incorporating a hybrid architecture combining sparse and linear attention, along with Manifold-Constrained Hyper-Connections (mHC).<\/p>\n<p>This repository provides the model converted into the GGUF format for use in inference environments such as llama.cpp, including the mmproj for the vision encoder and the MTP sidecar for speculative sampling.<\/p>\n<h2>Specifications<\/h2>\n<ul>\n<li>Total parameters: 320B<\/li>\n<li>Active parameters: 18B<\/li>\n<li>Architecture: MoE (288 routed experts per layer, 1 shared expert), mHC + DSA, 1 MTP layer<\/li>\n<\/ul>\n<h2>Performance<\/h2>\n<p>Since this model is a converted version in GGUF format, performance evaluations are based on the public information of the original model &#8220;GLM-5.3-Flash-BF16&#8221;. Note that slight variations from the original precision may occur due to quantization.<\/p>\n<p>According to explanations from the original model&#8217;s publishers, it delivers performance surpassing the previous-generation GLM-5.2 across various benchmarks and real-world workloads. Furthermore, it is reported to achieve performance levels approaching Claude Opus 4.8 on coding-related and agent-related benchmarks.<\/p>\n<p>The evaluation procedures for the original model cite the following benchmarks:<\/p>\n<ul>\n<li>HLE w\/ tools (full set)<\/li>\n<li>NL2Repo<\/li>\n<li>DeepSWE<\/li>\n<li>Terminal-Bench 2.1<\/li>\n<li>Agent\u2019s Last Exam<\/li>\n<li>Toolathlon Verified<\/li>\n<li>AutomationBench<\/li>\n<li>GDPval-AA v2<\/li>\n<li>BabyVision<\/li>\n<\/ul>\n<p>The original model features a parameter (<code>reasoning_effort<\/code>) to control the amount of reasoning during inference, allowing users to adjust inference capabilities by specifying one of three stages: <code>low<\/code>, <code>high<\/code>, or <code>max<\/code> (default).<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>The base model &#8220;GLM-5.3-Flash-BF16&#8221; is the first in the GLM-5 series to support native multimodal processing. It supports image-text input and conversation processing (<code>image-text-to-text<\/code>), making it useful for general tasks involving visual information.<\/p>\n<p>The main strengths and expected use cases derived from the original model&#8217;s design and evaluation details are as follows:<\/p>\n<h3>Coding and Autonomous Agent Support<\/h3>\n<p>The original model is optimized to demonstrate high capabilities in coding and agent-related benchmarks. Evaluations adopt &#8220;NL2Repo&#8221; for repository generation, &#8220;DeepSWE&#8221; for autonomously solving software engineering tasks, &#8220;Terminal-Bench 2.1&#8221; for evaluating terminal operations, &#8220;Toolathlon Verified&#8221; for measuring tool utilization capabilities, and &#8220;AutomationBench&#8221; for verifying automated workflows. This makes it suitable for complex agent use cases such as not only creating and editing code, but also executing development tools and automating autonomous workflows.<\/p>\n<h3>Efficient Long-Context Processing<\/h3>\n<p>Architecturally, a hybrid structure combining sparse and linear attention is introduced. This aims to significantly reduce inference overhead when handling long contexts while maintaining context-grasping capabilities. It is suited for scenarios where context lengths tend to grow large, such as analyzing massive codebases or handling agent tasks that include multi-step tool-calling histories.<\/p>\n<h3>Thinking Budget Control and Dialogue<\/h3>\n<p>The original model features a parameter (<code>reasoning_effort<\/code>) that controls the amount of reasoning during inference, allowing users to adjust inference cost and depth according to their use case. Additionally, the chat template provides a <code>clear_thinking<\/code> option to control the display of the thought process, and explicitly setting <code>clear_thinking=true<\/code> is recommended for general conversational use. Supported language tags include English (<code>en<\/code>) and Chinese (<code>zh<\/code>).<\/p>\n<h3>High-Speed Inference in Local Environments<\/h3>\n<p>In this GGUF version, in addition to the <code>mmproj<\/code> file for the vision encoder, an MTP (Multi-Token Prediction) sidecar that enables speculative decoding is included. This is expected to achieve faster text generation than standard inference when using a compatible inference engine.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 321.3B parameters (taken from the base model zai-org\/GLM-5.3-Flash-BF16)<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>More than 150GB of VRAM (multi-GPU or CPU offload required)<\/td>\n<td>Q2_K<\/td>\n<td>124.6GB<\/td>\n<td>149.5GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<p><!-- lmw:runnability --><\/p>\n<h2>Can You Run It Locally?<\/h2>\n<p><strong>Runs in Ollama, LM Studio and llama.cpp as-is.<\/strong><\/p>\n<p>It is distributed in GGUF, so no conversion is needed.<\/p>\n<p><strong>License \u2014 <code>other<\/code>:<\/strong> A custom license from the publisher. Check the original terms directly, including whether commercial use is permitted.<\/p>\n<p><strong>Compression:<\/strong> the Q2_K build measures 3.33 bits per weight \u2014 about 21% the size of the original 16-bit weights, calculated by this site from the actual file sizes.<\/p>\n<p><em>Compiled by this site&#8217;s code from the published formats, converted builds we have found, and each engine&#8217;s own model registry. &#8220;Not found&#8221; means we have not seen such a build, not that none exists. License summaries are not legal advice \u2014 check the publisher&#8217;s original terms before relying on them.<\/em><\/p>\n<p><!-- \/lmw:runnability --><\/p>\n<p><!-- lmw:lab --><\/p>\n<h2>Our Own Measurements<\/h2>\n<p>Values we checked ourselves on our server (no GPU) by actually reading this model&#8217;s files, without running the model \u2014 not figures copied from the model card. How we measure, and the results for every model: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/observations-en\/\">Our Measurements<\/a>.<\/p>\n<h3>Measurements We Did Not Take<\/h3>\n<p>We have not confirmed that this site meets the commercial-use terms of this model&#8217;s license (<code>other<\/code>). Because this site carries advertising, we did not run the model (no CPU run, answers, quantization comparison, conversion or generation). Only values we checked without running the model, such as token counts and the GGUF header, are shown.<\/p>\n<h3>Japanese Token Efficiency<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Tokenizer<\/th>\n<th>Tokens per 1,000 Japanese characters<\/th>\n<th>Ratio to the same text in English<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>This model<\/strong><\/td>\n<td><strong>733<\/strong><\/td>\n<td><strong>1.34\u00d7<\/strong><\/td>\n<\/tr>\n<tr>\n<td>Qwen3<\/td>\n<td>688<\/td>\n<td>1.26\u00d7<\/td>\n<\/tr>\n<tr>\n<td>Llama 3.2<\/td>\n<td>744<\/td>\n<td>1.36\u00d7<\/td>\n<\/tr>\n<tr>\n<td>Gemma 3<\/td>\n<td>564<\/td>\n<td>1.03\u00d7<\/td>\n<\/tr>\n<tr>\n<td>gpt-oss<\/td>\n<td>795<\/td>\n<td>1.45\u00d7<\/td>\n<\/tr>\n<tr>\n<td>LLM-jp-3<\/td>\n<td>497<\/td>\n<td>0.85\u00d7<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>It needs about 7% more tokens than the Qwen3 tokenizer for the same Japanese text, so only about 0.94\u00d7 as much Japanese fits in the same context length.<\/p>\n<p>Counted with the <code>tokenizer.json<\/code> of <a href=\"https:\/\/huggingface.co\/zai-org\/GLM-5.3-Flash-BF16\">zai-org\/GLM-5.3-Flash-BF16<\/a> on a fixed text we wrote ourselves (876 Japanese characters across news, conversation, technical docs, a formal email, travel writing and a recipe) and its English translation. Fewer tokens mean more Japanese fits in the context window.<\/p>\n<h3>Inside the GGUF File<\/h3>\n<p>File: <a href=\"https:\/\/huggingface.co\/ggml-org\/GLM-5.3-Flash-GGUF\/blob\/main\/GLM-5.3-Flash-Q2_K-00001-of-00002.gguf\">GLM-5.3-Flash-Q2_K-00001-of-00002.gguf<\/a> (124.62GB, <code>Q2_K<\/code>). We read only the header (metadata) of the file, not the weights.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Architecture (as named in the GGUF)<\/td>\n<td><code>glm5-next<\/code><\/td>\n<\/tr>\n<tr>\n<td>Maximum trained context length<\/td>\n<td>1,048,576 tokens<\/td>\n<\/tr>\n<tr>\n<td>Layers<\/td>\n<td>45<\/td>\n<\/tr>\n<tr>\n<td>Experts<\/td>\n<td>8 active out of 288<\/td>\n<\/tr>\n<tr>\n<td>Vocabulary size<\/td>\n<td>154,880<\/td>\n<\/tr>\n<tr>\n<td>Chat template<\/td>\n<td>Included (mentions tool calls, has a thinking switch)<\/td>\n<\/tr>\n<tr>\n<td>imatrix<\/td>\n<td>Not recorded in the file<\/td>\n<\/tr>\n<tr>\n<td>Weight types (share of parameters)<\/td>\n<td>Q2_K 64.8% \/ Q4_K 32.4% \/ Q8_0 2.7% \/ BF16 0.1% \/ other 0.0%<\/td>\n<\/tr>\n<tr>\n<td>Average bits per weight<\/td>\n<td>3.42 bits<\/td>\n<\/tr>\n<tr>\n<td>Embedding \/ output layer type<\/td>\n<td>Q8_0 \/ Q8_0<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>The quant name in the file name describes the file as a whole; in practice layers mix several types. The average bits per weight is the measured file data divided by the number of weights.<\/p>\n<p><!-- \/lmw:lab --><\/p>\n<p><!-- lmw:peers --><\/p>\n<h2>Recent Models in the Same Size Class<\/h2>\n<p><em>Models with <\/em><em>over 40B<\/em><em> parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site&#8217;s estimates; licenses are as stated on the model cards.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Parameters<\/th>\n<th>Smallest VRAM tier<\/th>\n<th>License<\/th>\n<th>Our article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Qwen\/Qwen3.8-Flash-Next<\/td>\n<td>180.0B<\/td>\n<td>\u2014<\/td>\n<td>qwen-community-1.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/05\/qwen-3-8-flash-next-released\/\">Qwen3.8-Flash-Next Multimodal MoE Model: ~402GB Memory<\/a> (2026-10-04)<\/td>\n<\/tr>\n<tr>\n<td>bartowski\/Intern-S2-397B-GGUF<\/td>\n<td>403.4B<\/td>\n<td>\u2014<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/intern-s2-397b-gguf-released\/\">Intern-S2-397B-GGUF Vision-Language Model: ~102GB Memory<\/a> (2026-09-15)<\/td>\n<\/tr>\n<tr>\n<td>deepseek-ai\/DeepSeek-V4.1-Flash<\/td>\n<td>763.2B<\/td>\n<td>\u2014<\/td>\n<td>mit<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/10\/deepseek-v4-1-flash-released\/\">DeepSeek-V4.1-Flash 552B Multimodal MoE Model: ~570GB Memory<\/a> (2026-09-10)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:peers --><\/p>\n<h2>How to Get It<\/h2>\n<p>This model is distributed in GGUF format on the Hugging Face repository &#8220;ggml-org\/GLM-5.3-Flash-GGUF&#8221;. Since it is not a gated model, you can download and use it without prior application or consent procedures.<\/p>\n<p>When using CLI tools from the official <code>llama.cpp<\/code> family, you can start a server by directly specifying the Hugging Face repository with the following command:<\/p>\n<pre><code class=\"language-bash\">llama serve -hf ggml-org\/GLM-5.3-Flash-GGUF\n<\/code><\/pre>\n<p>The internal specifications of the quantized versions provided in this repository have the following features:<\/p>\n<ul>\n<li><code>Q4_K<\/code>: All routed experts are quantized with Q4_K.<\/li>\n<li><code>Q2_K<\/code>: Composed of Q4_K for routed down experts and Q2_K for gate\/up experts.<\/li>\n<li>Smaller non-expert tensors, such as embedding layers, attention layers, shared experts, and dense FFNs, are kept as Q8_0 to maintain quality.<\/li>\n<li>Bundled with mmproj (Q8_0) for the vision encoder and MTP sidecars (Q8_0 \/ Q4_0) for speculative sampling.<\/li>\n<li>Note that low-bit quantization is noted as not having undergone calibration using imatrix (importance matrix).<\/li>\n<\/ul>\n<p>Regarding the license classification, the base model &#8220;GLM-5.3-Flash-BF16&#8221; is set under the MIT license, but this GGUF repository carries an <code>other<\/code> tag. Please check the repository&#8217;s license terms in advance before use.<\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/ggml-org-openjev-gguf-released-2\/\">OpenJev-GGUF 27.4B Decision-Making Specialized Model: 24GB+ VRAM<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/mimo-v2-6-flash-rl-gguf-released\/\">MiMo-V2.6-Flash-RL-GGUF Multimodal MoE Model: ~141GB Memory<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/testing-openjev-gguf-tiny-model-for-runtime-loaders\/\">Testing OpenJev GGUF: A Tiny Model for Runtime Loaders<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/mimo-v2-6-distill-qwen-9b-gguf\/\">MiMo-V2.6-Distill-Qwen-9B-GGUF: Our Test Answers, 12GB+ VRAM<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Find models by VRAM<\/strong> (This model does not fit even in 80GB; its smallest build needs about 150GB) \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference, including models over 80GB<\/a><\/li>\n<li><strong>Engines that run this model<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-llama-cpp-en\/\">llama.cpp<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-ollama-en\/\">Ollama<\/a><\/li>\n<li><strong>What Q2_K mean and where to get this model<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-gguf-en\/\">GGUF format guide and models<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">Quantization and model-format glossary<\/a><\/li>\n<li><strong>Learn about the publisher<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-zai-en\/\">Z.ai (Zhipu AI): models, licenses and articles<\/a><\/li>\n<li><strong>Other models for the same task<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/models-by-task-en\/#task-vision\">Other vision-language models<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li>ggml-org\/GLM-5.3-Flash-GGUF: <a href=\"https:\/\/huggingface.co\/ggml-org\/GLM-5.3-Flash-GGUF\">https:\/\/huggingface.co\/ggml-org\/GLM-5.3-Flash-GGUF<\/a><\/li>\n<li>zai-org\/GLM-5.3-Flash-BF16: <a href=\"https:\/\/huggingface.co\/zai-org\/GLM-5.3-Flash-BF16\">https:\/\/huggingface.co\/zai-org\/GLM-5.3-Flash-BF16<\/a><\/li>\n<\/ul>\n<p><!-- lmw:updates --><\/p>\n<h2>Update History<\/h2>\n<ul>\n<li>2026-10-05: Added our own measurements: Japanese token efficiency, what is inside the GGUF file.<\/li>\n<\/ul>\n<p><!-- \/lmw:updates --><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Discover ggml-org\/GLM-5.3-Flash-GGUF, a GGUF quantization of zai-org&#8217;s native multimodal MoE model, featuring vision encoders and MTP sidecars.<\/p>\n","protected":false},"author":1,"featured_media":9654,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[310],"tags":[161,1970,163,1469,2950,1073,165,1547,1952],"class_list":["post-9655","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-new-models","tag-ggml-en","tag-ggml-org-en","tag-gguf-en","tag-glm-5-3-flash-en","tag-glm-5-3-flash-gguf-en","tag-llama-cpp-en","tag-moe-en","tag-verified","tag-vlm-en"],"lang":"en","translations":{"en":9655,"ja":9653},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9655","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=9655"}],"version-history":[{"count":2,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9655\/revisions"}],"predecessor-version":[{"id":9905,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9655\/revisions\/9905"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/9654"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=9655"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=9655"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=9655"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}