{"id":645,"date":"2026-09-15T10:09:03","date_gmt":"2026-09-15T01:09:03","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/15\/migrating-large-prompts-to-local-ollama\/"},"modified":"2026-09-20T17:37:25","modified_gmt":"2026-09-20T08:37:25","slug":"migrating-large-prompts-to-local-ollama","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/migrating-large-prompts-to-local-ollama\/","title":{"rendered":"Challenges of Migrating AI Agents to Local Ollama Environments"},"content":{"rendered":"<h2>Overview<\/h2>\n<p>As developers attempt to migrate agents from commercial LLM APIs like Claude (Opus) to self-hosted environments like Ollama for privacy protection, technical challenges caused by massive pre-prompts are becoming a subject of debate. In particular, the focus is on the &#8220;thrashing&#8221; phenomenon, where agents exhibit unintended behavior due to context window limitations, as well as the nature of prompt design.<\/p>\n<h2>Why It Is Being Discussed<\/h2>\n<p>This discussion was sparked when a developer shared their experience of attempting to migrate an agent using a massive 35kb pre-prompt from a commercial API to a local Ollama environment. The post gained traction on Hacker News, garnering 139 points and 76 comments, attracting the interest of many engineers.<\/p>\n<p>In the background, there are strong concerns regarding data collection by frontend providers and the unintended use of session metadata. The developer argues for the necessity of &#8220;Bring Your Own Weights (BYOW)&#8221;\u2014performing inference on one&#8217;s own hardware to protect ideas and highly confidential work.<\/p>\n<h2>Discussion Points<\/h2>\n<p>The discussion is mainly centered around three technical points:<\/p>\n<h3>Context Depletion and Agent Behavior<\/h3>\n<p>While commercial APIs offer massive context windows, local environments face strict limitations. For instance, in a system with a 65k tokens context window, a 35kb prompt immediately consumes 14% of the total. When the context saturates, a phenomenon known as &#8220;thrashing&#8221; occurs, where the agent repeats the same tool calls, re-reads already loaded files, or rewrites completed work. It is reported that this puts the agent in a state where it cannot retain immediate instructions, resembling giving instructions to &#8220;a human who reincarnates every 90 seconds.&#8221;<\/p>\n<h3>Prompt Design Dependencies and the &#8220;Chain of Thought&#8221; Trap<\/h3>\n<p>It has been pointed out that the massive context windows provided by frontend providers may actually be masking poor prompt design. When a model uses Chain of Thought (CoT) for reasoning, a wide context allows the model to internally infer and compensate even for improperly constructed prompts. However, in local environments with restricted contexts, this &#8220;hidden compensation&#8221; fails, bringing design flaws in the prompt to the surface.<\/p>\n<h3>Hardware and Tool Constraints in Local Execution<\/h3>\n<p>Migrating to a self-hosted environment requires immense hardware resources. Additionally, software choices are debated, with technical considerations regarding optimizing the runtime environment, such as choosing between Ollama or the lower-layer llama.cpp.<\/p>\n<h2>Community Reactions<\/h2>\n<p>Various opinions are being exchanged within the community regarding prompt design philosophies and the realities of local environments.<\/p>\n<h3>Critical Perspectives on Prompt Design<\/h3>\n<p>Some participants pointed out harshly that using a massive 35kb prompt itself is confusing and inefficient for any LLM. Opinions suggested that excessively long prompts reduce the effective context where the model can accurately pay attention, and thus prompts should be split by a single purpose, advocating for more declarative agent definitions.<\/p>\n<h3>Concerns Over the Feasibility of Local Environments<\/h3>\n<p>Many comments addressed the high hardware requirements for operating local models. Discussions highlighted that even with expensive equipment, performance may not meet expectations, and securing practical context sizes locally remains difficult. Doubts were also raised as to whether relatively small models, such as 27b parameter models, can adequately perform advanced security tasks.<\/p>\n<h3>Trade-offs Between Privacy and Convenience<\/h3>\n<p>On the other hand, many voices agreed with the distrust toward commercial providers&#8217; data handling practices. Some opinions viewed the movement to avoid over-reliance on cloud computing and build one&#8217;s own environment as inevitable from a security standpoint. Furthermore, movements to explore concrete workarounds were shown, such as utilizing new tools to assist with context management or logging session state to disk to perform frequent handoffs.<\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/patrickmccanna.net\/notes-on-migrating-large-prompts-away-from-anthropic-openai-to-self-hosted-llms\/\">https:\/\/patrickmccanna.net\/notes-on-migrating-large-prompts-away-from-anthropic-openai-to-self-hosted-llms\/<\/a><\/li>\n<li><a href=\"https:\/\/news.ycombinator.com\/item?id=49697014\">https:\/\/news.ycombinator.com\/item?id=49697014<\/a><\/li>\n<\/ul>\n<p><!-- lmw:updates --><\/p>\n<h2>Update History<\/h2>\n<ul>\n<li>2026-09-16: Verified the content against the official primary source.<\/li>\n<li>2026-09-19: Rewrote the article from re-collected sources and restored it from draft to published.<\/li>\n<\/ul>\n<p><!-- \/lmw:updates --><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Discussing the technical hurdles, context thrashing, and prompt design issues when moving large agents from commercial LLM APIs to self-hosted Ollama.<\/p>\n","protected":false},"author":1,"featured_media":644,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[130],"tags":[1664,163,136,1667,1093,1547],"class_list":["post-645","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-community","tag-claude-en","tag-gguf-en","tag-llm-en","tag-local-llm-en","tag-ollama-en","tag-verified"],"lang":"en","translations":{"en":645,"ja":643},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/645","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=645"}],"version-history":[{"count":8,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/645\/revisions"}],"predecessor-version":[{"id":2317,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/645\/revisions\/2317"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/644"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=645"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=645"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=645"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}