<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Why Sending Only “Hi” Can Still Consume Thousands of Tokens]]></title><description><![CDATA[<p dir="auto">&lt;h2&gt;&lt;strong&gt;Introduction&lt;/strong&gt;&lt;/h2&gt;&lt;p&gt;&lt;/p&gt;&lt;p&gt;Recently, some users noticed an unusual situation:&lt;/p&gt;&lt;blockquote&gt;&lt;p&gt;“I only sent ‘Hi’, why did it consume 10,000+ tokens?”&lt;/p&gt;&lt;/blockquote&gt;&lt;p&gt;At first glance, this may seem unexpected. A two-character message should only require a few tokens, right?&lt;/p&gt;&lt;p&gt;The answer is: &lt;strong&gt;the model does not only process your visible message.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;Modern AI assistants, especially coding agents such as Codex, Claude Code, and other IDE-based AI tools, send much more information together with your message to help the model understand the task and provide better responses.&lt;/p&gt;&lt;hr&gt;&lt;p&gt;&lt;/p&gt;&lt;h2&gt;Your Message Is Only a Small Part of the Context&lt;/h2&gt;&lt;p&gt;When you send:&lt;/p&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;Hi&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;the actual request sent to the AI model may look more like this:&lt;/p&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;System Instructions<br />
+<br />
User Settings<br />
+<br />
Configured Skills<br />
+<br />
Available Tools<br />
+<br />
Project Information<br />
+<br />
File Context<br />
+<br />
Previous Conversation History<br />
+<br />
Your Message: "Hi"&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;The total size of all these components is the actual context processed by the model.&lt;/p&gt;&lt;p&gt;Therefore, even a very short message can consume a significant number of tokens.&lt;/p&gt;&lt;hr&gt;&lt;h2&gt;What Can Increase Context Usage?&lt;/h2&gt;&lt;h3&gt;1. Configured Skills&lt;/h3&gt;&lt;p&gt;If you have enabled custom skills or AI capabilities, their instructions need to be loaded into the model context.&lt;/p&gt;&lt;p&gt;For example:&lt;/p&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;Skill A instructions: 2,000 tokens<br />
Skill B instructions: 3,000 tokens<br />
Skill C instructions: 1,500 tokens&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Even before you type anything, these instructions may already occupy several thousand tokens.&lt;/p&gt;&lt;hr&gt;&lt;h3&gt;2. Tools and Agent Capabilities&lt;/h3&gt;&lt;p&gt;AI coding assistants usually provide tools such as:&lt;/p&gt;&lt;ul&gt;&lt;li&gt;&lt;p&gt;File reading&lt;/p&gt;&lt;/li&gt;&lt;li&gt;&lt;p&gt;Code search&lt;/p&gt;&lt;/li&gt;&lt;li&gt;&lt;p&gt;Terminal execution&lt;/p&gt;&lt;/li&gt;&lt;li&gt;&lt;p&gt;Git operations&lt;/p&gt;&lt;/li&gt;&lt;li&gt;&lt;p&gt;Browser access&lt;/p&gt;&lt;/li&gt;&lt;li&gt;&lt;p&gt;External integrations&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;The definitions and usage instructions of these tools also consume context space.&lt;/p&gt;&lt;hr&gt;&lt;h3&gt;3. Project and File Context&lt;/h3&gt;&lt;p&gt;When working inside an IDE or code repository, the AI client may automatically include:&lt;/p&gt;&lt;ul&gt;&lt;li&gt;&lt;p&gt;Project structure&lt;/p&gt;&lt;/li&gt;&lt;li&gt;&lt;p&gt;Open files&lt;/p&gt;&lt;/li&gt;&lt;li&gt;&lt;p&gt;Code snippets&lt;/p&gt;&lt;/li&gt;&lt;li&gt;&lt;p&gt;Dependencies&lt;/p&gt;&lt;/li&gt;&lt;li&gt;&lt;p&gt;Configuration files&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;This allows the AI to understand your project, but it also increases context usage.&lt;/p&gt;&lt;hr&gt;&lt;h3&gt;4. Previous Conversation History&lt;/h3&gt;&lt;p&gt;If you continue an existing conversation, the previous messages are usually included again so the model understands the context.&lt;/p&gt;&lt;p&gt;For example:&lt;/p&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;Previous discussion: 12,000 tokens</p>
<p dir="auto">New message:<br />
Hi&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;The new request may still contain the previous 12,000 tokens.&lt;/p&gt;&lt;hr&gt;&lt;h2&gt;How to Test Minimum Token Usage&lt;/h2&gt;&lt;p&gt;If you want to measure the token usage of a simple message, please try the following:&lt;/p&gt;&lt;ol&gt;&lt;li&gt;&lt;p&gt;Create a completely new conversation.&lt;/p&gt;&lt;/li&gt;&lt;li&gt;&lt;p&gt;Disable custom skills.&lt;/p&gt;&lt;/li&gt;&lt;li&gt;&lt;p&gt;Do not attach files.&lt;/p&gt;&lt;/li&gt;&lt;li&gt;&lt;p&gt;Do not open a project workspace.&lt;/p&gt;&lt;/li&gt;&lt;li&gt;&lt;p&gt;Send only:&lt;/p&gt;&lt;/li&gt;&lt;/ol&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;Hi&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;This will show the approximate minimum context usage.&lt;/p&gt;&lt;hr&gt;&lt;h2&gt;Context Window vs Token Usage&lt;/h2&gt;&lt;p&gt;It is important to distinguish between:&lt;/p&gt;&lt;h3&gt;Context Window&lt;/h3&gt;&lt;p&gt;The maximum amount of information a model can process in one request.&lt;/p&gt;&lt;p&gt;Example:&lt;/p&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;1M token context window&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;means the model can theoretically handle up to 1 million tokens.&lt;/p&gt;&lt;h3&gt;Token Usage&lt;/h3&gt;&lt;p&gt;The actual amount of tokens included in a specific request.&lt;/p&gt;&lt;p&gt;Example:&lt;/p&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;User message:<br />
Hi</p>
<p dir="auto">Actual request:<br />
System instructions + tools + history + "Hi"</p>
<p dir="auto">Total:<br />
15,000 tokens&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;A large context window does not mean every request will only consume the length of your message.&lt;/p&gt;&lt;hr&gt;&lt;h2&gt;Conclusion&lt;/h2&gt;&lt;p&gt;High token usage from a short message does not necessarily indicate a problem with the model or API.&lt;/p&gt;&lt;p&gt;In AI agent environments, the total context includes many hidden components required for intelligent operation. A simple “Hi” can still result in thousands of tokens being processed if the session contains skills, tools, project information, or previous history.&lt;/p&gt;&lt;p&gt;We recommend starting a clean session without additional context when testing raw token usage.&lt;/p&gt;&lt;p&gt;Thank you for your understanding and continued support.&lt;/p&gt;</p>
]]></description><link>https://nodebb.ai-router.dev/topic/387/why-sending-only-hi-can-still-consume-thousands-of-tokens</link><generator>RSS for Node</generator><lastBuildDate>Fri, 09 Oct 2026 15:30:23 GMT</lastBuildDate><atom:link href="https://nodebb.ai-router.dev/topic/387.rss" rel="self" type="application/rss+xml"/><pubDate>Sat, 25 Jul 2026 06:59:06 GMT</pubDate><ttl>60</ttl></channel></rss>