Why Sending Only “Hi” Can Still Consume Thousands of Tokens
-
<h2><strong>Introduction</strong></h2><p></p><p>Recently, some users noticed an unusual situation:</p><blockquote><p>“I only sent ‘Hi’, why did it consume 10,000+ tokens?”</p><p></p></blockquote><p>At first glance, this may seem unexpected. A two-character message should only require a few tokens, right?</p><p>The answer is: <strong>the model does not only process your visible message.</strong></p><p>Modern AI assistants, especially coding agents such as Codex, Claude Code, and other IDE-based AI tools, send much more information together with your message to help the model understand the task and provide better responses.</p><p></p><hr><p></p><h2>Your Message Is Only a Small Part of the Context</h2><p>When you send:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Hi</code></pre><p></p><p>the actual request sent to the AI model may look more like this:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>System Instructions
+
User Settings
+
Configured Skills
+
Available Tools
+
Project Information
+
File Context
+
Previous Conversation History
+
Your Message: "Hi"</code></pre><p></p><p>The total size of all these components is the actual context processed by the model.</p><p>Therefore, even a very short message can consume a significant number of tokens.</p><hr><p></p><h2><strong>What Can Increase Context Usage?</strong></h2><h3>1. Configured Skills</h3><p></p><p>If you have enabled custom skills or AI capabilities, their instructions need to be loaded into the model context.</p><p></p><p>For example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Skill A instructions: 2,000 tokens
Skill B instructions: 3,000 tokens
Skill C instructions: 1,500 tokens</code></pre><p>Even before you type anything, these instructions may already occupy several thousand tokens.</p><hr><p></p><h3>2. Tools and Agent Capabilities</h3><p></p><p>AI coding assistants usually provide tools such as:</p><ul><li><p>File reading</p></li><li><p>Code search</p></li><li><p>Terminal execution</p></li><li><p>Git operations</p></li><li><p>Browser access</p></li><li><p>External integrations</p></li></ul><p>The definitions and usage instructions of these tools also consume context space.</p><hr><p></p><h3>3. Project and File Context</h3><p></p><p>When working inside an IDE or code repository, the AI client may automatically include:</p><ul><li><p>Project structure</p></li><li><p>Open files</p></li><li><p>Code snippets</p></li><li><p>Dependencies</p></li><li><p>Configuration files</p></li></ul><p>This allows the AI to understand your project, but it also increases context usage.</p><hr><p></p><h3>4. Previous Conversation History</h3><p></p><p>If you continue an existing conversation, the previous messages are usually included again so the model understands the context.</p><p>For example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Previous discussion: 12,000 tokensNew message:
Hi</code></pre><p>The new request may still contain the previous 12,000 tokens.</p><hr><h2>How to Test Minimum Token Usage</h2><p>If you want to measure the token usage of a simple message, please try the following:</p><ol><li><p>Create a completely new conversation.</p></li><li><p>Disable custom skills.</p></li><li><p>Do not attach files.</p></li><li><p>Do not open a project workspace.</p></li><li><p>Send only:</p></li></ol><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Hi</code></pre><p>This will show the approximate minimum context usage.</p><hr><h2>Context Window vs Token Usage</h2><p>It is important to distinguish between:</p><h3>Context Window</h3><p>The maximum amount of information a model can process in one request.</p><p>Example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>1M token context window</code></pre><p>means the model can theoretically handle up to 1 million tokens.</p><h3>Token Usage</h3><p>The actual amount of tokens included in a specific request.</p><p>Example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>User message:
HiActual request:
System instructions + tools + history + "Hi"Total:
15,000 tokens</code></pre><p>A large context window does not mean every request will only consume the length of your message.</p><p></p><hr><p></p><h2><strong>Conclusion</strong></h2><p></p><p>High token usage from a short message does not necessarily indicate a problem with the model or API.</p><p></p><p>In AI agent environments, the total context includes many hidden components required for intelligent operation. A simple “Hi” can still result in thousands of tokens being processed if the session contains skills, tools, project information, or previous history.</p><p></p><p>We recommend starting a clean session without additional context when testing raw token usage.</p><p></p><p>Thank you for your understanding and continued support.</p>
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Register Login