Skip to content
  • 0 Votes
    1 Posts
    7 Views
    A
    <h2><strong>Introduction</strong></h2><p></p><p>Recently, some users noticed an unusual situation:</p><blockquote><p>“I only sent ‘Hi’, why did it consume 10,000+ tokens?”</p></blockquote><p>At first glance, this may seem unexpected. A two-character message should only require a few tokens, right?</p><p>The answer is: <strong>the model does not only process your visible message.</strong></p><p>Modern AI assistants, especially coding agents such as Codex, Claude Code, and other IDE-based AI tools, send much more information together with your message to help the model understand the task and provide better responses.</p><hr><p></p><h2>Your Message Is Only a Small Part of the Context</h2><p>When you send:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Hi</code></pre><p>the actual request sent to the AI model may look more like this:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>System Instructions + User Settings + Configured Skills + Available Tools + Project Information + File Context + Previous Conversation History + Your Message: "Hi"</code></pre><p>The total size of all these components is the actual context processed by the model.</p><p>Therefore, even a very short message can consume a significant number of tokens.</p><hr><h2>What Can Increase Context Usage?</h2><h3>1. Configured Skills</h3><p>If you have enabled custom skills or AI capabilities, their instructions need to be loaded into the model context.</p><p>For example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Skill A instructions: 2,000 tokens Skill B instructions: 3,000 tokens Skill C instructions: 1,500 tokens</code></pre><p>Even before you type anything, these instructions may already occupy several thousand tokens.</p><hr><h3>2. Tools and Agent Capabilities</h3><p>AI coding assistants usually provide tools such as:</p><ul><li><p>File reading</p></li><li><p>Code search</p></li><li><p>Terminal execution</p></li><li><p>Git operations</p></li><li><p>Browser access</p></li><li><p>External integrations</p></li></ul><p>The definitions and usage instructions of these tools also consume context space.</p><hr><h3>3. Project and File Context</h3><p>When working inside an IDE or code repository, the AI client may automatically include:</p><ul><li><p>Project structure</p></li><li><p>Open files</p></li><li><p>Code snippets</p></li><li><p>Dependencies</p></li><li><p>Configuration files</p></li></ul><p>This allows the AI to understand your project, but it also increases context usage.</p><hr><h3>4. Previous Conversation History</h3><p>If you continue an existing conversation, the previous messages are usually included again so the model understands the context.</p><p>For example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Previous discussion: 12,000 tokens New message: Hi</code></pre><p>The new request may still contain the previous 12,000 tokens.</p><hr><h2>How to Test Minimum Token Usage</h2><p>If you want to measure the token usage of a simple message, please try the following:</p><ol><li><p>Create a completely new conversation.</p></li><li><p>Disable custom skills.</p></li><li><p>Do not attach files.</p></li><li><p>Do not open a project workspace.</p></li><li><p>Send only:</p></li></ol><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Hi</code></pre><p>This will show the approximate minimum context usage.</p><hr><h2>Context Window vs Token Usage</h2><p>It is important to distinguish between:</p><h3>Context Window</h3><p>The maximum amount of information a model can process in one request.</p><p>Example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>1M token context window</code></pre><p>means the model can theoretically handle up to 1 million tokens.</p><h3>Token Usage</h3><p>The actual amount of tokens included in a specific request.</p><p>Example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>User message: Hi Actual request: System instructions + tools + history + "Hi" Total: 15,000 tokens</code></pre><p>A large context window does not mean every request will only consume the length of your message.</p><hr><h2>Conclusion</h2><p>High token usage from a short message does not necessarily indicate a problem with the model or API.</p><p>In AI agent environments, the total context includes many hidden components required for intelligent operation. A simple “Hi” can still result in thousands of tokens being processed if the session contains skills, tools, project information, or previous history.</p><p>We recommend starting a clean session without additional context when testing raw token usage.</p><p>Thank you for your understanding and continued support.</p>
  • 0 Votes
    1 Posts
    6 Views
    A
    <h2><strong>Introduction</strong></h2><p></p><p>Recently, some users noticed an unusual situation:</p><blockquote><p>“I only sent ‘Hi’, why did it consume 10,000+ tokens?”</p></blockquote><p>At first glance, this may seem unexpected. A two-character message should only require a few tokens, right?</p><p>The answer is: <strong>the model does not only process your visible message.</strong></p><p>Modern AI assistants, especially coding agents such as Codex, Claude Code, and other IDE-based AI tools, send much more information together with your message to help the model understand the task and provide better responses.</p><hr><p></p><h2>Your Message Is Only a Small Part of the Context</h2><p>When you send:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Hi</code></pre><p>the actual request sent to the AI model may look more like this:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>System Instructions + User Settings + Configured Skills + Available Tools + Project Information + File Context + Previous Conversation History + Your Message: "Hi"</code></pre><p>The total size of all these components is the actual context processed by the model.</p><p>Therefore, even a very short message can consume a significant number of tokens.</p><hr><h2>What Can Increase Context Usage?</h2><h3>1. Configured Skills</h3><p>If you have enabled custom skills or AI capabilities, their instructions need to be loaded into the model context.</p><p>For example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Skill A instructions: 2,000 tokens Skill B instructions: 3,000 tokens Skill C instructions: 1,500 tokens</code></pre><p>Even before you type anything, these instructions may already occupy several thousand tokens.</p><hr><h3>2. Tools and Agent Capabilities</h3><p>AI coding assistants usually provide tools such as:</p><ul><li><p>File reading</p></li><li><p>Code search</p></li><li><p>Terminal execution</p></li><li><p>Git operations</p></li><li><p>Browser access</p></li><li><p>External integrations</p></li></ul><p>The definitions and usage instructions of these tools also consume context space.</p><hr><h3>3. Project and File Context</h3><p>When working inside an IDE or code repository, the AI client may automatically include:</p><ul><li><p>Project structure</p></li><li><p>Open files</p></li><li><p>Code snippets</p></li><li><p>Dependencies</p></li><li><p>Configuration files</p></li></ul><p>This allows the AI to understand your project, but it also increases context usage.</p><hr><h3>4. Previous Conversation History</h3><p>If you continue an existing conversation, the previous messages are usually included again so the model understands the context.</p><p>For example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Previous discussion: 12,000 tokens New message: Hi</code></pre><p>The new request may still contain the previous 12,000 tokens.</p><hr><h2>How to Test Minimum Token Usage</h2><p>If you want to measure the token usage of a simple message, please try the following:</p><ol><li><p>Create a completely new conversation.</p></li><li><p>Disable custom skills.</p></li><li><p>Do not attach files.</p></li><li><p>Do not open a project workspace.</p></li><li><p>Send only:</p></li></ol><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Hi</code></pre><p>This will show the approximate minimum context usage.</p><hr><h2>Context Window vs Token Usage</h2><p>It is important to distinguish between:</p><h3>Context Window</h3><p>The maximum amount of information a model can process in one request.</p><p>Example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>1M token context window</code></pre><p>means the model can theoretically handle up to 1 million tokens.</p><h3>Token Usage</h3><p>The actual amount of tokens included in a specific request.</p><p>Example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>User message: Hi Actual request: System instructions + tools + history + "Hi" Total: 15,000 tokens</code></pre><p>A large context window does not mean every request will only consume the length of your message.</p><hr><h2>Conclusion</h2><p>High token usage from a short message does not necessarily indicate a problem with the model or API.</p><p>In AI agent environments, the total context includes many hidden components required for intelligent operation. A simple “Hi” can still result in thousands of tokens being processed if the session contains skills, tools, project information, or previous history.</p><p>We recommend starting a clean session without additional context when testing raw token usage.</p><p>Thank you for your understanding and continued support.</p>
  • 0 Votes
    1 Posts
    5 Views
    A
    <h2><strong>Panimula</strong></h2> <p></p> <p>Kamakailan, may ilang user na nakapansin ng isang hindi pangkaraniwang sitwasyon:</p> <blockquote> <p>“‘Hi’ lang ang ipinadala ko, bakit ito kumonsumo ng mahigit 10,000 token?”</p> <p></p> </blockquote> <p>Sa unang tingin, maaaring mukhang hindi ito inaasahan. Ang mensaheng may dalawang character ay dapat mangailangan lang ng kaunting token, tama ba?</p> <p>Ang sagot ay: <strong>hindi lang ang nakikita mong mensahe ang pinoproseso ng modelo.</strong></p> <p>Ang mga modernong AI assistant, lalo na ang mga coding agent gaya ng Codex, Claude Code, at iba pang IDE-based na AI tool, ay nagpapadala ng mas marami pang impormasyon kasabay ng iyong mensahe upang matulungan ang modelo na maunawaan ang gawain at makapagbigay ng mas mahusay na mga tugon.</p> <p></p> <hr> <p></p> <h2>Ang Iyong Mensahe ay Maliit na Bahagi Lang ng Konteksto</h2> <p>Kapag ipinadala mo ang:</p> <pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Hi</code></pre> <p></p> <p>ang aktuwal na request na ipinapadala sa AI model ay maaaring mas kahawig nito:</p> <pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>System Instructions + User Settings + Configured Skills + Available Tools + Project Information + File Context + Previous Conversation History + Your Message: "Hi"</code></pre> <p></p> <p>Ang kabuuang laki ng lahat ng bahaging ito ang tunay na kontekstong pinoproseso ng modelo.</p> <p>Dahil dito, kahit napakaikling mensahe ay maaaring kumonsumo ng malaking bilang ng token.</p> <hr> <p></p> <h2><strong>Ano ang Maaaring Magpalaki ng Context Usage?</strong></h2> <h3>1. Mga Naka-configure na Skill</h3> <p></p> <p>Kung naka-enable ang mga custom skill o AI capability, kailangang i-load ang kanilang mga instruction sa model context.</p> <p></p> <p>Halimbawa:</p> <pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Skill A instructions: 2,000 tokens Skill B instructions: 3,000 tokens Skill C instructions: 1,500 tokens</code></pre> <p>Kahit bago ka pa mag-type ng kahit ano, maaaring sumasakop na agad ang mga instruction na ito ng ilang libong token.</p> <hr> <p></p> <h3>2. Mga Tool at Kakayahan ng Agent</h3> <p></p> <p>Ang mga AI coding assistant ay karaniwang may mga tool gaya ng:</p> <ul> <li><p>Pagbasa ng file</p></li> <li><p>Paghahanap ng code</p></li> <li><p>Pagpapatakbo ng terminal</p></li> <li><p>Mga operasyon sa Git</p></li> <li><p>Pag-access sa browser</p></li> <li><p>Mga external integration</p></li> </ul> <p>Ang mga depinisyon at instruction sa paggamit ng mga tool na ito ay kumokonsumo rin ng context space.</p> <hr> <p></p> <h3>3. Konteksto ng Proyekto at mga File</h3> <p></p> <p>Kapag nagtatrabaho sa loob ng IDE o code repository, maaaring awtomatikong isama ng AI client ang:</p> <ul> <li><p>Istruktura ng proyekto</p></li> <li><p>Mga bukas na file</p></li> <li><p>Mga code snippet</p></li> <li><p>Mga dependency</p></li> <li><p>Mga configuration file</p></li> </ul> <p>Nakakatulong ito para maunawaan ng AI ang iyong proyekto, ngunit pinapataas din nito ang context usage.</p> <hr> <p></p> <h3>4. Nakaraang Kasaysayan ng Usapan</h3> <p></p> <p>Kung ipinagpapatuloy mo ang isang umiiral nang usapan, karaniwang isinasamang muli ang mga naunang mensahe upang maunawaan ng modelo ang konteksto.</p> <p>Halimbawa:</p> <pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Previous discussion: 12,000 tokens New message: Hi</code></pre> <p>Maaaring lamanin pa rin ng bagong request ang nakaraang 12,000 token.</p> <hr> <h2>Paano Subukan ang Pinakamababang Token Usage</h2> <p>Kung gusto mong sukatin ang token usage ng isang simpleng mensahe, subukan ang mga sumusunod:</p> <ol> <li><p>Gumawa ng ganap na bagong usapan.</p></li> <li><p>I-disable ang mga custom skill.</p></li> <li><p>Huwag mag-attach ng mga file.</p></li> <li><p>Huwag magbukas ng project workspace.</p></li> <li><p>Ipadala lamang ang:</p></li> </ol> <pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Hi</code></pre> <p>Ipapakita nito ang tinatayang pinakamababang context usage.</p> <hr> <h2>Context Window kumpara sa Token Usage</h2> <p>Mahalagang pag-ibahin ang mga sumusunod:</p> <h3>Context Window</h3> <p>Ang pinakamataas na dami ng impormasyong kayang iproseso ng isang modelo sa iisang request.</p> <p>Halimbawa:</p> <pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>1M token context window</code></pre> <p>na nangangahulugang kayang humawak ng modelo, sa teorya, ng hanggang 1 milyong token.</p> <h3>Token Usage</h3> <p>Ang aktuwal na dami ng token na kasama sa isang partikular na request.</p> <p>Halimbawa:</p> <pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>User message: Hi Actual request: System instructions + tools + history + "Hi" Total: 15,000 tokens</code></pre> <p>Ang malaking context window ay hindi nangangahulugang ang bawat request ay kakonsumo lamang ng habang katumbas ng iyong mensahe.</p> <p></p> <hr> <p></p> <h2><strong>Konklusyon</strong></h2> <p></p> <p>Ang mataas na token usage mula sa maikling mensahe ay hindi awtomatikong nangangahulugan na may problema sa modelo o API.</p> <p></p> <p>Sa mga AI agent environment, kasama sa kabuuang konteksto ang maraming nakatagong bahagi na kailangan para sa matalinong operasyon. Ang simpleng “Hi” ay maaari pa ring magresulta sa libo-libong token na napoproseso kung ang session ay may mga skill, tool, impormasyon ng proyekto, o nakaraang kasaysayan.</p> <p></p> <p>Iminumungkahi naming magsimula ng malinis na session na walang karagdagang konteksto kapag sinusubukan ang raw token usage.</p> <p></p> <p>Maraming salamat sa inyong pag-unawa at patuloy na suporta.</p>
  • 0 Votes
    1 Posts
    5 Views
    A
    <h2><strong>Εισαγωγή</strong></h2> <p></p> <p>Πρόσφατα, ορισμένοι χρήστες παρατήρησαν μια ασυνήθιστη κατάσταση:</p> <blockquote> <p>«Έστειλα μόνο ‘Hi’, γιατί κατανάλωσε 10.000+ tokens;»</p> <p></p> </blockquote> <p>Με την πρώτη ματιά, αυτό μπορεί να φαίνεται απροσδόκητο. Ένα μήνυμα δύο χαρακτήρων θα έπρεπε να απαιτεί μόνο λίγα tokens, σωστά;</p> <p>Η απάντηση είναι: <strong>το μοντέλο δεν επεξεργάζεται μόνο το ορατό μήνυμά σας.</strong></p> <p>Οι σύγχρονοι βοηθοί AI, ειδικά οι agents προγραμματισμού όπως το Codex, το Claude Code και άλλα εργαλεία AI που βασίζονται σε IDE, στέλνουν πολύ περισσότερες πληροφορίες μαζί με το μήνυμά σας, ώστε να βοηθήσουν το μοντέλο να κατανοήσει την εργασία και να παρέχει καλύτερες απαντήσεις.</p> <p></p> <hr> <p></p> <h2>Το μήνυμά σας είναι μόνο ένα μικρό μέρος του context</h2> <p>Όταν στέλνετε:</p> <pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Hi</code></pre> <p></p> <p>το πραγματικό αίτημα που αποστέλλεται στο μοντέλο AI μπορεί να μοιάζει περισσότερο με αυτό:</p> <pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Οδηγίες Συστήματος + Ρυθμίσεις Χρήστη + Διαμορφωμένες Δεξιότητες + Διαθέσιμα Εργαλεία + Πληροφορίες Έργου + Context Αρχείων + Ιστορικό Προηγούμενης Συνομιλίας + Το μήνυμά σας: "Hi"</code></pre> <p></p> <p>Το συνολικό μέγεθος όλων αυτών των στοιχείων είναι το πραγματικό context που επεξεργάζεται το μοντέλο.</p> <p>Επομένως, ακόμη και ένα πολύ σύντομο μήνυμα μπορεί να καταναλώσει σημαντικό αριθμό tokens.</p> <hr> <p></p> <h2><strong>Τι μπορεί να αυξήσει τη χρήση context;</strong></h2> <h3>1. Διαμορφωμένες Δεξιότητες</h3> <p></p> <p>Αν έχετε ενεργοποιήσει προσαρμοσμένες δεξιότητες ή δυνατότητες AI, οι οδηγίες τους πρέπει να φορτωθούν στο context του μοντέλου.</p> <p></p> <p>Για παράδειγμα:</p> <pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Οδηγίες Δεξιότητας A: 2.000 tokens Οδηγίες Δεξιότητας B: 3.000 tokens Οδηγίες Δεξιότητας 1.500 tokens</code></pre> <p>Ακόμη και πριν πληκτρολογήσετε οτιδήποτε, αυτές οι οδηγίες μπορεί ήδη να καταλαμβάνουν αρκετές χιλιάδες tokens.</p> <hr> <p></p> <h3>2. Εργαλεία και δυνατότητες agent</h3> <p></p> <p>Οι βοηθοί προγραμματισμού AI συνήθως παρέχουν εργαλεία όπως:</p> <ul> <li><p>Ανάγνωση αρχείων</p></li> <li><p>Αναζήτηση κώδικα</p></li> <li><p>Εκτέλεση εντολών στο τερματικό</p></li> <li><p>Λειτουργίες Git</p></li> <li><p>Πρόσβαση σε browser</p></li> <li><p>Εξωτερικές ενσωματώσεις</p></li> </ul> <p>Οι ορισμοί και οι οδηγίες χρήσης αυτών των εργαλείων καταναλώνουν επίσης χώρο στο context.</p> <hr> <p></p> <h3>3. Context έργου και αρχείων</h3> <p></p> <p>Όταν εργάζεστε μέσα σε ένα IDE ή αποθετήριο κώδικα, ο client AI μπορεί να συμπεριλάβει αυτόματα:</p> <ul> <li><p>Δομή έργου</p></li> <li><p>Ανοιχτά αρχεία</p></li> <li><p>Αποσπάσματα κώδικα</p></li> <li><p>Εξαρτήσεις</p></li> <li><p>Αρχεία ρυθμίσεων</p></li> </ul> <p>Αυτό επιτρέπει στο AI να κατανοεί το έργο σας, αλλά αυξάνει επίσης τη χρήση context.</p> <hr> <p></p> <h3>4. Ιστορικό προηγούμενης συνομιλίας</h3> <p></p> <p>Αν συνεχίζετε μια υπάρχουσα συνομιλία, τα προηγούμενα μηνύματα συνήθως συμπεριλαμβάνονται ξανά ώστε το μοντέλο να κατανοεί το context.</p> <p>Για παράδειγμα:</p> <pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Προηγούμενη συζήτηση: 12.000 tokens Νέο μήνυμα: Hi</code></pre> <p>Το νέο αίτημα μπορεί και πάλι να περιέχει τα προηγούμενα 12.000 tokens.</p> <hr> <h2>Πώς να δοκιμάσετε την ελάχιστη χρήση tokens</h2> <p>Αν θέλετε να μετρήσετε τη χρήση tokens ενός απλού μηνύματος, δοκιμάστε τα εξής:</p> <ol> <li><p>Δημιουργήστε μια εντελώς νέα συνομιλία.</p></li> <li><p>Απενεργοποιήστε τις προσαρμοσμένες δεξιότητες.</p></li> <li><p>Μην επισυνάψετε αρχεία.</p></li> <li><p>Μην ανοίξετε workspace έργου.</p></li> <li><p>Στείλτε μόνο:</p></li> </ol> <pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Hi</code></pre> <p>Αυτό θα δείξει την κατά προσέγγιση ελάχιστη χρήση context.</p> <hr> <h2>Context Window έναντι Χρήσης Tokens</h2> <p>Είναι σημαντικό να γίνεται διάκριση μεταξύ των εξής:</p> <h3>Context Window</h3> <p>Η μέγιστη ποσότητα πληροφοριών που μπορεί να επεξεργαστεί ένα μοντέλο σε ένα αίτημα.</p> <p>Παράδειγμα:</p> <pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Context window 1M tokens</code></pre> <p>σημαίνει ότι το μοντέλο μπορεί θεωρητικά να διαχειριστεί έως και 1 εκατομμύριο tokens.</p> <h3>Χρήση Tokens</h3> <p>Η πραγματική ποσότητα tokens που περιλαμβάνεται σε ένα συγκεκριμένο αίτημα.</p> <p>Παράδειγμα:</p> <pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Μήνυμα χρήστη: Hi Πραγματικό αίτημα: Οδηγίες συστήματος + εργαλεία + ιστορικό + "Hi" Σύνολο: 15.000 tokens</code></pre> <p>Ένα μεγάλο context window δεν σημαίνει ότι κάθε αίτημα θα καταναλώνει μόνο το μήκος του μηνύματός σας.</p> <p></p> <hr> <p></p> <h2><strong>Συμπέρασμα</strong></h2> <p></p> <p>Η υψηλή χρήση tokens από ένα σύντομο μήνυμα δεν υποδεικνύει απαραίτητα πρόβλημα με το μοντέλο ή το API.</p> <p></p> <p>Σε περιβάλλοντα AI agent, το συνολικό context περιλαμβάνει πολλά κρυφά στοιχεία που απαιτούνται για έξυπνη λειτουργία. Ένα απλό «Hi» μπορεί και πάλι να έχει ως αποτέλεσμα την επεξεργασία χιλιάδων tokens, αν η συνεδρία περιλαμβάνει δεξιότητες, εργαλεία, πληροφορίες έργου ή προηγούμενο ιστορικό.</p> <p></p> <p>Συνιστούμε να ξεκινάτε μια καθαρή συνεδρία χωρίς επιπλέον context όταν δοκιμάζετε την ακατέργαστη χρήση tokens.</p> <p></p> <p>Σας ευχαριστούμε για την κατανόηση και τη συνεχή υποστήριξή σας.</p>
  • 0 Votes
    1 Posts
    6 Views
    A
    <h2><strong>Introduction</strong></h2><p></p><p>Recently, some users noticed an unusual situation:</p><blockquote><p>“I only sent ‘Hi’, why did it consume 10,000+ tokens?”</p><p></p></blockquote><p>At first glance, this may seem unexpected. A two-character message should only require a few tokens, right?</p><p>The answer is: <strong>the model does not only process your visible message.</strong></p><p>Modern AI assistants, especially coding agents such as Codex, Claude Code, and other IDE-based AI tools, send much more information together with your message to help the model understand the task and provide better responses.</p><p></p><hr><p></p><h2>Your Message Is Only a Small Part of the Context</h2><p>When you send:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Hi</code></pre><p></p><p>the actual request sent to the AI model may look more like this:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>System Instructions + User Settings + Configured Skills + Available Tools + Project Information + File Context + Previous Conversation History + Your Message: "Hi"</code></pre><p></p><p>The total size of all these components is the actual context processed by the model.</p><p>Therefore, even a very short message can consume a significant number of tokens.</p><hr><p></p><h2><strong>What Can Increase Context Usage?</strong></h2><h3>1. Configured Skills</h3><p></p><p>If you have enabled custom skills or AI capabilities, their instructions need to be loaded into the model context.</p><p></p><p>For example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Skill A instructions: 2,000 tokens Skill B instructions: 3,000 tokens Skill C instructions: 1,500 tokens</code></pre><p>Even before you type anything, these instructions may already occupy several thousand tokens.</p><hr><p></p><h3>2. Tools and Agent Capabilities</h3><p></p><p>AI coding assistants usually provide tools such as:</p><ul><li><p>File reading</p></li><li><p>Code search</p></li><li><p>Terminal execution</p></li><li><p>Git operations</p></li><li><p>Browser access</p></li><li><p>External integrations</p></li></ul><p>The definitions and usage instructions of these tools also consume context space.</p><hr><p></p><h3>3. Project and File Context</h3><p></p><p>When working inside an IDE or code repository, the AI client may automatically include:</p><ul><li><p>Project structure</p></li><li><p>Open files</p></li><li><p>Code snippets</p></li><li><p>Dependencies</p></li><li><p>Configuration files</p></li></ul><p>This allows the AI to understand your project, but it also increases context usage.</p><hr><p></p><h3>4. Previous Conversation History</h3><p></p><p>If you continue an existing conversation, the previous messages are usually included again so the model understands the context.</p><p>For example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Previous discussion: 12,000 tokens New message: Hi</code></pre><p>The new request may still contain the previous 12,000 tokens.</p><hr><h2>How to Test Minimum Token Usage</h2><p>If you want to measure the token usage of a simple message, please try the following:</p><ol><li><p>Create a completely new conversation.</p></li><li><p>Disable custom skills.</p></li><li><p>Do not attach files.</p></li><li><p>Do not open a project workspace.</p></li><li><p>Send only:</p></li></ol><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>Hi</code></pre><p>This will show the approximate minimum context usage.</p><hr><h2>Context Window vs Token Usage</h2><p>It is important to distinguish between:</p><h3>Context Window</h3><p>The maximum amount of information a model can process in one request.</p><p>Example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>1M token context window</code></pre><p>means the model can theoretically handle up to 1 million tokens.</p><h3>Token Usage</h3><p>The actual amount of tokens included in a specific request.</p><p>Example:</p><pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"><code>User message: Hi Actual request: System instructions + tools + history + "Hi" Total: 15,000 tokens</code></pre><p>A large context window does not mean every request will only consume the length of your message.</p><p></p><hr><p></p><h2><strong>Conclusion</strong></h2><p></p><p>High token usage from a short message does not necessarily indicate a problem with the model or API.</p><p></p><p>In AI agent environments, the total context includes many hidden components required for intelligent operation. A simple “Hi” can still result in thousands of tokens being processed if the session contains skills, tools, project information, or previous history.</p><p></p><p>We recommend starting a clean session without additional context when testing raw token usage.</p><p></p><p>Thank you for your understanding and continued support.</p>