<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Explanation of Tiered Pricing for GPT-5.6 Sol Ultra-Long Contexts]]></title><description><![CDATA[<p dir="auto">&lt;p&gt;According to OpenAI's latest pricing rules, models such as GPT-5.6 Sol use tiered pricing for ultra-long-context requests.&lt;/p&gt;&lt;p&gt;When the context of a single request exceeds 272K tokens, input, cached reads, and output may be charged according to higher-tier prices. This rule comes from OpenAI's official pricing and is not a temporary markup or abnormal charge from AI-ROUTER.&lt;/p&gt;&lt;p&gt;&lt;/p&gt;&lt;p&gt;Using GPT-5.6 Sol's standard pricing as an example:&lt;/p&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;   Item                Within 272K          Over 272K&lt;/code&gt;&lt;/pre&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;  ━━━━━━━━━━  ━━━━━━━━━━━━━━━━━━━  ━━━━━━━━━━━━━━━━━&lt;/code&gt;&lt;/pre&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;   Input            $5 / 1M tokens    $10 / 1M tokens&lt;/code&gt;&lt;/pre&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;  ──────────  ───────────────────  ─────────────────&lt;/code&gt;&lt;/pre&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;   Cached read    $0.50 / 1M tokens     $1 / 1M tokens&lt;/code&gt;&lt;/pre&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;  ──────────  ───────────────────  ─────────────────&lt;/code&gt;&lt;/pre&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;   Output          $30 / 1M tokens    $45 / 1M tokens&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;&lt;/p&gt;&lt;p&gt;The final cost is still calculated based on the specific group, account multiplier, and other billing settings.&lt;/p&gt;&lt;p&gt;&lt;/p&gt;&lt;p&gt;An x2 flag in usage records indicates that the request triggered tiered pricing for long contexts. This flag primarily indicates that input and cached reads entered a higher tier; it does not mean that all charges for the request are simply doubled. Output may use a different tier multiplier.&lt;/p&gt;&lt;p&gt;&lt;/p&gt;&lt;p&gt;&lt;strong&gt; ## How to Avoid Triggering Tiered Pricing&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;If you do not need a 1M-token context window, we recommend adjusting Codex's context limit back to 272K or below.&lt;/p&gt;&lt;p&gt;&lt;/p&gt;&lt;p&gt;Open:&lt;/p&gt;&lt;p&gt;  ~/.codex/config.toml&lt;/p&gt;&lt;p&gt;&lt;/p&gt;&lt;p&gt;Change the 1M configuration from the original article:&lt;/p&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;  model = "gpt-5.6-sol"&lt;/code&gt;&lt;/pre&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;  model_context_window = 1000000&lt;/code&gt;&lt;/pre&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;  model_auto_compact_token_limit = 900000&lt;/code&gt;&lt;/pre&gt;&lt;p&gt; &lt;/p&gt;&lt;p&gt;to:&lt;/p&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;  model = "gpt-5.6-sol"&lt;/code&gt;&lt;/pre&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;  model_context_window = 272000&lt;/code&gt;&lt;/pre&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;  model_auto_compact_token_limit = 250000&lt;/code&gt;&lt;/pre&gt;&lt;p&gt; &lt;/p&gt;&lt;p&gt;Save the changes, restart Codex, and start a new session.&lt;/p&gt;&lt;p&gt;&lt;/p&gt;&lt;p&gt;If you only want the change to apply temporarily to a single session, use:&lt;/p&gt;&lt;pre class="rounded-lg bg-slate-950 px-4 py-3 text-slate-100"&gt;&lt;code&gt;codex -m gpt-5.6-sol -c model_context_window=272000 -c model_auto_compact_token_limit=250000&lt;/code&gt;&lt;/pre&gt;&lt;p&gt; We also recommend:&lt;/p&gt;&lt;p&gt;  - Enable automatic compaction promptly;&lt;/p&gt;&lt;p&gt;  - Split ultra-long tasks across multiple sessions;&lt;/p&gt;&lt;p&gt;  - Reduce the amount of tool output returned at once;&lt;/p&gt;&lt;p&gt;  - Regularly summarize and clean up older conversation content.&lt;/p&gt;&lt;p&gt;&lt;/p&gt;&lt;p&gt;If you genuinely need a 1M-token context window, you can continue using the original configuration, but usage beyond 272K will be charged according to OpenAI's higher-tier pricing.&lt;/p&gt;&lt;p&gt;&lt;/p&gt;&lt;p&gt;For the original configuration instructions, see: &lt;a target="_blank" rel="noopener noreferrer nofollow" href="<a href="https://ai-router.dev/blog/post-g-677aa3e043912838-how-to-enable-a-1m-token-context-window-in-codex-for-gpt-5-6-sol" rel="nofollow ugc">https://ai-router.dev/blog/post-g-677aa3e043912838-how-to-enable-a-1m-token-context-window-in-codex-for-gpt-5-6-sol</a>"&gt;How to Enable a 1M-Token Context Window in Codex for GPT-5.6 Sol &lt;/a&gt;&lt;/p&gt;&lt;p&gt;&lt;/p&gt;&lt;p&gt;Thank you for your understanding and support.&lt;/p&gt;</p>
]]></description><link>https://nodebb.ai-router.dev/topic/674/explanation-of-tiered-pricing-for-gpt-5.6-sol-ultra-long-contexts</link><generator>RSS for Node</generator><lastBuildDate>Fri, 09 Oct 2026 15:30:34 GMT</lastBuildDate><atom:link href="https://nodebb.ai-router.dev/topic/674.rss" rel="self" type="application/rss+xml"/><pubDate>Mon, 24 Aug 2026 20:14:55 GMT</pubDate><ttl>60</ttl></channel></rss>