· AI

OpenAI Says GPT-6 Sol's Half-Price Rate Is Permanent. That Word Matters More Than the Discount.

A durable price floor is more useful to a business than a deep temporary discount, and OpenAI just offered one, so plan around it while keeping the exit door open.

OpenAI released GPT-6 Sol and GPT-6 Luna on September 22 at $2 and $10 per million input and output tokens for Sol, and $0.10 and $0.50 for Luna, roughly half of what the GPT-5.6 versions cost that morning (VentureBeat). The company credited better inference and caching for the cut. Hours earlier, Anthropic shipped Claude Opus 5.5 at $4 and $20, which it says is 20 percent cheaper per token than Opus 5 and about 40 percent cheaper to run on typical work because it answers with fewer tokens.

Everyone led with the 50 percent. I think that is the least reliable part of the story. The part worth budgeting on is one sentence lower down: an OpenAI spokesperson told VentureBeat these rates are permanent, not promotional or introductory.

Why the discount is fuzzier than it reads

Start with the baseline. GPT-5.6 Sol's $4 and $20 rate was itself a cut. OpenAI dropped it from $5 and $30 on August 21 and described it as available at least through November 21, according to TokenCost's pricing tracker. Measured against that original list price, Sol's new rate is 60 percent off input and 67 percent off output, bigger than the headline. But TokenCost also notes that OpenRouter was already reselling GPT-5.6 Sol at $2 and $10. For anyone buying through that channel, the per-token price did not move at all. They got a different model at the same rate.

Then the benchmarks. Every comparison in the launch was run by OpenAI, and the rival it chose was Opus 5, not the Opus 5.5 that shipped the same morning. VentureBeat noted there is no public head-to-head test of the two under identical settings yet. BuildFastWithAI's roundup goes further, reporting that Sol's best scores on DeepSWE and OSWorld 2.0 (68.8 and 64.4) come in below GPT-5.6 Sol's (72.7 and 66.2). That is one outlet's tally and I could not confirm it against OpenAI's own page, so hold it loosely. If it holds, part of the saving is a step down in capability, and for some workloads that trade is a bad one.

The case that "permanent" means nothing

The strongest skeptical reading goes like this. "Permanent" came from an unnamed spokesperson in one interview, not a contract. Price commitments in this market get walked back: TokenCost documents Google doubling Gemini 3.6 and 3.7 Flash to $1.50 and $7.50 on January 1, 2027. And if frontier labs are pricing below cost to win developer share, which is what I suspected in Tuesday's post on chip prices, a spokesperson's promise lasts exactly as long as the investors funding the loss.

That case is fair, and I hold part of it. It underrates two things.

First, OpenAI has put its credibility behind a specific claim that is easy to check. Raising Sol's price in six months would hand every competitor a ready-made story about bait and switch, at a time when Anthropic has cancelled its own planned Sonnet 5 increase (per TokenCost) and Xiaomi sells the open-weight MiMo-V2.6-Pro at $0.435 and $0.87 (per VentureBeat). The market is doing the enforcing. No regulator had to demand price stability. Competitors made reversing course expensive.

Second, the efficiency story has evidence behind it. Storyboard18 reports a 90 percent discount on cached input, and Anthropic's own pitch for Opus 5.5 is about spending fewer tokens per answer, not just charging less per token. Two labs pointing to specific engineering gains on the same day reads less like a coordinated subsidy and more like the work actually getting cheaper to serve, though neither lab publishes the margins that would prove it. I was partly wrong on Tuesday about how much of the drop is subsidy. Some of it, I still suspect, but less than I implied.

What to do with it

Here is the math on a modest workload of 200 million input tokens and 40 million output tokens a month. On GPT-5.6 Sol at $4 and $20, that is $1,600. On GPT-6 Sol, $800. If half your input is a repeated system prompt that hits the cache, input drops from $400 to about $220, so the bill lands near $620.

That saving is real, and a permanent rate means you can write it into an annual plan. Just do not buy it blind:

  1. Run your own evaluation on your actual tasks before you migrate. A cheaper model that scores lower on your work is not a discount.
  2. Structure prompts so the stable part repeats and caches. The 90 percent cached rate is worth more than the headline cut on most chat and agent workloads.
  3. Keep the model swappable behind one internal interface. If "permanent" turns out to mean eighteen months, switching should cost you a config change, not a rewrite.

Picking which model handles which work, and on what pricing assumptions, is a large share of what we do in MojoAI advisory. If you are about to move production traffic to GPT-6 Sol and have not tested it on your own tasks yet, let's talk first.

Sources

References used in this article. Links also appear alongside the relevant claims.

Let's talk

Tell us what's on your mind.

You don't need a polished brief to reach out. A two-line email about what's bugging you is plenty; we'll tell you straight if we're the right fit, and what we'd tackle first.

We'll scope the work around your workflow, goals, and timeline before quoting anything, so you know what's included before committing.

LocationBoca Raton, Florida
CoverageSouth Florida + remote nationwide
Status Now accepting clients