On 30 July 2026, OpenAI announced a sharp recalibration of its model pricing. GPT-5.6 Luna, the fastest and most affordable tier in the GPT-5.6 family, dropped by 80% to $0.20 per million input tokens and $1.20 per million output tokens. GPT-5.6 Terra, the balanced mid-tier model, fell by 20% to $2 per million input tokens and $12 per million output tokens. A new Sol Fast mode was also introduced, offering up to 2.5x faster processing at twice the standard price for latency-sensitive applications.
The immediate reaction in developer communities focused on the competitive pressure from open-weight models and the ongoing price war between frontier AI providers. That framing is accurate but commercially incomplete. The more important question for marketing and GTM leaders is what an 80% reduction in the cost of a capable, tool-using model means for the economics of agentic workflows — and whether the cost-per-token metric is the right unit of measurement for those workflows at all.
The GPT-5.6 Three-Tier Architecture
GPT-5.6 is structured as a three-tier family designed to match model capability to task requirements. Sol is the premium tier, optimised for complex reasoning, multi-step planning and tasks where accuracy is the primary constraint. Terra is the balanced tier, suitable for everyday work tasks that require solid reasoning without the full cost of the premium model. Luna is the fast, lightweight tier, optimised for high-volume, latency-sensitive tasks such as classification, summarisation, structured extraction and tool calls within larger workflows.
The pricing structure before the 30 July cut positioned Luna at approximately $1.00 per million input tokens — already competitive with comparable models from other providers. At $0.20 per million input tokens, Luna is now priced at a level where the cost of running a high-volume agentic workflow is measured in dollars rather than tens of dollars per thousand interactions. For a marketing workflow that processes 10,000 content items per month — classifying intent, extracting entities, generating structured summaries — the monthly Luna cost at the new pricing is approximately $2 in input tokens plus output costs, depending on content length.
Multi-Model Routing: The Architecture That Makes Price Cuts Matter
The commercial value of the Luna price cut is not primarily in running Luna for everything. It is in enabling a multi-model routing architecture where the right model is used for each step of a workflow based on the complexity and accuracy requirements of that step.
A practical agentic marketing workflow might use Sol for strategic planning and campaign brief generation, where reasoning quality directly affects downstream output quality. It would use Terra for content drafting, audience analysis and competitive research, where a balance of quality and cost is appropriate. And it would use Luna for high-volume execution tasks: classifying inbound leads, extracting structured data from documents, generating metadata, routing content to the correct workflow branch, and validating outputs against defined criteria.
In this architecture, the 80% Luna price cut reduces the cost of the high-volume execution layer — the part of the workflow that runs most frequently and at the highest token volume. The planning and reasoning layer, which runs less frequently and at lower volume, remains at Sol pricing. The net effect on total workflow cost depends on the ratio of execution to planning tokens, but for most production marketing workflows, execution tokens significantly outnumber planning tokens.
The Metric That Actually Matters: Cost Per Accepted Outcome
The most common mistake in AI cost modelling for marketing workflows is treating cost per token as the primary metric. Token cost is a useful input, but it is not the commercially relevant unit. The relevant unit is cost per accepted commercial outcome — the cost of producing an output that meets the quality threshold required for the business to use it without additional human review or rework.
A workflow that produces 1,000 content summaries at $0.50 per summary but requires human review and correction for 40% of outputs has an effective cost of $0.50 plus the labour cost of reviewing and correcting 400 items. A workflow that produces the same 1,000 summaries at $1.20 per summary but requires correction for only 5% of outputs has a lower total cost despite the higher token price. The quality threshold — the acceptance rate — is the variable that determines whether a cheaper model actually reduces total workflow cost.
For Luna specifically, the acceptance rate question is: at what task types and complexity levels does Luna produce outputs that meet the quality threshold for direct use? The answer varies by task. For binary classification, entity extraction from structured documents and format conversion, Luna's acceptance rate is typically high. For nuanced content generation, complex reasoning chains and tasks requiring contextual judgment, the acceptance rate drops and the cost advantage over Terra or Sol narrows or reverses when rework costs are included.
Practical Implications for B2B Marketing Workflows
The Luna price cut makes several previously marginal use cases economically viable. Real-time lead scoring and routing — classifying inbound leads against ICP criteria and routing them to the appropriate sales sequence — can now run at Luna pricing with a cost per lead classification well below $0.01. Content metadata generation — extracting topics, entities, intent signals and structured tags from large content libraries — becomes cost-effective at scale. Quality assurance passes — checking outputs from other models or human writers against defined criteria — can run continuously rather than on a sample basis.
For B2B businesses evaluating their first agentic workflow deployment, the price reduction lowers the barrier to experimentation. A pilot that processes 100,000 items per month to validate acceptance rates and workflow design now costs less than $50 in Luna tokens, making the cost of learning negligible relative to the cost of the workflow design and integration work. The constraint on agentic marketing adoption is no longer primarily token cost — it is workflow design, integration complexity and the organisational readiness to act on AI-generated outputs.
Key Takeaways
- OpenAI cut GPT-5.6 Luna pricing by 80% on 30 July 2026 to $0.20 per million input tokens and $1.20 per million output tokens; Terra fell 20% to $2/$12 per million tokens.
- The price cut enables multi-model routing architectures where Sol handles planning, Terra handles drafting and Luna handles high-volume execution — reducing total workflow cost significantly.
- Cost per token is the wrong primary metric; cost per accepted commercial outcome — which accounts for acceptance rates and rework costs — is the commercially relevant unit.
- Luna is well-suited to binary classification, entity extraction, metadata generation and quality assurance passes; it is less suited to nuanced content generation and complex reasoning chains.
- The constraint on agentic marketing adoption is no longer primarily token cost — it is workflow design, integration complexity and organisational readiness to act on AI outputs.
On 30 July 2026, OpenAI announced a sharp recalibration of its model pricing. GPT-5.6 Luna, the fastest and most affordable tier in the GPT-5.6 family, dropped by 80% to $0.20 per million input tokens and $1.20 per million output tokens. GPT-5.6 Terra, the balanced mid-tier model, fell by 20% to $2 per million input tokens and $12 per million output tokens. A new Sol Fast mode was also introduced, offering up to 2.5x faster processing at twice the standard price for latency-sensitive applications.
The immediate reaction in developer communities focused on the competitive pressure from open-weight models and the ongoing price war between frontier AI providers. That framing is accurate but commercially incomplete. The more important question for marketing and GTM leaders is what an 80% reduction in the cost of a capable, tool-using model means for the economics of agentic workflows — and whether the cost-per-token metric is the right unit of measurement for those workflows at all.
The GPT-5.6 Three-Tier Architecture
GPT-5.6 is structured as a three-tier family designed to match model capability to task requirements. Sol is the premium tier, optimised for complex reasoning, multi-step planning and tasks where accuracy is the primary constraint. Terra is the balanced tier, suitable for everyday work tasks that require solid reasoning without the full cost of the premium model. Luna is the fast, lightweight tier, optimised for high-volume, latency-sensitive tasks such as classification, summarisation, structured extraction and tool calls within larger workflows.
The pricing structure before the 30 July cut positioned Luna at approximately $1.00 per million input tokens — already competitive with comparable models from other providers. At $0.20 per million input tokens, Luna is now priced at a level where the cost of running a high-volume agentic workflow is measured in dollars rather than tens of dollars per thousand interactions. For a marketing workflow that processes 10,000 content items per month — classifying intent, extracting entities, generating structured summaries — the monthly Luna cost at the new pricing is approximately $2 in input tokens plus output costs, depending on content length.
Multi-Model Routing: The Architecture That Makes Price Cuts Matter
The commercial value of the Luna price cut is not primarily in running Luna for everything. It is in enabling a multi-model routing architecture where the right model is used for each step of a workflow based on the complexity and accuracy requirements of that step.
A practical agentic marketing workflow might use Sol for strategic planning and campaign brief generation, where reasoning quality directly affects downstream output quality. It would use Terra for content drafting, audience analysis and competitive research, where a balance of quality and cost is appropriate. And it would use Luna for high-volume execution tasks: classifying inbound leads, extracting structured data from documents, generating metadata, routing content to the correct workflow branch, and validating outputs against defined criteria.
In this architecture, the 80% Luna price cut reduces the cost of the high-volume execution layer — the part of the workflow that runs most frequently and at the highest token volume. The planning and reasoning layer, which runs less frequently and at lower volume, remains at Sol pricing. The net effect on total workflow cost depends on the ratio of execution to planning tokens, but for most production marketing workflows, execution tokens significantly outnumber planning tokens.
The Metric That Actually Matters: Cost Per Accepted Outcome
The most common mistake in AI cost modelling for marketing workflows is treating cost per token as the primary metric. Token cost is a useful input, but it is not the commercially relevant unit. The relevant unit is cost per accepted commercial outcome — the cost of producing an output that meets the quality threshold required for the business to use it without additional human review or rework.
A workflow that produces 1,000 content summaries at $0.50 per summary but requires human review and correction for 40% of outputs has an effective cost of $0.50 plus the labour cost of reviewing and correcting 400 items. A workflow that produces the same 1,000 summaries at $1.20 per summary but requires correction for only 5% of outputs has a lower total cost despite the higher token price. The quality threshold — the acceptance rate — is the variable that determines whether a cheaper model actually reduces total workflow cost.
For Luna specifically, the acceptance rate question is: at what task types and complexity levels does Luna produce outputs that meet the quality threshold for direct use? The answer varies by task. For binary classification, entity extraction from structured documents and format conversion, Luna's acceptance rate is typically high. For nuanced content generation, complex reasoning chains and tasks requiring contextual judgment, the acceptance rate drops and the cost advantage over Terra or Sol narrows or reverses when rework costs are included.
Practical Implications for B2B Marketing Workflows
The Luna price cut makes several previously marginal use cases economically viable. Real-time lead scoring and routing — classifying inbound leads against ICP criteria and routing them to the appropriate sales sequence — can now run at Luna pricing with a cost per lead classification well below $0.01. Content metadata generation — extracting topics, entities, intent signals and structured tags from large content libraries — becomes cost-effective at scale. Quality assurance passes — checking outputs from other models or human writers against defined criteria — can run continuously rather than on a sample basis.
For B2B businesses evaluating their first agentic workflow deployment, the price reduction lowers the barrier to experimentation. A pilot that processes 100,000 items per month to validate acceptance rates and workflow design now costs less than $50 in Luna tokens, making the cost of learning negligible relative to the cost of the workflow design and integration work. The constraint on agentic marketing adoption is no longer primarily token cost — it is workflow design, integration complexity and the organisational readiness to act on AI-generated outputs.
Key Takeaways
- OpenAI cut GPT-5.6 Luna pricing by 80% on 30 July 2026 to $0.20 per million input tokens and $1.20 per million output tokens; Terra fell 20% to $2/$12 per million tokens.
- The price cut enables multi-model routing architectures where Sol handles planning, Terra handles drafting and Luna handles high-volume execution — reducing total workflow cost significantly.
- Cost per token is the wrong primary metric; cost per accepted commercial outcome — which accounts for acceptance rates and rework costs — is the commercially relevant unit.
- Luna is well-suited to binary classification, entity extraction, metadata generation and quality assurance passes; it is less suited to nuanced content generation and complex reasoning chains.
- The constraint on agentic marketing adoption is no longer primarily token cost — it is workflow design, integration complexity and organisational readiness to act on AI outputs.






