GPT-5.6 Luna API Pricing, Context Window & Benchmarks

Verify GPT-5.6 Luna API pricing, model IDs, context and output limits, availability, benchmarks, strengths and ideal workloads.
GPT-5.6 Luna API Pricing, Context Window & Benchmarks
What if a frontier-model API could process 1 million input tokens for just $0.20? That headline makes GPT-5.6 Luna API pricing unusually compelling for high-volume applications—but it also sits amid confusing claims about a “GPT-6 Luna” launch, conflicting context-window figures, and search snippets that mix Luna’s prices with those of more expensive models.
Here is the verified picture as of September 2026: OpenAI’s current API documentation identifies the public model as GPT-5.6 Luna, with the model ID gpt-5.6-luna, and describes it as designed for “cost-sensitive, high-volume workloads.” OpenAI lists GPT-5.6 Luna at $0.20 per million input tokens, $0.02 per million cached input tokens, and $1.20 per million output tokens as of September 2026. OpenAI’s GPT-5.6 announcement also states that Luna’s price was reduced by 80%, positioning it as the GPT-5.6 family’s “most cost-efficient” and “fastest” option.
The reported GPT-6 Luna launch requires an important qualification. Although the launch has been confirmed editorially, the primary OpenAI materials available for verification currently associate the Luna name with GPT-5.6 and the Astra name with GPT-6. A distinct GPT-6 Luna model page, API identifier, price sheet, or release announcement has not been independently verified from the cited primary sources. Developers should therefore avoid assuming that gpt-6-luna is a valid production model ID until OpenAI documents it explicitly.
That distinction matters because a naming mistake can break deployments, while misunderstanding token limits or cached-input pricing can distort cost forecasts by thousands of dollars at scale. OpenAI’s builder guide reports a GPT-5.6 Luna “Extra High” score of 84.04% at a cost of $1.33, but the benchmark name and test conditions must be checked before treating that figure as a universal performance measure. Likewise, widely repeated 1.05-million- and 1.1-million-token context-window claims should remain provisional until the exact shared input-output limits are confirmed in current model documentation.
This guide separates confirmed specifications from unresolved claims, covering availability, API use, model IDs, context limits, pricing, benchmarks, strengths, limitations, and ideal workloads. For developers comparing providers, CallMissed’s OpenAI-compatible gateway offers access to 136 models through one API key and balance, making it possible to evaluate models without rebuilding each integration.
Is GPT-6 Luna real? OpenAI currently lists GPT-5.6 Luna

Yes—GPT-6 Luna’s launch existence is supported by an OpenAI-domain product result dated September 22, 2026. The surfaced result is titled “Introducing GPT-6 Sol and Luna,” which supersedes the earlier conclusion that Luna was documented only as GPT-5.6.
However, the available search did not return a corresponding OpenAI developer model page, API reference or pricing entry for GPT-6 Luna. The announcement evidence therefore confirms the product name, but not its API specifications.
What does OpenAI officially call Luna?
OpenAI now appears to use Luna in two distinct model generations:
- GPT-5.6 Luna, the previously documented cost-efficient model with the API identifier
gpt-5.6-luna. - GPT-6 Luna, named alongside GPT-6 Sol in an OpenAI-domain product snippet dated September 22, 2026.
These models must not be treated as interchangeable. In particular, GPT-5.6 Luna’s model ID, pricing, context limits and benchmark results do not establish the corresponding details for GPT-6 Luna.
What is verified about GPT-6 Luna?
| Claim | Status as of September 22, 2026 | Evidence |
|---|---|---|
| OpenAI has introduced a product named GPT-6 Luna | Supported | OpenAI-domain result titled “Introducing GPT-6 Sol and Luna” |
| GPT-6 Luna is distinct from GPT-5.6 Luna | Supported by the naming | The new result identifies GPT-6, while existing documentation separately names GPT-5.6 Luna |
Public API model ID is gpt-6-luna | Unverified | No matching developer model page or API reference was returned |
| GPT-6 Luna API pricing | Unverified | No GPT-6 Luna pricing entry was available in the surfaced primary documentation |
| GPT-6 Luna context window | Unverified | No matching technical model page was returned |
| GPT-6 Luna benchmark results | Unverified | No GPT-6 Luna-specific benchmark documentation was located |
| General API availability or rollout tiers | Unverified | The product snippet alone does not establish endpoint access or account eligibility |
How does GPT-5.6 Luna differ?
The existing GPT-5.6 Luna documentation remains evidence about GPT-5.6 Luna only. It identifies the production model as gpt-5.6-luna and lists pricing of $0.20 per million input tokens, $0.02 per million cached input tokens and $1.20 per million output tokens.
Likewise, OpenAI’s reported 80% price reduction and the builder-guide result of 84.04% at a cost of $1.33 at “Extra High” belong to GPT-5.6 Luna. They should not be quoted as GPT-6 Luna prices or benchmarks without new GPT-6-specific documentation.
Should developers use gpt-6-luna in API requests?
Not solely on the basis of the launch snippet. A product announcement does not prove that a predictable API identifier is active. Until OpenAI exposes GPT-6 Luna in its developer catalogue or API documentation, developers should:
- Check the live provider model list before sending requests.
- Use
gpt-5.6-lunaonly when intentionally selecting GPT-5.6 Luna. - Avoid assuming that
gpt-6-lunais the production identifier. - Validate pricing, context limits and supported features against a GPT-6 Luna-specific primary source.
- Run a small test request before changing production routing.
The accurate conclusion as of September 22, 2026 is that GPT-6 Luna’s launch is supported by an official OpenAI-domain product snippet, while its API model ID, pricing, context window, benchmarks and developer availability remain unverified from the available primary documentation.
How do GPT-5.6 Luna, Sol, Terra and GPT-6 Astra fit together?

GPT-5.6 Luna, Sol and Terra are three members of OpenAI’s GPT-5.6 family, while GPT-6 Astra belongs to the newer GPT-6 generation. As of September 22, 2026, OpenAI’s primary documentation does not show GPT-6 Luna as a separate fourth tier; it identifies Luna as gpt-5.6-luna and Astra as the model powering GPT-6 Pro.
What is the GPT-5.6 Luna, Sol and Terra hierarchy?
The GPT-5.6 lineup is best understood as a cost, speed and capability spectrum, not as three successive releases:
- GPT-5.6 Luna targets cost-sensitive, high-volume processing. OpenAI calls Luna the GPT-5.6 family’s “most cost-efficient” and “fastest” model.
- GPT-5.6 Sol occupies a more expensive tier for workloads where stronger model capability can justify higher token costs.
- GPT-5.6 Terra is another GPT-5.6 tier, but the supplied primary-source excerpts do not establish its current model ID, exact pricing or intended workload clearly enough to state them as verified facts.
- GPT-6 Astra sits outside that family as a GPT-6 model and powers the GPT-6 Pro experience described by OpenAI.
This distinction prevents a common implementation error: treating Luna, Sol, Terra and Astra as interchangeable aliases or assuming that gpt-6-luna follows automatically from gpt-5.6-luna.
How different are Luna and Sol API costs?
GPT-5.6 Luna has a substantially lower per-token price than GPT-5.6 Sol. According to OpenAI’s model documentation dated September 2026:
- GPT-5.6 Luna: $0.20 per million input tokens and $1.20 per million output tokens.
- GPT-5.6 Sol: $4 per million input tokens and $20 per million output tokens.
- Luna’s listed input price is therefore 20 times lower than Sol’s, while its output price is approximately 16.7 times lower.
- OpenAI says Sol’s current pricing represents a 20% input-price reduction and a 33% output-price reduction.
- OpenAI’s GPT-5.6 announcement says Luna received an 80% price reduction, while Terra received a 20% reduction.
These figures make Luna the natural starting point for classification, extraction, routing, summarisation and other high-throughput jobs. Sol may be more economical when better results reduce retries, human review or multi-step prompting, but that trade-off must be tested on the application’s own evaluations.
Where does GPT-6 Astra fit?
GPT-6 Astra is the higher-generation option, not a renamed GPT-5.6 Luna tier. OpenAI’s Help Center states that GPT-6 Pro, powered by GPT-6 Astra, is available in ChatGPT on the Pro $100, Pro $200, Business and Enterprise plans as of September 2026; Enterprise access can also depend on organisational settings.
Developers should keep two access layers separate:
- ChatGPT availability determines which users can select GPT-6 Pro in the ChatGPT interface.
- API availability requires a documented API model ID, endpoint eligibility, rate limits and an unambiguous price row.
- GPT-6 Luna availability remains unverified until OpenAI publishes those details under that specific name.
The practical selection order is therefore: choose Luna for economical scale, evaluate Sol or Terra where additional capability warrants greater cost, and assess Astra for GPT-6-class workloads—without presuming that benchmark results, context limits or API identifiers transfer between tiers.
What launch dates, specifications and availability are confirmed?

The editorially confirmed GPT-6 Luna launch does not yet have a matching public OpenAI specification page, model ID, price sheet, or independently verifiable release date. As of September 22, 2026, the production-ready Luna model documented by OpenAI is GPT-5.6 Luna, available through the OpenAI Responses API as gpt-5.6-luna.
Which GPT-6 Luna launch details are confirmed?
The table separates details verified in current OpenAI primary sources from claims that remain editorially confirmed or unresolved.
| Item | Confirmed detail | Availability or value | Verification status |
|---|---|---|---|
| GPT-6 Luna launch | The launch is confirmed by the site administrator | Exact public launch date not established | Editorially confirmed; not independently verified |
| Public Luna model | GPT-5.6 Luna | OpenAI API documentation was live as of September 22, 2026 | Verified by OpenAI |
| Production model ID | gpt-5.6-luna | Listed for API use | Verified by OpenAI |
| API access | OpenAI Responses API | Intended for developer workloads | Verified by OpenAI Models documentation |
| Token pricing | $0.20 input, $0.02 cached input, $1.20 output per million tokens | Standard GPT-5.6 Luna API rates | Verified by OpenAI Pricing |
| Context window | Public claims cite approximately 1.05 million or 1.1 million tokens | Exact shared input-output limit remains unresolved | Not independently verified |
OpenAI describes GPT-5.6 Luna as designed for “cost-sensitive, high-volume workloads.” OpenAI’s GPT-5.6 announcement also calls Luna the family’s “most cost-efficient” and “fastest” model and reports that its price was reduced by 80%.
What specifications can developers safely use?
For implementation and budgeting, developers can currently rely on these documented values as of September 2026:
- Model name: GPT-5.6 Luna
- API identifier:
gpt-5.6-luna - Input price: $0.20 per 1 million tokens
- Cached-input price: $0.02 per 1 million tokens
- Output price: $1.20 per 1 million tokens
- Documented interface: OpenAI Responses API
- Primary positioning: fast, economical processing for high-volume applications
At those rates, processing 100 million uncached input tokens would cost $20, excluding output tokens and other platform charges. The same quantity served entirely as cached input would cost $2, illustrating why prompt caching can materially change production economics.
Which launch claims remain unresolved?
Three details should not yet be treated as production facts:
gpt-6-lunais not a verified model ID. Sending it in an API request could produce an invalid-model error unless OpenAI subsequently documents the identifier.- The exact GPT-6 Luna release date is unavailable in the cited primary materials. The administrator’s launch confirmation establishes the editorial status, not a public OpenAI publication date.
- The context limit is unsettled. Neither the widely repeated 1.05-million-token figure nor the rounded 1.1-million-token claim should be used for architecture decisions without checking the live model schema.
OpenAI’s public naming currently places Luna under GPT-5.6 and Astra under GPT-6. OpenAI’s Help Center states that GPT-6 Pro is powered by GPT-6 Astra and is available in ChatGPT on Pro $100, Pro $200, Business, and Enterprise plans as of September 2026; that statement does not independently establish GPT-6 Luna availability in ChatGPT or the API.
How much does the GPT-5.6 Luna API cost?

OpenAI charges $0.20 per million fresh input tokens, $0.02 per million cached input tokens, and $1.20 per million output tokens for gpt-5.6-luna as of September 2026. No independently verified price exists for a separate gpt-6-luna API model, so developers should not use GPT-5.6 Luna’s rates as confirmed GPT-6 pricing.
What are the confirmed GPT-5.6 Luna token prices?
OpenAI’s GPT-5.6 Luna model page describes Luna as intended for “cost-sensitive, high-volume workloads.” OpenAI also says it reduced GPT-5.6 Luna’s price by 80%, although that announcement should not be interpreted as confirming prices for an undocumented GPT-6 Luna model.
| Usage scenario | Input cost | Output cost | Estimated total |
|---|---|---|---|
| 1M fresh input tokens | $0.20 | $0.00 | $0.20 |
| 1M cached input tokens | $0.02 | $0.00 | $0.02 |
| 1M output tokens | $0.00 | $1.20 | $1.20 |
| 1M fresh input + 1M output | $0.20 | $1.20 | $1.40 |
| 10M fresh input + 1M output | $2.00 | $1.20 | $3.20 |
| 100M fresh input + 10M output | $20.00 | $12.00 | $32.00 |
Calculations use OpenAI’s standard gpt-5.6-luna API rates published as of September 2026; they exclude unrelated platform, storage, tool, search, or infrastructure charges.
The pricing formula is:
Total cost = (fresh input tokens ÷ 1,000,000 × $0.20) + (cached input tokens ÷ 1,000,000 × $0.02) + (output tokens ÷ 1,000,000 × $1.20).
How much cheaper is cached input?
GPT-5.6 Luna cached input costs 90% less than fresh input, according to OpenAI’s September 2026 pricing. Processing one billion fully cache-eligible input tokens would therefore cost $20, compared with $200 for one billion fresh input tokens.
Caching can materially reduce costs when applications repeatedly send the same stable content, such as:
- Long system instructions
- Shared policy or product documentation
- Reused codebase context
- Fixed schemas and tool definitions
- Common few-shot examples
Developers should not assume every prompt token receives the cached rate. Only input recognized as cache-eligible is billed at $0.02 per million tokens; new or changed prompt content remains fresh input.
How does Luna pricing compare with GPT-5.6 Sol?
OpenAI lists GPT-5.6 Sol at $4 per million input tokens and $20 per million output tokens as of September 2026. At published standard rates, Luna’s $0.20 input price is 95% lower than Sol’s, while Luna’s $1.20 output price is 94% lower than Sol’s.
That gap makes Luna financially attractive for classification, extraction, summarization, routing, and other high-volume tasks where its quality is sufficient. The trade-off should still be tested against accuracy and reliability requirements rather than decided on token price alone.
For multi-model evaluations, CallMissed’s OpenAI-compatible developer API provides one key and one balance for 136 models, allowing teams to compare model economics without rewriting an existing OpenAI SDK integration. Regardless of gateway, production budgets should include output length, cache-hit rates, retries, and fallback-model usage—not just the headline input price.
Is the context window 1.05M or 1.1M tokens?

Neither 1.05 million nor 1.1 million tokens is independently confirmed as the exact GPT-5.6 Luna context window in the cited OpenAI materials. As of September 22, 2026, treat both figures as provisional descriptions—not production-safe limits—until OpenAI publishes an explicit token count for gpt-5.6-luna.
Why do sources report both 1.05M and 1.1M tokens?
The discrepancy may be a matter of rounding rather than two different model versions. For example, an underlying limit of 1,050,000 tokens could be displayed as:
- 1.05M tokens when reported to two decimal places.
- Approximately 1.1M tokens when rounded to one decimal place.
- “Over one million tokens” in nontechnical launch material.
That explanation is plausible, but it is not proof that 1,050,000 is Luna’s actual limit. A model snapshot change, documentation error, or confusion between an advertised input capacity and a shared input-output budget could produce the same conflicting figures.
OpenAI’s GPT-5.6 Luna model page confirms that Luna targets “cost-sensitive, high-volume workloads” and identifies the production API model as gpt-5.6-luna as of September 2026. However, the cited model-page and pricing-page extracts do not state an exact context-window limit, so neither 1.05M nor 1.1M can be attributed conclusively to OpenAI from the available evidence.
Does the context window include output tokens?
Developers should assume the advertised context figure may represent a shared token budget unless OpenAI explicitly separates maximum input and maximum output limits. In a shared window, prompt tokens, retrieved documents, tool results, conversation history, reasoning-related tokens where applicable, and generated output all compete for capacity.
A simplified planning formula is:
usable input ≤ total context limit − reserved output − system and tool overhead
If a workload sends 1,000,000 input tokens and reserves 50,000 tokens for generation, it already requires at least 1,050,000 tokens of combined capacity before allowing for system instructions or tool schemas. That request might fit a genuine 1.1M shared window but fail against a 1.05M ceiling once overhead is counted.
Which limit should developers use in production?
Until OpenAI documents the exact number, use a deliberately conservative operating limit:
- Do not hard-code 1.1M based on rounded launch coverage.
- Keep prompts materially below 1.05M, with room for output and orchestration overhead.
- Count tokens using the tokenizer supported for
gpt-5.6-luna. - Test against the production API and record the model snapshot, request size, and error response.
- Recheck OpenAI’s model documentation before raising application limits.
API acceptance tests can reveal what works at a particular moment, but they do not replace a documented guarantee; routing changes and model snapshots can alter observed behavior.
What is the verified conclusion?
The defensible specification is currently “approximately one million tokens, exact limit unverified.” Use 1.05M only as a provisional lower claim and 1.1M only as an approximate upper description—not as an assured request budget. The exact maximum input, maximum output, and shared-context rules remain unresolved pending explicit OpenAI documentation for GPT-5.6 Luna.
What do the GPT-5.6 Luna benchmarks actually show?

The verified benchmark evidence supports a narrow conclusion: GPT-5.6 Luna can deliver strong task performance at unusually low API cost, but the published 84.04% result does not establish universal model quality. No independently verified benchmark suite or score is currently available for a distinct GPT-6 Luna model as of September 22, 2026.
What does the 84.04% GPT-5.6 Luna score mean?
OpenAI’s Builder’s Guide to GPT-5.6 reports that GPT-5.6 Luna with “Extra High” reasoning scored 84.04% at a cost of $1.33 at launch. OpenAI describes this result as delivering “essentially the same performance,” although the available search context does not identify the benchmark, comparison baseline, dataset size, or scoring protocol.
That makes the result promising but incomplete. Four details materially affect its interpretation:
- Reasoning setting: “Extra High” suggests a compute-intensive configuration rather than the model’s default behavior.
- Cost date: OpenAI labels $1.33 as the cost “at launch,” not necessarily the cost under September 2026 pricing.
- Benchmark scope: An aggregate percentage cannot reveal performance on coding, multilingual prompts, extraction, tool use, or long-context retrieval individually.
- Repeatability: The cited excerpt does not provide prompt templates, sampling settings, run count, variance, or error bars.
OpenAI subsequently reduced GPT-5.6 Luna’s price by 80%, according to its GPT-5.6 announcement. However, the benchmark cannot simply be repriced to $0.266 by multiplying $1.33 by 20%, because the published context does not disclose the run’s input, cached-input, and output-token mix.
Do the benchmarks prove GPT-5.6 Luna matches larger models?
No single benchmark proves parity across workloads. OpenAI’s 84.04% result indicates a favorable accuracy-to-cost trade-off under one reported evaluation configuration, not equivalent capability on every task.
OpenAI separately calls GPT-5.6 Luna the GPT-5.6 family’s “most cost-efficient” and “fastest” model. Those are product-positioning claims from OpenAI, and the cited materials do not provide a standardized tokens-per-second figure or latency distribution that readers can reproduce.
The evidence therefore supports three conclusions:
- Luna is optimized for economics: OpenAI prices
gpt-5.6-lunaat $0.20 per million input tokens, $0.02 per million cached input tokens, and $1.20 per million output tokens as of September 2026. - Higher reasoning can materially affect results: The published 84.04% score specifically uses the Extra High setting.
- GPT-6 comparisons remain unsupported: OpenAI’s verified GPT-6 materials name GPT-6 Astra, while no primary-source GPT-6 Luna benchmark has been identified.
How should developers benchmark Luna themselves?
A practical evaluation should measure the production task rather than repeat one headline score. Use a fixed test set and compare:
- Task success rate against human-approved answers
- Total cost per successful result, including retries
- Median and 95th-percentile response time
- Output-token usage at each reasoning level
- Tool-call accuracy and structured-output validity
- Long-context retrieval accuracy at multiple document lengths
Run each case several times with identical settings, then report both the mean and failure distribution. Until OpenAI publishes fuller methodology—or documents a separate gpt-6-luna model—the 84.04% at $1.33 figure should be treated as a useful vendor-reported data point, not a comprehensive GPT-6 Luna verdict.
How do developers access GPT-5.6 Luna through the API?

Developers should access Luna with the production model ID gpt-5.6-luna through OpenAI’s Responses API—not with the unverified identifier gpt-6-luna. As of September 22, 2026, OpenAI’s model documentation lists GPT-5.6 Luna for API workloads, while no primary OpenAI source reviewed for this explainer documents a separate GPT-6 Luna endpoint.
What model ID should developers use for GPT-5.6 Luna?
Use the exact, versioned model identifier:
gpt-5.6-lunaOpenAI’s GPT-5.6 Luna model page describes the model as designed for “cost-sensitive, high-volume workloads.” OpenAI’s general model catalog says current models are available through the Responses API.
Do not substitute names seen in headlines or launch discussions:
- Verified:
gpt-5.6-luna - Not independently verified:
gpt-6-luna - Different documented model: GPT-6 Astra
- Not interchangeable: GPT-5.6 Sol, Terra and Luna
Model names are routing identifiers, not descriptive labels. Sending an unrecognized ID can produce an API error, while silently replacing Luna with another model can materially change cost and behavior.
How do you make a GPT-5.6 Luna API request?
A minimal request follows the standard OpenAI Responses API pattern:
curl https://api.openai.com/v1/responses \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-luna",
"input": "Summarize the key risks in this customer-support transcript."
}'A practical implementation sequence is:
- Create or select an OpenAI API project and configure its billing.
- Store the API key in a secret manager rather than source code or browser-side JavaScript.
- Call
/v1/responseswithmodelset togpt-5.6-luna. - Record input, cached-input and output token usage separately.
- Handle unavailable-model and rate-limit errors instead of automatically switching to an unidentified model.
- Pin the model ID in configuration so a naming change does not require an application release.
Developers should confirm the model appears in their project before production rollout because a public documentation page does not necessarily prove availability in every account, region or organizational policy configuration.
How should developers calculate GPT-5.6 Luna API costs?
OpenAI lists GPT-5.6 Luna at $0.20 per million input tokens, $0.02 per million cached input tokens and $1.20 per million output tokens as of September 22, 2026. OpenAI’s GPT-5.6 announcement says Luna received an 80% price reduction and calls it the family’s “most cost-efficient” and “fastest” model.
For example, a workload using 10 million uncached input tokens and 1 million output tokens would cost:
- Input: 10 × $0.20 = $2.00
- Output: 1 × $1.20 = $1.20
- Total model-token cost: $3.20
Cached-input pricing applies only when requests qualify for OpenAI’s caching mechanism; developers should not assume every repeated token receives the $0.02-per-million rate.
For multi-model testing, CallMissed’s OpenAI-compatible developer API provides one API key and one balance for 136 models, with existing SDKs able to switch by changing the base URL. Teams should still verify the live catalog before assuming that any newly announced model is available through an intermediary.
Where is Luna strongest, and what are its limitations?

GPT-5.6 Luna is strongest where throughput, low token cost, and adequate frontier-model capability matter more than maximum reasoning depth. Its biggest limitation is not necessarily technical performance but verification: as of September 22, 2026, OpenAI has not published primary documentation confirming a separate GPT-6 Luna model, API ID, context limit, or benchmark suite.
What workloads suit GPT-5.6 Luna best?
OpenAI describes GPT-5.6 Luna as designed for “cost-sensitive, high-volume workloads” and calls it the GPT-5.6 family’s “fastest” and “most cost-efficient” model. OpenAI does not provide a universal latency figure, so “fastest” should be understood as a family-level positioning claim rather than a guaranteed response time.
Luna is particularly attractive for:
- Document classification and extraction: Processing invoices, tickets, product records, or compliance documents where per-token economics determine whether automation is viable.
- Search and retrieval pipelines: Generating query rewrites, ranking candidates, extracting citations, or synthesizing retrieved passages.
- Customer-support automation: Drafting replies, summarising conversations, detecting intent, and routing requests at high volume.
- Batch-like application workflows: Although no dedicated batch API is confirmed in the supplied OpenAI sources, developers can submit ordinary API requests through their own queues and concurrency controls.
- Agent subtasks: Luna can handle routine tool selection, structured outputs, summarisation, and validation while a more capable model is reserved for difficult reasoning.
At $0.20 per million input tokens as of September 2026, processing 100 million uncached input tokens would cost $20 before output charges. The same volume served from cache would cost $2, based on OpenAI’s $0.02-per-million cached-input rate.
How strong are GPT-5.6 Luna’s benchmarks?
The clearest published result is promising but incomplete. OpenAI’s Builder’s Guide reports that GPT-5.6 Luna with “Extra High” effort scored 84.04% at a cost of $1.33 at launch.
That result should not be treated as proof that Luna scores 84.04% across coding, mathematics, agents, or factuality. The benchmark name, dataset composition, scoring method, token usage, and comparison conditions must accompany the number before teams can reproduce it or use it for procurement.
A defensible evaluation should measure:
- Task success rate on representative production inputs.
- Cost per successful task, not merely cost per token.
- Tail latency, including the 95th and 99th percentiles.
- Schema-valid output rate for structured workflows.
- Human escalation or correction rate for customer-facing tasks.
What are Luna’s main limitations?
The most important constraints are:
- Unverified GPT-6 identity:
gpt-6-lunashould not be placed in production configuration until OpenAI documents that identifier. - Unresolved context capacity: Reported 1.05-million- and 1.1-million-token windows remain provisional; neither should be presented as a confirmed usable input allowance.
- Output-heavy economics: OpenAI charges $1.20 per million output tokens, six times the uncached input rate, so verbose generation can dominate the bill.
- Limited reproducible benchmark evidence: One headline score cannot establish performance across every domain.
- No guaranteed latency: “Fastest” does not specify region, concurrency, prompt length, service tier, or percentile latency.
- Likely capability trade-offs: Luna’s cost-focused positioning suggests that complex research, difficult coding, or high-stakes reasoning should be tested against GPT-5.6 Sol, Terra, and GPT-6 Astra rather than assigned automatically.
In practice, Luna is a compelling high-volume workhorse, but developers should validate accuracy, context behavior, and model availability against live OpenAI documentation before deployment.
Which model should you choose for each workload?

Choose GPT-5.6 Luna for high-volume, price-sensitive tasks; GPT-5.6 Sol when stronger reasoning justifies higher token costs; and GPT-6 Astra only when frontier capability is worth a carefully benchmarked premium. Do not configure production systems with gpt-6-luna until OpenAI publishes that model ID in its API documentation.
Which OpenAI model fits each workload?
| Workload | Recommended model | Why it fits | Cost or verification note |
|---|---|---|---|
| Classification, tagging and routing at scale | GPT-5.6 Luna | Optimized for fast, cost-sensitive, high-volume processing | OpenAI lists gpt-5.6-luna at $0.20/M input and $1.20/M output tokens as of September 2026 |
| Customer-support drafts and conversational assistants | GPT-5.6 Luna | Low token prices support frequent, short interactions and multiple prompt iterations | Cached input costs $0.02/M tokens, according to OpenAI’s September 2026 model page |
| Document extraction and large RAG pipelines | GPT-5.6 Luna, with testing | Cheap input makes retrieval results, policies and document batches economical to process | Treat reported 1.05M–1.1M context limits as provisional until OpenAI confirms the shared input-output limit |
| Complex coding, analysis and multi-step workflows | GPT-5.6 Sol | A better candidate where task quality matters more than minimum cost | OpenAI lists Sol at $4/M input and $20/M output tokens as of September 2026 |
| Highest-stakes frontier reasoning | GPT-6 Astra | Evaluate for difficult tasks where better outcomes could outweigh substantially higher costs | OpenAI associates GPT-6 with Astra; verify the current endpoint, tier and pricing before deployment |
| Mixed or unpredictable production traffic | Luna first, stronger-model fallback | Handles routine requests cheaply while escalating only difficult cases | Requires application-level routing, evaluation thresholds and cost monitoring |
How large is the price difference in practice?
Consider a monthly pipeline processing 10 million uncached input tokens and 1 million output tokens:
- GPT-5.6 Luna:
(10 × $0.20) + (1 × $1.20) = $3.20 - GPT-5.6 Sol:
(10 × $4) + (1 × $20) = $60
At those published September 2026 rates, Sol costs $56.80 more for that token mix. This does not prove that Luna is always the better economic choice: a more capable model can still cost less overall if it reduces retries, human review or workflow failures.
Caching can change the calculation further. If all 10 million Luna input tokens qualified for cached pricing, input processing would cost $0.20 rather than $2, although actual cache eligibility depends on prompt structure and repeated prefixes.
How should teams make the final selection?
Use a representative evaluation set rather than selecting from headline benchmarks alone:
- Measure task success, not just stylistic preference.
- Record total input, cached-input and output tokens.
- Include retries, tool calls and human corrections in effective cost.
- Test latency under realistic concurrency.
- Confirm the exact production model ID before release.
OpenAI reports that GPT-5.6 Luna achieved 84.04% at a cost of $1.33 in an “Extra High” configuration in its builder guide, but that result should not be generalized without the benchmark definition and test conditions.
For multi-model evaluation, CallMissed’s OpenAI-compatible developer API provides one key and balance across 136 models, including caller-selected fallback models and request logs. That architecture can simplify controlled comparisons, but each candidate should still be validated against the workload’s own quality, latency and cost thresholds.
Frequently Asked Questions

Did GPT-6 Luna officially launch, and is it available through the OpenAI API?
What is GPT-5.6 Luna API pricing in September 2026?
What is the correct GPT-5.6 Luna model ID for API requests?
gpt-5.6-luna, and OpenAI’s model catalogue says GPT-5.6 models are accessible through the Responses API as of September 22, 2026. A separate gpt-6-luna identifier has not been independently verified, so using that presumed name in production could return a model-not-found error or create an unsafe dependency on undocumented behavior.What is the verified GPT-5.6 Luna context window?
How strong are GPT-5.6 Luna API pricing and benchmarks compared with larger models?
What workloads are best suited to GPT-5.6 Luna, and how should developers test it?
Conclusion
The practical conclusion is clear: GPT-5.6 Luna—not a separately documented GPT-6 Luna—is the production model developers can currently verify and use. Editorial confirmation of the GPT-6 Luna launch should not override OpenAI’s published model catalog, API identifiers, or pricing documentation.
- Use the verified model ID
gpt-5.6-luna. As of September 22, 2026, OpenAI associates GPT-6 with Astra and has not published a distinctgpt-6-lunaproduction identifier. - Budget using documented API rates. OpenAI lists GPT-5.6 Luna at $0.20 per million input tokens, $0.02 per million cached input tokens, and $1.20 per million output tokens as of September 2026.
- Treat context claims cautiously. The reported 1.05-million- and 1.1-million-token windows remain provisional until OpenAI confirms the precise shared input-output constraints.
- Interpret benchmarks in context. OpenAI reports an 84.04% “Extra High” score at $1.33, but developers should verify the benchmark, settings, latency, and workload fit before extrapolating.
Next, watch for an official GPT-6 Luna model page, release note, price sheet, context specification, and stable API ID. Developers can also explore CallMissed, an OpenAI-compatible AI gateway offering 136 models through one API key and balance, to compare available models without rebuilding integrations.
Will GPT-6 Luna become a documented production tier—or will Luna remain part of the GPT-5.6 family?
Related Reading
- GPT-OSS 120B Open Source Performance on Coding: Specs, Benchmarks, Pricing, and API Access
- GPT-6 API Guide: Sol, Luna and Astra Facts for 2026
- Claude Fable 5.1 vs GPT-6 Astra Pricing: Real API Costs in 2026
Sources
Discussion
Related Posts
Ready to automate customer conversations?
Launch AI voice agents and WhatsApp bots with CallMissed — one API, 22+ Indian languages.



