Skip to content

Explore CallMissed

model launch explainer

GPT-5.6 Luna API Pricing, Context Window & Benchmarks

CallMissed logo
CallMissed Team
·23 min read
GPT-5.6 Luna API Pricing, Context Window & Benchmarks

Verify GPT-5.6 Luna API pricing, model IDs, context and output limits, availability, benchmarks, strengths and ideal workloads.

CallMissed logo

CallMissed

AI Communication Platform

Build AI-powered voice agents, WhatsApp bots, and customer engagement workflows.

Try free

GPT-5.6 Luna API Pricing, Context Window & Benchmarks

What if a frontier-model API could process 1 million input tokens for just $0.20? That headline makes GPT-5.6 Luna API pricing unusually compelling for high-volume applications—but it also sits amid confusing claims about a “GPT-6 Luna” launch, conflicting context-window figures, and search snippets that mix Luna’s prices with those of more expensive models.

Here is the verified picture as of September 2026: OpenAI’s current API documentation identifies the public model as GPT-5.6 Luna, with the model ID gpt-5.6-luna, and describes it as designed for “cost-sensitive, high-volume workloads.” OpenAI lists GPT-5.6 Luna at $0.20 per million input tokens, $0.02 per million cached input tokens, and $1.20 per million output tokens as of September 2026. OpenAI’s GPT-5.6 announcement also states that Luna’s price was reduced by 80%, positioning it as the GPT-5.6 family’s “most cost-efficient” and “fastest” option.

The reported GPT-6 Luna launch requires an important qualification. Although the launch has been confirmed editorially, the primary OpenAI materials available for verification currently associate the Luna name with GPT-5.6 and the Astra name with GPT-6. A distinct GPT-6 Luna model page, API identifier, price sheet, or release announcement has not been independently verified from the cited primary sources. Developers should therefore avoid assuming that gpt-6-luna is a valid production model ID until OpenAI documents it explicitly.

That distinction matters because a naming mistake can break deployments, while misunderstanding token limits or cached-input pricing can distort cost forecasts by thousands of dollars at scale. OpenAI’s builder guide reports a GPT-5.6 Luna “Extra High” score of 84.04% at a cost of $1.33, but the benchmark name and test conditions must be checked before treating that figure as a universal performance measure. Likewise, widely repeated 1.05-million- and 1.1-million-token context-window claims should remain provisional until the exact shared input-output limits are confirmed in current model documentation.

This guide separates confirmed specifications from unresolved claims, covering availability, API use, model IDs, context limits, pricing, benchmarks, strengths, limitations, and ideal workloads. For developers comparing providers, CallMissed’s OpenAI-compatible gateway offers access to 136 models through one API key and balance, making it possible to evaluate models without rebuilding each integration.

Is GPT-6 Luna real? OpenAI currently lists GPT-5.6 Luna

A technology journalist and software engineer review two official-looking documentation pages on adjacent monitors in a
A technology journalist and software engineer review two official-looking documentation pages on adjacent monitors in a

Yes—GPT-6 Luna’s launch existence is supported by an OpenAI-domain product result dated September 22, 2026. The surfaced result is titled “Introducing GPT-6 Sol and Luna,” which supersedes the earlier conclusion that Luna was documented only as GPT-5.6.

However, the available search did not return a corresponding OpenAI developer model page, API reference or pricing entry for GPT-6 Luna. The announcement evidence therefore confirms the product name, but not its API specifications.

What does OpenAI officially call Luna?

OpenAI now appears to use Luna in two distinct model generations:

  • GPT-5.6 Luna, the previously documented cost-efficient model with the API identifier gpt-5.6-luna.
  • GPT-6 Luna, named alongside GPT-6 Sol in an OpenAI-domain product snippet dated September 22, 2026.

These models must not be treated as interchangeable. In particular, GPT-5.6 Luna’s model ID, pricing, context limits and benchmark results do not establish the corresponding details for GPT-6 Luna.

What is verified about GPT-6 Luna?

ClaimStatus as of September 22, 2026Evidence
OpenAI has introduced a product named GPT-6 LunaSupportedOpenAI-domain result titled “Introducing GPT-6 Sol and Luna”
GPT-6 Luna is distinct from GPT-5.6 LunaSupported by the namingThe new result identifies GPT-6, while existing documentation separately names GPT-5.6 Luna
Public API model ID is gpt-6-lunaUnverifiedNo matching developer model page or API reference was returned
GPT-6 Luna API pricingUnverifiedNo GPT-6 Luna pricing entry was available in the surfaced primary documentation
GPT-6 Luna context windowUnverifiedNo matching technical model page was returned
GPT-6 Luna benchmark resultsUnverifiedNo GPT-6 Luna-specific benchmark documentation was located
General API availability or rollout tiersUnverifiedThe product snippet alone does not establish endpoint access or account eligibility

How does GPT-5.6 Luna differ?

The existing GPT-5.6 Luna documentation remains evidence about GPT-5.6 Luna only. It identifies the production model as gpt-5.6-luna and lists pricing of $0.20 per million input tokens, $0.02 per million cached input tokens and $1.20 per million output tokens.

Likewise, OpenAI’s reported 80% price reduction and the builder-guide result of 84.04% at a cost of $1.33 at “Extra High” belong to GPT-5.6 Luna. They should not be quoted as GPT-6 Luna prices or benchmarks without new GPT-6-specific documentation.

Should developers use gpt-6-luna in API requests?

Not solely on the basis of the launch snippet. A product announcement does not prove that a predictable API identifier is active. Until OpenAI exposes GPT-6 Luna in its developer catalogue or API documentation, developers should:

  • Check the live provider model list before sending requests.
  • Use gpt-5.6-luna only when intentionally selecting GPT-5.6 Luna.
  • Avoid assuming that gpt-6-luna is the production identifier.
  • Validate pricing, context limits and supported features against a GPT-6 Luna-specific primary source.
  • Run a small test request before changing production routing.

The accurate conclusion as of September 22, 2026 is that GPT-6 Luna’s launch is supported by an official OpenAI-domain product snippet, while its API model ID, pricing, context window, benchmarks and developer availability remain unverified from the available primary documentation.

How do GPT-5.6 Luna, Sol, Terra and GPT-6 Astra fit together?

A clean vertical model-family map titled OPENAI MODEL NAMING MAP
A clean vertical model-family map titled OPENAI MODEL NAMING MAP

GPT-5.6 Luna, Sol and Terra are three members of OpenAI’s GPT-5.6 family, while GPT-6 Astra belongs to the newer GPT-6 generation. As of September 22, 2026, OpenAI’s primary documentation does not show GPT-6 Luna as a separate fourth tier; it identifies Luna as gpt-5.6-luna and Astra as the model powering GPT-6 Pro.

What is the GPT-5.6 Luna, Sol and Terra hierarchy?

The GPT-5.6 lineup is best understood as a cost, speed and capability spectrum, not as three successive releases:

  1. GPT-5.6 Luna targets cost-sensitive, high-volume processing. OpenAI calls Luna the GPT-5.6 family’s “most cost-efficient” and “fastest” model.
  2. GPT-5.6 Sol occupies a more expensive tier for workloads where stronger model capability can justify higher token costs.
  3. GPT-5.6 Terra is another GPT-5.6 tier, but the supplied primary-source excerpts do not establish its current model ID, exact pricing or intended workload clearly enough to state them as verified facts.
  4. GPT-6 Astra sits outside that family as a GPT-6 model and powers the GPT-6 Pro experience described by OpenAI.

This distinction prevents a common implementation error: treating Luna, Sol, Terra and Astra as interchangeable aliases or assuming that gpt-6-luna follows automatically from gpt-5.6-luna.

How different are Luna and Sol API costs?

GPT-5.6 Luna has a substantially lower per-token price than GPT-5.6 Sol. According to OpenAI’s model documentation dated September 2026:

  • GPT-5.6 Luna: $0.20 per million input tokens and $1.20 per million output tokens.
  • GPT-5.6 Sol: $4 per million input tokens and $20 per million output tokens.
  • Luna’s listed input price is therefore 20 times lower than Sol’s, while its output price is approximately 16.7 times lower.
  • OpenAI says Sol’s current pricing represents a 20% input-price reduction and a 33% output-price reduction.
  • OpenAI’s GPT-5.6 announcement says Luna received an 80% price reduction, while Terra received a 20% reduction.

These figures make Luna the natural starting point for classification, extraction, routing, summarisation and other high-throughput jobs. Sol may be more economical when better results reduce retries, human review or multi-step prompting, but that trade-off must be tested on the application’s own evaluations.

Where does GPT-6 Astra fit?

GPT-6 Astra is the higher-generation option, not a renamed GPT-5.6 Luna tier. OpenAI’s Help Center states that GPT-6 Pro, powered by GPT-6 Astra, is available in ChatGPT on the Pro $100, Pro $200, Business and Enterprise plans as of September 2026; Enterprise access can also depend on organisational settings.

Developers should keep two access layers separate:

  • ChatGPT availability determines which users can select GPT-6 Pro in the ChatGPT interface.
  • API availability requires a documented API model ID, endpoint eligibility, rate limits and an unambiguous price row.
  • GPT-6 Luna availability remains unverified until OpenAI publishes those details under that specific name.

The practical selection order is therefore: choose Luna for economical scale, evaluate Sol or Terra where additional capability warrants greater cost, and assess Astra for GPT-6-class workloads—without presuming that benchmark results, context limits or API identifiers transfer between tiers.

What launch dates, specifications and availability are confirmed?

A structured specification matrix titled GPT-5.6 LUNA: VERIFIED SPECIFICATION CHECKLIST with rows labelled Official name,
A structured specification matrix titled GPT-5.6 LUNA: VERIFIED SPECIFICATION CHECKLIST with rows labelled Official name,

The editorially confirmed GPT-6 Luna launch does not yet have a matching public OpenAI specification page, model ID, price sheet, or independently verifiable release date. As of September 22, 2026, the production-ready Luna model documented by OpenAI is GPT-5.6 Luna, available through the OpenAI Responses API as gpt-5.6-luna.

Which GPT-6 Luna launch details are confirmed?

The table separates details verified in current OpenAI primary sources from claims that remain editorially confirmed or unresolved.

ItemConfirmed detailAvailability or valueVerification status
GPT-6 Luna launchThe launch is confirmed by the site administratorExact public launch date not establishedEditorially confirmed; not independently verified
Public Luna modelGPT-5.6 LunaOpenAI API documentation was live as of September 22, 2026Verified by OpenAI
Production model IDgpt-5.6-lunaListed for API useVerified by OpenAI
API accessOpenAI Responses APIIntended for developer workloadsVerified by OpenAI Models documentation
Token pricing$0.20 input, $0.02 cached input, $1.20 output per million tokensStandard GPT-5.6 Luna API ratesVerified by OpenAI Pricing
Context windowPublic claims cite approximately 1.05 million or 1.1 million tokensExact shared input-output limit remains unresolvedNot independently verified

OpenAI describes GPT-5.6 Luna as designed for “cost-sensitive, high-volume workloads.” OpenAI’s GPT-5.6 announcement also calls Luna the family’s “most cost-efficient” and “fastest” model and reports that its price was reduced by 80%.

What specifications can developers safely use?

For implementation and budgeting, developers can currently rely on these documented values as of September 2026:

  • Model name: GPT-5.6 Luna
  • API identifier: gpt-5.6-luna
  • Input price: $0.20 per 1 million tokens
  • Cached-input price: $0.02 per 1 million tokens
  • Output price: $1.20 per 1 million tokens
  • Documented interface: OpenAI Responses API
  • Primary positioning: fast, economical processing for high-volume applications

At those rates, processing 100 million uncached input tokens would cost $20, excluding output tokens and other platform charges. The same quantity served entirely as cached input would cost $2, illustrating why prompt caching can materially change production economics.

Which launch claims remain unresolved?

Three details should not yet be treated as production facts:

  1. gpt-6-luna is not a verified model ID. Sending it in an API request could produce an invalid-model error unless OpenAI subsequently documents the identifier.
  2. The exact GPT-6 Luna release date is unavailable in the cited primary materials. The administrator’s launch confirmation establishes the editorial status, not a public OpenAI publication date.
  3. The context limit is unsettled. Neither the widely repeated 1.05-million-token figure nor the rounded 1.1-million-token claim should be used for architecture decisions without checking the live model schema.

OpenAI’s public naming currently places Luna under GPT-5.6 and Astra under GPT-6. OpenAI’s Help Center states that GPT-6 Pro is powered by GPT-6 Astra and is available in ChatGPT on Pro $100, Pro $200, Business, and Enterprise plans as of September 2026; that statement does not independently establish GPT-6 Luna availability in ChatGPT or the API.

How much does the GPT-5.6 Luna API cost?

A detailed API pricing dashboard titled GPT-5.6 LUNA API PRICING
A detailed API pricing dashboard titled GPT-5.6 LUNA API PRICING

OpenAI charges $0.20 per million fresh input tokens, $0.02 per million cached input tokens, and $1.20 per million output tokens for gpt-5.6-luna as of September 2026. No independently verified price exists for a separate gpt-6-luna API model, so developers should not use GPT-5.6 Luna’s rates as confirmed GPT-6 pricing.

What are the confirmed GPT-5.6 Luna token prices?

OpenAI’s GPT-5.6 Luna model page describes Luna as intended for “cost-sensitive, high-volume workloads.” OpenAI also says it reduced GPT-5.6 Luna’s price by 80%, although that announcement should not be interpreted as confirming prices for an undocumented GPT-6 Luna model.

Usage scenarioInput costOutput costEstimated total
1M fresh input tokens$0.20$0.00$0.20
1M cached input tokens$0.02$0.00$0.02
1M output tokens$0.00$1.20$1.20
1M fresh input + 1M output$0.20$1.20$1.40
10M fresh input + 1M output$2.00$1.20$3.20
100M fresh input + 10M output$20.00$12.00$32.00

Calculations use OpenAI’s standard gpt-5.6-luna API rates published as of September 2026; they exclude unrelated platform, storage, tool, search, or infrastructure charges.

The pricing formula is:

Total cost = (fresh input tokens ÷ 1,000,000 × $0.20) + (cached input tokens ÷ 1,000,000 × $0.02) + (output tokens ÷ 1,000,000 × $1.20).

How much cheaper is cached input?

GPT-5.6 Luna cached input costs 90% less than fresh input, according to OpenAI’s September 2026 pricing. Processing one billion fully cache-eligible input tokens would therefore cost $20, compared with $200 for one billion fresh input tokens.

Caching can materially reduce costs when applications repeatedly send the same stable content, such as:

  • Long system instructions
  • Shared policy or product documentation
  • Reused codebase context
  • Fixed schemas and tool definitions
  • Common few-shot examples

Developers should not assume every prompt token receives the cached rate. Only input recognized as cache-eligible is billed at $0.02 per million tokens; new or changed prompt content remains fresh input.

How does Luna pricing compare with GPT-5.6 Sol?

OpenAI lists GPT-5.6 Sol at $4 per million input tokens and $20 per million output tokens as of September 2026. At published standard rates, Luna’s $0.20 input price is 95% lower than Sol’s, while Luna’s $1.20 output price is 94% lower than Sol’s.

That gap makes Luna financially attractive for classification, extraction, summarization, routing, and other high-volume tasks where its quality is sufficient. The trade-off should still be tested against accuracy and reliability requirements rather than decided on token price alone.

For multi-model evaluations, CallMissed’s OpenAI-compatible developer API provides one key and one balance for 136 models, allowing teams to compare model economics without rewriting an existing OpenAI SDK integration. Regardless of gateway, production budgets should include output length, cache-hit rates, retries, and fallback-model usage—not just the headline input price.

Is the context window 1.05M or 1.1M tokens?

A token-capacity diagram titled CONTEXT LIMITS EXPLAINED showing one long horizontal container labelled Documented context
A token-capacity diagram titled CONTEXT LIMITS EXPLAINED showing one long horizontal container labelled Documented context

Neither 1.05 million nor 1.1 million tokens is independently confirmed as the exact GPT-5.6 Luna context window in the cited OpenAI materials. As of September 22, 2026, treat both figures as provisional descriptions—not production-safe limits—until OpenAI publishes an explicit token count for gpt-5.6-luna.

Why do sources report both 1.05M and 1.1M tokens?

The discrepancy may be a matter of rounding rather than two different model versions. For example, an underlying limit of 1,050,000 tokens could be displayed as:

  • 1.05M tokens when reported to two decimal places.
  • Approximately 1.1M tokens when rounded to one decimal place.
  • “Over one million tokens” in nontechnical launch material.

That explanation is plausible, but it is not proof that 1,050,000 is Luna’s actual limit. A model snapshot change, documentation error, or confusion between an advertised input capacity and a shared input-output budget could produce the same conflicting figures.

OpenAI’s GPT-5.6 Luna model page confirms that Luna targets “cost-sensitive, high-volume workloads” and identifies the production API model as gpt-5.6-luna as of September 2026. However, the cited model-page and pricing-page extracts do not state an exact context-window limit, so neither 1.05M nor 1.1M can be attributed conclusively to OpenAI from the available evidence.

Does the context window include output tokens?

Developers should assume the advertised context figure may represent a shared token budget unless OpenAI explicitly separates maximum input and maximum output limits. In a shared window, prompt tokens, retrieved documents, tool results, conversation history, reasoning-related tokens where applicable, and generated output all compete for capacity.

A simplified planning formula is:

usable input ≤ total context limit − reserved output − system and tool overhead

If a workload sends 1,000,000 input tokens and reserves 50,000 tokens for generation, it already requires at least 1,050,000 tokens of combined capacity before allowing for system instructions or tool schemas. That request might fit a genuine 1.1M shared window but fail against a 1.05M ceiling once overhead is counted.

Which limit should developers use in production?

Until OpenAI documents the exact number, use a deliberately conservative operating limit:

  1. Do not hard-code 1.1M based on rounded launch coverage.
  2. Keep prompts materially below 1.05M, with room for output and orchestration overhead.
  3. Count tokens using the tokenizer supported for gpt-5.6-luna.
  4. Test against the production API and record the model snapshot, request size, and error response.
  5. Recheck OpenAI’s model documentation before raising application limits.

API acceptance tests can reveal what works at a particular moment, but they do not replace a documented guarantee; routing changes and model snapshots can alter observed behavior.

What is the verified conclusion?

The defensible specification is currently “approximately one million tokens, exact limit unverified.” Use 1.05M only as a provisional lower claim and 1.1M only as an approximate upper description—not as an assured request budget. The exact maximum input, maximum output, and shared-context rules remain unresolved pending explicit OpenAI documentation for GPT-5.6 Luna.

What do the GPT-5.6 Luna benchmarks actually show?

A benchmark evidence board titled READ GPT-5.6 LUNA BENCHMARKS CAREFULLY divided into three columns: OpenAI-reported
A benchmark evidence board titled READ GPT-5.6 LUNA BENCHMARKS CAREFULLY divided into three columns: OpenAI-reported

The verified benchmark evidence supports a narrow conclusion: GPT-5.6 Luna can deliver strong task performance at unusually low API cost, but the published 84.04% result does not establish universal model quality. No independently verified benchmark suite or score is currently available for a distinct GPT-6 Luna model as of September 22, 2026.

What does the 84.04% GPT-5.6 Luna score mean?

OpenAI’s Builder’s Guide to GPT-5.6 reports that GPT-5.6 Luna with “Extra High” reasoning scored 84.04% at a cost of $1.33 at launch. OpenAI describes this result as delivering “essentially the same performance,” although the available search context does not identify the benchmark, comparison baseline, dataset size, or scoring protocol.

That makes the result promising but incomplete. Four details materially affect its interpretation:

  • Reasoning setting: “Extra High” suggests a compute-intensive configuration rather than the model’s default behavior.
  • Cost date: OpenAI labels $1.33 as the cost “at launch,” not necessarily the cost under September 2026 pricing.
  • Benchmark scope: An aggregate percentage cannot reveal performance on coding, multilingual prompts, extraction, tool use, or long-context retrieval individually.
  • Repeatability: The cited excerpt does not provide prompt templates, sampling settings, run count, variance, or error bars.

OpenAI subsequently reduced GPT-5.6 Luna’s price by 80%, according to its GPT-5.6 announcement. However, the benchmark cannot simply be repriced to $0.266 by multiplying $1.33 by 20%, because the published context does not disclose the run’s input, cached-input, and output-token mix.

Do the benchmarks prove GPT-5.6 Luna matches larger models?

No single benchmark proves parity across workloads. OpenAI’s 84.04% result indicates a favorable accuracy-to-cost trade-off under one reported evaluation configuration, not equivalent capability on every task.

OpenAI separately calls GPT-5.6 Luna the GPT-5.6 family’s “most cost-efficient” and “fastest” model. Those are product-positioning claims from OpenAI, and the cited materials do not provide a standardized tokens-per-second figure or latency distribution that readers can reproduce.

The evidence therefore supports three conclusions:

  1. Luna is optimized for economics: OpenAI prices gpt-5.6-luna at $0.20 per million input tokens, $0.02 per million cached input tokens, and $1.20 per million output tokens as of September 2026.
  2. Higher reasoning can materially affect results: The published 84.04% score specifically uses the Extra High setting.
  3. GPT-6 comparisons remain unsupported: OpenAI’s verified GPT-6 materials name GPT-6 Astra, while no primary-source GPT-6 Luna benchmark has been identified.

How should developers benchmark Luna themselves?

A practical evaluation should measure the production task rather than repeat one headline score. Use a fixed test set and compare:

  • Task success rate against human-approved answers
  • Total cost per successful result, including retries
  • Median and 95th-percentile response time
  • Output-token usage at each reasoning level
  • Tool-call accuracy and structured-output validity
  • Long-context retrieval accuracy at multiple document lengths

Run each case several times with identical settings, then report both the mean and failure distribution. Until OpenAI publishes fuller methodology—or documents a separate gpt-6-luna model—the 84.04% at $1.33 figure should be treated as a useful vendor-reported data point, not a comprehensive GPT-6 Luna verdict.

How do developers access GPT-5.6 Luna through the API?

A six-step developer workflow titled FROM MODEL DOCS TO PRODUCTION flowing left to right with arrows: 1
A six-step developer workflow titled FROM MODEL DOCS TO PRODUCTION flowing left to right with arrows: 1

Developers should access Luna with the production model ID gpt-5.6-luna through OpenAI’s Responses API—not with the unverified identifier gpt-6-luna. As of September 22, 2026, OpenAI’s model documentation lists GPT-5.6 Luna for API workloads, while no primary OpenAI source reviewed for this explainer documents a separate GPT-6 Luna endpoint.

What model ID should developers use for GPT-5.6 Luna?

Use the exact, versioned model identifier:

text
gpt-5.6-luna

OpenAI’s GPT-5.6 Luna model page describes the model as designed for “cost-sensitive, high-volume workloads.” OpenAI’s general model catalog says current models are available through the Responses API.

Do not substitute names seen in headlines or launch discussions:

  • Verified: gpt-5.6-luna
  • Not independently verified: gpt-6-luna
  • Different documented model: GPT-6 Astra
  • Not interchangeable: GPT-5.6 Sol, Terra and Luna

Model names are routing identifiers, not descriptive labels. Sending an unrecognized ID can produce an API error, while silently replacing Luna with another model can materially change cost and behavior.

How do you make a GPT-5.6 Luna API request?

A minimal request follows the standard OpenAI Responses API pattern:

bash
curl https://api.openai.com/v1/responses \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-luna",
    "input": "Summarize the key risks in this customer-support transcript."
  }'

A practical implementation sequence is:

  1. Create or select an OpenAI API project and configure its billing.
  2. Store the API key in a secret manager rather than source code or browser-side JavaScript.
  3. Call /v1/responses with model set to gpt-5.6-luna.
  4. Record input, cached-input and output token usage separately.
  5. Handle unavailable-model and rate-limit errors instead of automatically switching to an unidentified model.
  6. Pin the model ID in configuration so a naming change does not require an application release.

Developers should confirm the model appears in their project before production rollout because a public documentation page does not necessarily prove availability in every account, region or organizational policy configuration.

How should developers calculate GPT-5.6 Luna API costs?

OpenAI lists GPT-5.6 Luna at $0.20 per million input tokens, $0.02 per million cached input tokens and $1.20 per million output tokens as of September 22, 2026. OpenAI’s GPT-5.6 announcement says Luna received an 80% price reduction and calls it the family’s “most cost-efficient” and “fastest” model.

For example, a workload using 10 million uncached input tokens and 1 million output tokens would cost:

  • Input: 10 × $0.20 = $2.00
  • Output: 1 × $1.20 = $1.20
  • Total model-token cost: $3.20

Cached-input pricing applies only when requests qualify for OpenAI’s caching mechanism; developers should not assume every repeated token receives the $0.02-per-million rate.

For multi-model testing, CallMissed’s OpenAI-compatible developer API provides one API key and one balance for 136 models, with existing SDKs able to switch by changing the base URL. Teams should still verify the live catalog before assuming that any newly announced model is available through an intermediary.

Where is Luna strongest, and what are its limitations?

A balanced workload decision matrix titled GPT-5.6 LUNA: FIT AND TRADE-OFFS
A balanced workload decision matrix titled GPT-5.6 LUNA: FIT AND TRADE-OFFS

GPT-5.6 Luna is strongest where throughput, low token cost, and adequate frontier-model capability matter more than maximum reasoning depth. Its biggest limitation is not necessarily technical performance but verification: as of September 22, 2026, OpenAI has not published primary documentation confirming a separate GPT-6 Luna model, API ID, context limit, or benchmark suite.

What workloads suit GPT-5.6 Luna best?

OpenAI describes GPT-5.6 Luna as designed for “cost-sensitive, high-volume workloads” and calls it the GPT-5.6 family’s “fastest” and “most cost-efficient” model. OpenAI does not provide a universal latency figure, so “fastest” should be understood as a family-level positioning claim rather than a guaranteed response time.

Luna is particularly attractive for:

  • Document classification and extraction: Processing invoices, tickets, product records, or compliance documents where per-token economics determine whether automation is viable.
  • Search and retrieval pipelines: Generating query rewrites, ranking candidates, extracting citations, or synthesizing retrieved passages.
  • Customer-support automation: Drafting replies, summarising conversations, detecting intent, and routing requests at high volume.
  • Batch-like application workflows: Although no dedicated batch API is confirmed in the supplied OpenAI sources, developers can submit ordinary API requests through their own queues and concurrency controls.
  • Agent subtasks: Luna can handle routine tool selection, structured outputs, summarisation, and validation while a more capable model is reserved for difficult reasoning.

At $0.20 per million input tokens as of September 2026, processing 100 million uncached input tokens would cost $20 before output charges. The same volume served from cache would cost $2, based on OpenAI’s $0.02-per-million cached-input rate.

How strong are GPT-5.6 Luna’s benchmarks?

The clearest published result is promising but incomplete. OpenAI’s Builder’s Guide reports that GPT-5.6 Luna with “Extra High” effort scored 84.04% at a cost of $1.33 at launch.

That result should not be treated as proof that Luna scores 84.04% across coding, mathematics, agents, or factuality. The benchmark name, dataset composition, scoring method, token usage, and comparison conditions must accompany the number before teams can reproduce it or use it for procurement.

A defensible evaluation should measure:

  1. Task success rate on representative production inputs.
  2. Cost per successful task, not merely cost per token.
  3. Tail latency, including the 95th and 99th percentiles.
  4. Schema-valid output rate for structured workflows.
  5. Human escalation or correction rate for customer-facing tasks.

What are Luna’s main limitations?

The most important constraints are:

  • Unverified GPT-6 identity: gpt-6-luna should not be placed in production configuration until OpenAI documents that identifier.
  • Unresolved context capacity: Reported 1.05-million- and 1.1-million-token windows remain provisional; neither should be presented as a confirmed usable input allowance.
  • Output-heavy economics: OpenAI charges $1.20 per million output tokens, six times the uncached input rate, so verbose generation can dominate the bill.
  • Limited reproducible benchmark evidence: One headline score cannot establish performance across every domain.
  • No guaranteed latency: “Fastest” does not specify region, concurrency, prompt length, service tier, or percentile latency.
  • Likely capability trade-offs: Luna’s cost-focused positioning suggests that complex research, difficult coding, or high-stakes reasoning should be tested against GPT-5.6 Sol, Terra, and GPT-6 Astra rather than assigned automatically.

In practice, Luna is a compelling high-volume workhorse, but developers should validate accuracy, context behavior, and model availability against live OpenAI documentation before deployment.

Which model should you choose for each workload?

A practical model-selection table titled WHAT THIS MEANS FOR DEVELOPERS with columns Workload, Likely starting point, Why,
A practical model-selection table titled WHAT THIS MEANS FOR DEVELOPERS with columns Workload, Likely starting point, Why,

Choose GPT-5.6 Luna for high-volume, price-sensitive tasks; GPT-5.6 Sol when stronger reasoning justifies higher token costs; and GPT-6 Astra only when frontier capability is worth a carefully benchmarked premium. Do not configure production systems with gpt-6-luna until OpenAI publishes that model ID in its API documentation.

Which OpenAI model fits each workload?

WorkloadRecommended modelWhy it fitsCost or verification note
Classification, tagging and routing at scaleGPT-5.6 LunaOptimized for fast, cost-sensitive, high-volume processingOpenAI lists gpt-5.6-luna at $0.20/M input and $1.20/M output tokens as of September 2026
Customer-support drafts and conversational assistantsGPT-5.6 LunaLow token prices support frequent, short interactions and multiple prompt iterationsCached input costs $0.02/M tokens, according to OpenAI’s September 2026 model page
Document extraction and large RAG pipelinesGPT-5.6 Luna, with testingCheap input makes retrieval results, policies and document batches economical to processTreat reported 1.05M–1.1M context limits as provisional until OpenAI confirms the shared input-output limit
Complex coding, analysis and multi-step workflowsGPT-5.6 SolA better candidate where task quality matters more than minimum costOpenAI lists Sol at $4/M input and $20/M output tokens as of September 2026
Highest-stakes frontier reasoningGPT-6 AstraEvaluate for difficult tasks where better outcomes could outweigh substantially higher costsOpenAI associates GPT-6 with Astra; verify the current endpoint, tier and pricing before deployment
Mixed or unpredictable production trafficLuna first, stronger-model fallbackHandles routine requests cheaply while escalating only difficult casesRequires application-level routing, evaluation thresholds and cost monitoring

How large is the price difference in practice?

Consider a monthly pipeline processing 10 million uncached input tokens and 1 million output tokens:

  • GPT-5.6 Luna: (10 × $0.20) + (1 × $1.20) = $3.20
  • GPT-5.6 Sol: (10 × $4) + (1 × $20) = $60

At those published September 2026 rates, Sol costs $56.80 more for that token mix. This does not prove that Luna is always the better economic choice: a more capable model can still cost less overall if it reduces retries, human review or workflow failures.

Caching can change the calculation further. If all 10 million Luna input tokens qualified for cached pricing, input processing would cost $0.20 rather than $2, although actual cache eligibility depends on prompt structure and repeated prefixes.

How should teams make the final selection?

Use a representative evaluation set rather than selecting from headline benchmarks alone:

  1. Measure task success, not just stylistic preference.
  2. Record total input, cached-input and output tokens.
  3. Include retries, tool calls and human corrections in effective cost.
  4. Test latency under realistic concurrency.
  5. Confirm the exact production model ID before release.

OpenAI reports that GPT-5.6 Luna achieved 84.04% at a cost of $1.33 in an “Extra High” configuration in its builder guide, but that result should not be generalized without the benchmark definition and test conditions.

For multi-model evaluation, CallMissed’s OpenAI-compatible developer API provides one key and balance across 136 models, including caller-selected fallback models and request logs. That architecture can simplify controlled comparisons, but each candidate should still be validated against the workload’s own quality, latency and cost thresholds.

Frequently Asked Questions

A visual FAQ board titled GPT-5.6 LUNA FAQ arranged as six expandable question cards
A visual FAQ board titled GPT-5.6 LUNA FAQ arranged as six expandable question cards
Did GPT-6 Luna officially launch, and is it available through the OpenAI API?
The GPT-6 Luna launch has been confirmed editorially, but OpenAI’s primary documentation does not independently verify a public model page, API price, or production identifier for it as of September 22, 2026. OpenAI currently associates GPT-6 with GPT-6 Astra, while the documented Luna model remains GPT-5.6 Luna; developers should distinguish the confirmed launch report from specifications that OpenAI has published directly.
What is GPT-5.6 Luna API pricing in September 2026?
OpenAI lists GPT-5.6 Luna API pricing at $0.20 per million input tokens, $0.02 per million cached input tokens, and $1.20 per million output tokens as of September 2026. For example, processing 10 million uncached input tokens and generating 1 million output tokens would cost $3.20, excluding other platform or tool charges; OpenAI says Luna received an 80% price reduction.
What is the correct GPT-5.6 Luna model ID for API requests?
OpenAI’s documented production model ID is gpt-5.6-luna, and OpenAI’s model catalogue says GPT-5.6 models are accessible through the Responses API as of September 22, 2026. A separate gpt-6-luna identifier has not been independently verified, so using that presumed name in production could return a model-not-found error or create an unsafe dependency on undocumented behavior.
What is the verified GPT-5.6 Luna context window?
No exact GPT-5.6 Luna context window should be treated as confirmed here because the available primary-source excerpts do not establish the precise shared input-output limit. Claims of approximately 1.05 million or 1.1 million tokens remain provisional until OpenAI’s current model documentation explicitly states the total context window, maximum output, and whether tool results or reasoning tokens consume that allowance.
How strong are GPT-5.6 Luna API pricing and benchmarks compared with larger models?
OpenAI’s Builder’s Guide to GPT-5.6 reports that GPT-5.6 Luna at Extra High scored 84.04% at a cost of $1.33 at launch, but the excerpted source does not identify enough methodology to generalize that result across every task. Buyers should compare models using their own prompts, output-quality rubric, latency measurements, and complete workload cost rather than interpreting one aggregate percentage as universal proof of capability.
What workloads are best suited to GPT-5.6 Luna, and how should developers test it?
OpenAI describes GPT-5.6 Luna as its fastest and most cost-efficient GPT-5.6 option, aimed at cost-sensitive, high-volume workloads such as classification, extraction, routing, summarization, and repetitive content transformation. Developers should test accuracy, structured-output reliability, cache-hit rates, and total output volume before deployment; as of September 2026, CallMissed’s OpenAI-compatible gateway provides one API key and balance for 136 models, allowing teams to evaluate fallback candidates without rewriting OpenAI-style integrations.

Conclusion

The practical conclusion is clear: GPT-5.6 Luna—not a separately documented GPT-6 Luna—is the production model developers can currently verify and use. Editorial confirmation of the GPT-6 Luna launch should not override OpenAI’s published model catalog, API identifiers, or pricing documentation.

  • Use the verified model ID gpt-5.6-luna. As of September 22, 2026, OpenAI associates GPT-6 with Astra and has not published a distinct gpt-6-luna production identifier.
  • Budget using documented API rates. OpenAI lists GPT-5.6 Luna at $0.20 per million input tokens, $0.02 per million cached input tokens, and $1.20 per million output tokens as of September 2026.
  • Treat context claims cautiously. The reported 1.05-million- and 1.1-million-token windows remain provisional until OpenAI confirms the precise shared input-output constraints.
  • Interpret benchmarks in context. OpenAI reports an 84.04% “Extra High” score at $1.33, but developers should verify the benchmark, settings, latency, and workload fit before extrapolating.

Next, watch for an official GPT-6 Luna model page, release note, price sheet, context specification, and stable API ID. Developers can also explore CallMissed, an OpenAI-compatible AI gateway offering 136 models through one API key and balance, to compare available models without rebuilding integrations.

Will GPT-6 Luna become a documented production tier—or will Luna remain part of the GPT-5.6 family?

Sources

Discussion

Your email is used only to identify you — it is never shown publicly.

Loading discussion…

Related Posts

Ready to automate customer conversations?

Launch AI voice agents and WhatsApp bots with CallMissed — one API, 22+ Indian languages.