Skip to main content
OpenAIFastHigh

GPT-6 Luna

The smallest, lowest-cost tier of OpenAI's GPT-6 generation, at $0.10/$0.50 per million tokens โ€” built for high-volume, latency-sensitive work such as classification, extraction, routing and short-form generation. Cached inputs at a 90% discount. Prompts above 272K input tokens are billed at 2x input / 1.5x output for the whole request.

186/929credits
input / output ยท per 1M tokens
Lowest-cost tier of the GPT-6 generation
1M token context window
128K max output tokens
Knowledge cutoff: May 18, 2026
Cached input pricing (90% discount)
Adjustable reasoning effort (low to max)
Function calling and native vision support

Run it right now

Test this model instantly in the Console Playground โ€” no code required

Sign in to try

Use with AI Assistant

Copy usage instructions for Claude, ChatGPT, or other AI

Model Specifications

Context Window
1M
tokens
Max Output
128K
tokens
Training Cutoff
2026-05
Compatible SDK
OpenAI

Capabilities

Vision
Function Calling
Streaming
JSON Mode
System Prompt

Token Pricing (per 1M tokens)

Token TypeCreditsUSD Equivalent
Input Tokens186$0.12
Output Tokens929$0.62
Cached Tokens18.58$0.01

* 1,500 credits โ‰ˆ $1 (actual charges may vary based on usage)

Quick Start

curl -X POST "https://api.core.today/llm/openai/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
  "model": "gpt-6-luna",
  "messages": [
    {
      "role": "system",
      "content": "Classify the support ticket into one of: billing, bug, feature_request, account. Reply with the label only."
    },
    {
      "role": "user",
      "content": "I was charged twice for my subscription this month."
    }
  ],
  "reasoning_effort": "low",
  "max_completion_tokens": 200
}'

Parameters

ParameterTypeRequiredDefaultDescription
messagesarrayYes-Array of message objects with role and content
modelstringYesgpt-6-lunaModel identifier
max_completion_tokensintegerNo4096Maximum tokens in response (up to 128000). Note: use max_completion_tokens, not max_tokens
reasoning_effortstringNomediumReasoning effort level: low, medium, high, xhigh, or max
lowmediumhighxhighmax
temperaturefloatNo1.0Sampling temperature (0-2)
streambooleanNofalseEnable Server-Sent Events streaming

Examples

High-volume Classification

Low-cost ticket triage with GPT-6 Luna

curl -X POST "https://api.core.today/llm/openai/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
  "model": "gpt-6-luna",
  "messages": [
    {
      "role": "system",
      "content": "Classify the support ticket into one of: billing, bug, feature_request, account. Reply with the label only."
    },
    {
      "role": "user",
      "content": "I was charged twice for my subscription this month."
    }
  ],
  "reasoning_effort": "low",
  "max_completion_tokens": 200
}'

Cached Repeated Context

Reuse a large system prompt with 90% cached input discount

curl -X POST "https://api.core.today/llm/openai/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
  "model": "gpt-6-luna",
  "messages": [
    {
      "role": "system",
      "content": "<large repeated system prompt or codebase context>"
    },
    {
      "role": "user",
      "content": "Summarize the open TODOs and rank them by risk."
    }
  ],
  "temperature": 0.3,
  "max_completion_tokens": 4000
}'

Tips & Best Practices

1Luna ($0.10/$0.50 per M) is the cheapest GPT-6 tier โ€” escalate to GPT-6 Sol ($2/$10) or Astra ($10/$50) when answers need deeper reasoning
2Use cached inputs ($0.01/M, 90% off) for repeated system prompts and RAG context
3Keep prompts under 272K input tokens where possible โ€” above that the whole request bills at 2x input / 1.5x output
4Keep reasoning_effort low for classification and extraction to minimize latency and output tokens
5128K output enables single-shot long-form generation
6Lower temperature (0.2-0.5) for coding and analytical tasks
7Cache writes bill at 1.25x the input rate โ€” implicit caching writes by default, so most of a large first prompt is billed as a cache write, then reads are 90% off
8Fast mode (service_tier: fast or priority) bills token charges at 2x โ€” it is settled on the tier the response reports, so a request downgraded to standard is billed at the standard rate

Use Cases

High-volume classification and routing
Structured data extraction
Short-form generation and summarization
Sub-agent and tool-call steps in agent pipelines
Repeated context with cached inputs (RAG, codebases)