# GPT-6 Sol - Core.Today AI API > The mid tier of OpenAI's GPT-6 generation (in the API since September 14, 2026) — strong reasoning, coding and multimodal performance at $2/$10 per million tokens, a fifth of GPT-6 Astra. 1.05M token context window, 128K max output tokens, cached inputs at a 90% discount. Prompts above 272K input tokens are billed at 2x input / 1.5x output for the whole request. - **Provider**: OpenAI - **Model ID**: gpt-6-sol - **Category**: LLM - **Credits**: 2 per request - **Speed**: Fast - **Quality**: Ultra ## Model Specifications - **Context Window**: 1.1M tokens - **Max Output**: 128K tokens - **Training Cutoff**: 2026-04 - **Supported Formats**: text, json, markdown - **Compatible SDK**: OpenAI ### Capabilities - Vision (image input) - Function Calling - Streaming - JSON Mode - System Prompt ### Token Pricing (per 1M tokens) - **Input Tokens**: 3,716 credits ($2.48) - **Output Tokens**: 18,580 credits ($12.39) - **Cached Tokens**: 371.6 credits ($0.25) ## Features - Mid tier of the GPT-6 generation - 1.05M token context window - 128K max output tokens - Knowledge cutoff: April 20, 2026 - Cached input pricing (90% discount) - Adjustable reasoning effort (low to max) - Function calling and native vision support ## Use Cases - Frontier reasoning and analysis - Complex agentic workflows with tool use - Large-codebase understanding and refactoring - Long document processing up to 1M tokens - Repeated context with cached inputs (RAG, codebases) ## API Endpoint Base URL: https://api.core.today Endpoint: POST /llm/openai/v1/chat/completions ## Authentication Header: Authorization: Bearer YOUR_API_KEY Note: LLM endpoints use OpenAI-compatible format with Authorization Bearer token. ## Input Parameters ### Required - **messages**: array - Array of message objects with role and content - **model**: string (default: gpt-6-sol) - Model identifier ### Optional - **max_completion_tokens**: integer (default: 4096) - Maximum tokens in response (up to 128000). Note: use max_completion_tokens, not max_tokens - **reasoning_effort**: string (default: medium) - Reasoning effort level: low, medium, high, xhigh, or max Options: low, medium, high, xhigh, max - **temperature**: float (default: 1.0) - Sampling temperature (0-2) - **stream**: boolean (default: false) - Enable Server-Sent Events streaming ## Examples ### Frontier Agentic Coding Multi-step code reasoning with GPT-6 Sol ```json { "model": "gpt-6-sol", "messages": [ { "role": "system", "content": "You are a senior software engineer. Think step by step." }, { "role": "user", "content": "Design a migration plan from a monolithic Express API to modular services, then generate the first service's code with tests." } ], "reasoning_effort": "high", "max_completion_tokens": 8000 } ``` ### Cached Repeated Context Reuse a large system prompt with 90% cached input discount ```json { "model": "gpt-6-sol", "messages": [ { "role": "system", "content": "" }, { "role": "user", "content": "Summarize the open TODOs and rank them by risk." } ], "temperature": 0.3, "max_completion_tokens": 4000 } ``` ## Response Format ```json { "id": "chatcmpl-abc123", "object": "chat.completion", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Response text here" }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 100, "completion_tokens": 50, "total_tokens": 150 } } ``` ## Tips - Sol ($2/$10 per M) covers most workloads — escalate to GPT-6 Astra ($10/$50) only for the hardest tasks, drop to GPT-6 Luna ($0.10/$0.50) for simple, high-volume calls - Use cached inputs ($0.20/M, 90% off) for repeated system prompts and RAG context - Keep prompts under 272K input tokens where possible — above that the whole request bills at 2x input / 1.5x output - reasoning_effort 'high', 'xhigh' or 'max' for the most complex tasks - 128K output enables single-shot long-form generation - Lower temperature (0.2-0.5) for coding and analytical tasks - Cache writes bill at 1.25x the input rate — implicit caching writes by default, so most of a large first prompt is billed as a cache write, then reads are 90% off - Fast mode (service_tier: fast or priority) bills token charges at 2x — it is settled on the tier the response reports, so a request downgraded to standard is billed at the standard rate ## Documentation https://platform.openai.com/docs/models