# GPT-5.4 Pro - Core.Today AI API > Maximum-capability tier of the GPT-5.4 family (dated snapshot). Uses more compute to think longer and deliver the most reliable answers on the hardest reasoning, coding and analysis tasks. Served only via the OpenAI Responses API. $30/$180 per million tokens (no cached-input discount); prompts above 272K input tokens are billed at 2x input / 1.5x output for the whole request ($60/$270 per million tokens). - **Provider**: OpenAI - **Model ID**: gpt-5.4-pro-2026-03-05 - **Category**: LLM - **Credits**: 195 per 1K tokens (avg) - **Speed**: Slow - **Quality**: Ultra ## Model Specifications - **Context Window**: 1M tokens - **Max Output**: 128K tokens - **Training Cutoff**: 2025-08 - **Supported Formats**: text, json, markdown - **Compatible SDK**: OpenAI ### Capabilities - Vision (image input) - Function Calling - Streaming - JSON Mode - System Prompt ### Token Pricing (per 1M tokens) - **Input Tokens**: 55,740 credits ($37.16) - **Output Tokens**: 334,440 credits ($223) ## Features - Maximum-capability GPT-5.4 tier - 1M token context window - 128K max output tokens - Extended thinking for hardest problems - Responses API only (no chat completions) - Native vision and function calling ## Use Cases - Hardest multi-step reasoning problems - High-stakes code review and architecture - Scientific and mathematical analysis - Long-document synthesis at maximum quality ## API Endpoint Base URL: https://api.core.today Endpoint: POST /llm/openai/v1/responses ## Authentication Header: Authorization: Bearer YOUR_API_KEY Note: LLM endpoints use OpenAI-compatible format with Authorization Bearer token. ## Input Parameters ### Required - **model**: string (default: gpt-5.4-pro-2026-03-05) - Model identifier (dated snapshot id) - **input**: string | array - Responses API input: a plain string or an array of input items (messages, tool results) ### Optional - **instructions**: string - System-level instructions for the run - **max_output_tokens**: integer - Maximum tokens in the response, including internal reasoning tokens (Responses API field; not max_tokens) - **reasoning**: object - Reasoning options, e.g. { effort: "high" } - **stream**: boolean (default: false) - Enable Server-Sent Events streaming ## Examples ### Hardest Reasoning Maximum-effort analysis with GPT-5.4 Pro via the Responses API ```json { "model": "gpt-5.4-pro-2026-03-05", "input": "Review this database migration plan for a zero-downtime rollout across 40 services and identify every failure mode with mitigations.", "max_output_tokens": 16000 } ``` ## Response Format ```json { "id": "chatcmpl-abc123", "object": "chat.completion", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Response text here" }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 100, "completion_tokens": 50, "total_tokens": 150 } } ``` ## Tips - Works with the OpenAI SDK - set base_url to https://ai.api.core.today/llm/openai/v1 and use client.responses.create() - Dated snapshot id - pin this exact id for reproducible behavior (the alias gpt-5.4-pro resolves to it) - Served only via the Responses API (POST /llm/openai/v1/responses) — chat completions is not supported for this model - Reserve for the hardest problems - standard GPT-5.4 ($2.50/$15) is far cheaper for routine tasks - Responses take longer due to extended internal thinking - budget generous max_output_tokens, reasoning tokens count toward the limit - Keep prompts under 272K input tokens where possible — above that the whole request bills at 2x input / 1.5x output ($60/$270 per 1M) ## Documentation https://platform.openai.com/docs/models