LLM API
OpenAI, Anthropic, Google의 대화형 AI 모델을 동일한 API 키로 사용합니다. 기존 OpenAI SDK와 100% 호환됩니다.
주요 모델 상세 정보
각 모델의 상세한 파라미터, 예제 코드, 활용 팁을 확인하세요.
Claude Haiku 4.5
Anthropic
Fast, cost-effective model for everyday tasks. Great balance of speed, intelligence, and cost for high-volume applications.
Claude Opus 5
Anthropic
Anthropic's latest flagship Opus model, with a 1M-token context window by default and 128K max output tokens. Same pricing as Opus 4.5–4.8 ($5/$25 per M tokens) with prompt caching (read $0.50/M, write $6.25/M) and web search. Compatible with the Anthropic Messages and OpenAI Chat Completions formats.
Claude Opus 4.8
Anthropic
Anthropic's most capable Opus-tier model, with a 1M-token context window (200K on some surfaces), 128K max output tokens, and knowledge cutoff to January 2026. Builds on Opus 4.7 with stronger long-horizon agentic coding, better tool triggering, and adaptive thinking that reasons only when a turn needs it. Compatible with the Anthropic Messages and OpenAI Chat Completions formats.
Claude Opus 4.7
Anthropic
Anthropic's latest flagship model with reliable knowledge cutoff to January 2026 and 128K max output tokens. Builds on Opus 4.6 with improved reasoning, coding, and instruction-following while staying compatible with the Anthropic Messages and OpenAI Chat Completions formats.
Claude Opus 4.6
Anthropic
Anthropic's most capable model. Delivers breakthrough performance in reasoning, coding, and complex analysis with enhanced safety and instruction following.
Claude Opus 4.5
Anthropic
Anthropic's most powerful model for highly complex tasks. Exceptional at research, analysis, and creative projects requiring deep expertise.
Claude Sonnet 5
Anthropic
Anthropic's newest Sonnet model, tuned for the best balance of speed, cost, and intelligence. Features a 1M-token context window, 64K max output tokens, and knowledge cutoff to August 2025. Supports adaptive thinking that reasons only when a turn needs it, vision, and tool use. Compatible with the Anthropic Messages and OpenAI Chat Completions formats.
Claude Sonnet 4.5
Anthropic
Anthropic's most intelligent and capable Sonnet model. Best-in-class for complex reasoning, nuanced understanding, and coding tasks with exceptional instruction following.
Claude Sonnet 4
Anthropic
Balanced Sonnet 4 model offering strong reasoning and coding abilities at an efficient price point. Ideal for everyday production workloads that need a good mix of speed and intelligence.
Gemini 3.5 Flash
Google's latest Flash-tier workhorse model. Flat pricing across all input modalities (text, image, video, audio, PDF) at $1.50/$9 per million tokens with no long-context premium, a 1M token context window, and 65,536 max output tokens. Cached inputs get a 90% discount.
Gemini 3.1 Flash Image Preview
Gemini 3.1 Flash with native image generation capabilities. Can generate images directly in chat responses alongside text. Features separate pricing for text and image output tokens.
Gemini 3.1 Flash Lite
Stable release of the ultra-lightweight Gemini 3.1 Flash Lite. The most cost-effective Gemini model at $0.25/$1.50 per million tokens, with cached input (90% discount), audio input, and batch (50% discount) support. Ideal for high-throughput, budget-conscious applications.
Gemini 3.1 Flash Lite Preview
Ultra-lightweight variant of Gemini 3.1 Flash. The most cost-effective Gemini model with support for cached input and audio input. Ideal for high-throughput, budget-conscious applications.
Gemini 3.1 Flash Live Preview
Gemini 3.1 Flash optimized for real-time interactions and live streaming scenarios. Features low-latency responses with audio input support at dedicated pricing.
Gemini 3.1 Pro Preview
Google's latest and most capable Gemini model in preview. Features dynamic pricing that adjusts based on context length, with enhanced pricing for inputs over 200K tokens.
Gemini 3 Flash
Google's most advanced reasoning model with state-of-the-art multimodal understanding, PhD-level reasoning, and leading coding performance.
Gemini 3 Pro Image Preview
Google's premium image generation model within the Gemini 3 Pro family. Generates high-quality images directly in chat with the highest fidelity among Gemini image models. Image output tokens are priced at 10x text output tokens.
Gemini 3 Pro Preview
Google's most powerful Gemini model in preview. Features breakthrough reasoning, coding, and multimodal capabilities with the largest context window.
Gemini 2.5 Flash
Google's fast and efficient model with built-in thinking capabilities. Great balance of speed, reasoning, and cost for high-volume applications.
Gemini 2.5 Pro
Google's most capable model with state-of-the-art reasoning and 1M token context. Excels at complex coding, math, and multi-document analysis.
Gemini 2.0 Flash
Google's fastest and most capable model. Features a massive 1M token context window, native multimodal support, and real-time capabilities.
Gemini 2.0 Flash Lite
Ultra-lightweight version of Gemini 2.0 Flash optimized for maximum speed and minimal cost. Perfect for high-volume, latency-sensitive applications.
Gemini Embedding 001
Google's text embedding model for generating vector representations. Optimized for semantic search, clustering, and similarity tasks.
GPT-5.6 Luna
OpenAI
The fast, low-cost tier of OpenAI's GPT-5.6 family (GA July 2026). Luna is built for high-throughput workloads at $1/$6 per million tokens, with a 1M token context window, 128K max output tokens, and cached inputs at a 90% discount.
GPT-5.6 Sol
OpenAI
The flagship tier of OpenAI's GPT-5.6 family (GA July 2026). Sol delivers the strongest reasoning, coding, and multimodal performance of the generation with a 1M token context window, 128K max output tokens, and cached inputs at a 90% discount.
GPT-5.6 Terra
OpenAI
The balanced tier of OpenAI's GPT-5.6 family (GA July 2026). Terra offers near-flagship quality at half the price of Sol, with a 1M token context window, 128K max output tokens, and cached inputs at a 90% discount.
GPT-5.5
OpenAI
OpenAI's newest flagship model with a 1.05M token context window and 128K max output tokens. Supports cached inputs at 10× discount and improved reasoning, coding, and multimodal performance over the GPT-5.4 series.
GPT-5.4
OpenAI
OpenAI's newest flagship model with 1M context window and 128K output tokens. Delivers top-tier reasoning across all domains with adjustable reasoning effort levels from none to xhigh.
GPT-5.4 Mini
OpenAI
Fast and cost-efficient variant of GPT-5.4 with 400K context window and 128K output tokens. Excellent balance of performance and affordability for everyday tasks.
GPT-5.4 Nano
OpenAI
Ultra-lightweight and fastest GPT-5.4 variant with 400K context and 128K output. Designed for high-throughput, low-latency applications at minimal cost. Supports MCP for tool integration.
GPT-5.2
OpenAI
OpenAI's latest and most advanced GPT model. Delivers state-of-the-art performance across reasoning, coding, and creative tasks with enhanced capabilities.
GPT-5.1 (2025-11-13)
OpenAI
Dated snapshot of GPT-5.1 for reproducible results. Supports cached input tokens for cost savings on repeated context. Ideal for production deployments requiring model version pinning.
GPT-5
OpenAI
OpenAI's latest flagship model. Delivers exceptional performance across reasoning, coding, and creative tasks with a massive 1M token context window and 32K output tokens. Supports vision, function calling, and JSON mode.
GPT-5 Mini
OpenAI
Fast and efficient variant of GPT-5. Delivers strong performance across reasoning, coding, and creative tasks with a 1M token context window and 32K output tokens, at a fraction of the cost of GPT-5.
GPT-5 Nano
OpenAI
Ultra-fast and lightweight variant of GPT-5. Designed for high-throughput, low-latency applications with a 1M token context window and 32K output tokens at minimal cost.
GPT-4.1
OpenAI
OpenAI's most capable model for coding and instruction following. Features a 1M token context window, 32K output tokens, and major improvements in coding, complex prompts, and long-context tasks. 20% cheaper than GPT-4o on output.
GPT-4.1 Mini
OpenAI
A significant leap in small model performance. Matches or exceeds GPT-4o in intelligence while reducing latency by nearly half and cost by 83%. Ideal balance of speed, quality, and affordability.
GPT-4.1 Nano
OpenAI
OpenAI's fastest and cheapest model. Optimized for classification, autocompletion, and low-latency tasks. Ultra-affordable at $0.10/1M input tokens.
GPT-4o
OpenAI
OpenAI's flagship multimodal model. Industry-leading performance in reasoning, coding, and creative tasks with native vision capabilities and structured output support.
GPT-4o Mini
OpenAI
Cost-effective, fast model with strong performance. Best for high-volume tasks where speed and cost matter more than absolute capability.
GPT Audio Mini
OpenAI
Lightweight multimodal model with native audio input/output capabilities. Optimized for voice-based interactions and audio processing tasks.
MiniMax M2.7
MiniMax
MiniMax's flagship M2-series language model, served through an OpenAI-compatible API. Strong multilingual capability (notably Chinese and English) at a very low price point ($0.30/$1.20 per million tokens) with prompt cache reads at $0.06/M.
OpenAI o4-mini
OpenAI
Fast, cost-effective reasoning model optimized for coding and STEM tasks. Provides strong reasoning at a fraction of the cost of larger reasoning models.
OpenAI o3-mini
OpenAI
Efficient reasoning model that delivers strong performance at lower cost. Ideal for tasks requiring reasoning without the overhead of larger models.
OpenAI o1
OpenAI
OpenAI's most advanced reasoning model. Uses extended thinking time to solve complex problems in science, coding, and math with exceptional accuracy.
가격 단위: 크레딧/토큰 기준입니다. 예: 1,000 토큰 입력, 500 토큰 출력 시 gpt-4o-mini는 0.3 + 0.6 = 0.9 크레딧
GPT-5 / GPT-5.2 / O-Series 주의사항
GPT-5, GPT-5.2, o1, o3 등 추론(Reasoning) 모델은 일반 모델과 파라미터가 다릅니다:
max_tokens→max_completion_tokens사용temperature,top_p지원 안 함- 새 파라미터:
reasoning_effort(minimal/low/medium/high)
OpenAI (GPT)
GPT-4o / GPT-4.1 (일반 모델)
curl -X POST https://api.core.today/llm/openai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer cdt_your_api_key" \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 1000,
"temperature": 0.7
}'GPT-5 / O-Series (추론 모델)
curl -X POST https://api.core.today/llm/openai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer cdt_your_api_key" \
-d '{
"model": "gpt-5",
"messages": [{"role": "user", "content": "Explain quantum computing"}],
"max_completion_tokens": 16000,
"reasoning_effort": "medium"
}'GPT-5 전용 파라미터
max_completion_tokens- 최대 출력 토큰 (max_tokens 대신 사용)reasoning_effort- 추론 수준: minimal, low, medium, high
Codex 모델 (코드 특화)
Codex 모델은 Responses API 전용입니다
gpt-5.1-codex, gpt-5.1-codex-mini 등 Codex 모델은 /v1/chat/completions를 지원하지 않습니다. 대신 /v1/responses 엔드포인트를 사용해야 합니다.
게이트웨이는 경로를 그대로 프록시하므로, 클라이언트에서 엔드포인트 경로만 변경하면 됩니다: /llm/openai/v1/responses → OpenAI /v1/responses
# Codex 모델: /v1/responses 엔드포인트 사용
curl -X POST https://api.core.today/llm/openai/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer cdt_your_api_key" \
-d '{
"model": "gpt-5.1-codex",
"instructions": "You are a helpful coding assistant.",
"input": "Write a Python function to merge two sorted lists",
"max_output_tokens": 16000
}'messages→input(문자열 또는 메시지 배열)- 시스템 프롬프트:
instructions파라미터 사용 - 출력 토큰 제한:
max_output_tokens사용 - 스트리밍 시 이벤트 형식:
response.output_text.delta
gpt-5.1-codex, gpt-5.1-codex-mini, gpt-5.2-codex 등 모델 이름에 "codex"가 포함된 모델| 모델 | Input | Output |
|---|---|---|
| 모델 | Input | Output |
|---|---|---|
| 모델 | Input | Output |
|---|---|---|
| 모델 | Input | Output |
|---|---|---|
Anthropic (Claude)
curl -X POST https://api.core.today/llm/anthropic/v1/messages \
-H "Content-Type: application/json" \
-H "Authorization: Bearer cdt_your_api_key" \
-d '{
"model": "claude-sonnet-4",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "Explain quantum computing simply."}
]
}'Claude Opus 4.7 주의사항
claude-opus-4-7은 temperature 파라미터가 deprecated 되었습니다. 요청 바디에 포함하면 Anthropic이 400: `temperature` is deprecated for this model 으로 거절합니다 — 이 모델 호출 시 temperature 필드를 제거하세요.
| 모델 | Input | Output |
|---|---|---|
| 모델 | Input | Output |
|---|---|---|
| 모델 | Input | Output |
|---|---|---|
Google (Gemini)
curl -X POST "https://api.core.today/llm/gemini/v1beta/models/gemini-2.5-pro:generateContent" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer cdt_your_api_key" \
-d '{
"contents": [
{
"parts": [{"text": "Write a haiku about programming"}]
}
]
}'| 모델 | Input | Output |
|---|---|---|
gemini-embedding-001 (임베딩 전용, Input 0.0003) | gemini-3-pro-preview-longcontext (Input 0.0080, Output 0.0360)스트리밍 응답
실시간으로 응답을 받으려면 stream: true를 추가하세요:
curl -X POST https://api.core.today/llm/openai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer cdt_your_api_key" \
-d '{
"model": "gpt-5",
"messages": [{"role": "user", "content": "Tell me a long story"}],
"stream": true
}'프롬프트 캐싱 (Claude 자동 비용 절감)
Claude 모델을 호출할 때 매번 똑같이 반복되는 앞부분(시스템 프롬프트·대화 히스토리)을 게이트웨이가 자동으로 캐시합니다. 두 번째 요청부터 그 부분은 다시 계산되지 않고 캐시에서 읽혀, 입력 토큰가의 10% 가격으로 청구됩니다. 별도 설정이나 코드 변경은 필요 없습니다.
| 캐시되는 부분 | 조건 |
|---|---|
시스템 프롬프트 (system) | 항상 — 같은 시스템 프롬프트를 쓰는 모든 요청 |
| 대화 히스토리 (직전 답변까지) | 멀티턴 대화일 때 — 이전 assistant 답변이 포함된 요청 |
| 새로 보낸 마지막 질문 | 캐시하지 않음 — 매번 바뀌므로 정상 가격 |
같은 시스템 프롬프트를 재사용하거나 대화를 계속 이어가면 캐시가 잘 적중합니다:
1번째 요청: [긴 시스템 프롬프트] + "질문 A" → 시스템 프롬프트를 캐시에 저장
2번째 요청: [긴 시스템 프롬프트] + "질문 B" → 캐시에서 읽음 (10% 가격)
3번째 요청: [긴 시스템 프롬프트] + "질문 C" → 캐시에서 읽음usage.cache_read_input_tokens가 0보다 크면 캐시가 적중한 것입니다. (cache_creation_input_tokens는 캐시에 새로 저장된 토큰)- Claude(Anthropic) 요청에만 적용됩니다. OpenAI는 서버가 자동 캐싱, Gemini는 별도 방식입니다.
- 앞부분이 최소 캐시 길이(대략 1,024토큰)에 못 미치면 캐시가 적용되지 않습니다 — 이때는 추가 비용도 없습니다.
- 캐시는 마지막 사용 후 약 5분간 유지되며, 그 안에 다시 쓰이면 자동 연장됩니다.
- 요청에 직접
cache_control을 넣은 경우, 그 설정을 그대로 사용합니다(게이트웨이가 손대지 않음).
비용 계산 예시
1,000 토큰 입력, 500 토큰 출력 기준:
| 모델 | 계산 | 총 비용 |
|---|---|---|
| gpt-4o-mini | 0.3 + 0.6 | 0.9 크레딧 |
| gpt-5 | 2.5 + 10.0 | 12.5 크레딧 |
| claude-3-haiku | 0.5 + 1.25 | 1.75 크레딧 |
| claude-sonnet-4 | 6.0 + 15.0 | 21.0 크레딧 |
| gemini-2.0-flash | 0.2 + 0.4 | 0.6 크레딧 |