Skip to main content
Core.Today
Model APIsOpenAI Compatible

LLM API

OpenAI, Anthropic, Google, xAI Grok, DeepSeek, Meta와 OpenRouter 경유 오픈 모델(Qwen, Kimi, GLM, Llama 등)을 동일한 API 키로 사용합니다. 기존 OpenAI SDK와 100% 호환됩니다.

Give this page to your AI — all model specs as LLM-friendly text
llms.txt ↗

주요 모델 상세 정보

각 모델의 상세한 파라미터, 예제 코드, 활용 팁을 확인하세요.

Claude Haiku 4.5

Anthropic

1 credits/1K

Fast, cost-effective model for everyday tasks. Great balance of speed, intelligence, and cost for high-volume applications.

빠름Ultra
상세 보기

Claude Opus 5

Anthropic

5 credits/1K

Anthropic's latest flagship Opus model, with a 1M-token context window by default and 128K max output tokens. Same pricing as Opus 4.5–4.8 ($5/$25 per M tokens) with prompt caching (read $0.50/M, write $6.25/M) and web search. Compatible with the Anthropic Messages and OpenAI Chat Completions formats.

빠름Ultra
상세 보기

Claude Opus 4.8

Anthropic

5 credits/1K

Anthropic's most capable Opus-tier model, with a 1M-token context window (200K on some surfaces), 128K max output tokens, and knowledge cutoff to January 2026. Builds on Opus 4.7 with stronger long-horizon agentic coding, better tool triggering, and adaptive thinking that reasons only when a turn needs it. Compatible with the Anthropic Messages and OpenAI Chat Completions formats.

빠름Ultra
상세 보기

Claude Opus 4.7

Anthropic

5 credits/1K

Anthropic's latest flagship model with reliable knowledge cutoff to January 2026 and 128K max output tokens. Builds on Opus 4.6 with improved reasoning, coding, and instruction-following while staying compatible with the Anthropic Messages and OpenAI Chat Completions formats.

빠름Ultra
상세 보기

Claude Opus 4.6

Anthropic

5 credits/1K

Anthropic's most capable model. Delivers breakthrough performance in reasoning, coding, and complex analysis with enhanced safety and instruction following.

빠름Ultra
상세 보기

Claude Opus 4.5

Anthropic

5 credits/1K

Anthropic's most powerful model for highly complex tasks. Exceptional at research, analysis, and creative projects requiring deep expertise.

빠름Ultra
상세 보기

Claude Sonnet 5

Anthropic

4 credits/1K

Anthropic's newest Sonnet model, tuned for the best balance of speed, cost, and intelligence. Features a 1M-token context window, 64K max output tokens, and knowledge cutoff to August 2025. Supports adaptive thinking that reasons only when a turn needs it, vision, and tool use. Compatible with the Anthropic Messages and OpenAI Chat Completions formats.

빠름Ultra
상세 보기

Claude Sonnet 4.6

Anthropic

17 credits/1K

The successor to Claude Sonnet 4.5. Anthropic's frontier Sonnet-tier model for complex reasoning, nuanced understanding, and coding tasks with exceptional instruction following, at the same token pricing as Sonnet 4.5.

빠름Ultra
상세 보기

Claude Sonnet 4.5

Anthropic

4 credits/1K

Anthropic's most intelligent and capable Sonnet model. Best-in-class for complex reasoning, nuanced understanding, and coding tasks with exceptional instruction following.

빠름Ultra
상세 보기

DeepSeek V4 Flash

DeepSeek

0.4 credits/1K

DeepSeek's fast, low-cost V4 model (checkpoint V4-Flash-0731). Thinking mode is on by default, with a 1M-token context window and heavily discounted cache reads. Replaces the retired deepseek-chat / deepseek-reasoner.

빠름Ultra
상세 보기

DeepSeek V4 Pro

DeepSeek

1.2 credits/1K

DeepSeek's strongest V4 model (checkpoint V4-Pro-0813). Advanced reasoning with thinking mode on by default, a 1M-token context window, and discounted cache reads - still far cheaper than comparable frontier models.

빠름Ultra
상세 보기

Gemini 3.5 Flash

Google

2 credits/1K

Google's latest Flash-tier workhorse model. Flat pricing across all input modalities (text, image, video, audio, PDF) at $1.50/$9 per million tokens with no long-context premium, a 1M token context window, and 65,536 max output tokens. Cached inputs get a 90% discount.

빠름Ultra
상세 보기

Gemini 3.1 Flash Image Preview

Google

500 credits/1K

Gemini 3.1 Flash with native image generation capabilities. Can generate images directly in chat responses alongside text. Features separate pricing for text and image output tokens.

빠름Ultra
상세 보기

Gemini 3.1 Flash Lite

Google

1 credits/1K

Stable release of the ultra-lightweight Gemini 3.1 Flash Lite. The most cost-effective Gemini model at $0.25/$1.50 per million tokens, with cached input (90% discount), audio input, and batch (50% discount) support. Ideal for high-throughput, budget-conscious applications.

빠름Ultra
상세 보기

Gemini 3.1 Flash Lite Preview

Google

100 credits/1K

Ultra-lightweight variant of Gemini 3.1 Flash. The most cost-effective Gemini model with support for cached input and audio input. Ideal for high-throughput, budget-conscious applications.

빠름Ultra
상세 보기

Gemini 3.1 Pro Preview

Google

500 credits/1K

Google's latest and most capable Gemini model in preview. Features dynamic pricing that adjusts based on context length, with enhanced pricing for inputs over 200K tokens.

빠름Ultra
상세 보기

Gemini 3 Flash Preview

Google

500 credits/1K

Google's most advanced reasoning model with state-of-the-art multimodal understanding, PhD-level reasoning, and leading coding performance.

빠름Ultra
상세 보기

Gemini 3 Pro Image Preview

Google

500 credits/1K

Google's premium image generation model within the Gemini 3 Pro family. Generates high-quality images directly in chat with the highest fidelity among Gemini image models. Image output tokens are priced at 10x text output tokens.

빠름Ultra
상세 보기

Gemini 3 Pro Preview

Google

4 credits/1K

Google's most powerful Gemini model in preview. Features breakthrough reasoning, coding, and multimodal capabilities with the largest context window.

빠름Ultra
상세 보기

Gemini 2.5 Flash

Google

1 credits/1K

Google's fast and efficient model with built-in thinking capabilities. Great balance of speed, reasoning, and cost for high-volume applications.

빠름Ultra
상세 보기

Gemini 2.5 Flash Lite

Google

0.5 credits/1K

The cheapest tier of the Gemini 2.5 family, optimized for high-volume, latency-sensitive workloads. Delivers 2.5-generation quality at a fraction of the cost, ideal for classification, extraction, and real-time chat at scale.

빠름Ultra
상세 보기

Gemini 2.5 Pro

Google

3 credits/1K

Google's most capable model with state-of-the-art reasoning and 1M token context. Excels at complex coding, math, and multi-document analysis.

빠름Ultra
상세 보기

Gemini 2.0 Flash

Google

1 credits/1K

Google's fastest and most capable model. Features a massive 1M token context window, native multimodal support, and real-time capabilities.

빠름Ultra
상세 보기

Gemini 2.0 Flash Lite

Google

0.5 credits/1K

Ultra-lightweight version of Gemini 2.0 Flash optimized for maximum speed and minimal cost. Perfect for high-volume, latency-sensitive applications.

빠름Ultra
상세 보기

Gemini Embedding 001

Google

0.1 credits/1K

Google's text embedding model for generating vector representations. Optimized for semantic search, clustering, and similarity tasks.

빠름Ultra
상세 보기

GLM 5.2

Z.ai

1.8 credits/1K

Z.ai's GLM 5.2 served through the OpenRouter aggregator - a large-scale text reasoning model with a 1M-token context window, suited for long-horizon agent workflows and project-level software engineering. Billed at OpenRouter's actual usage cost.

빠름Ultra
상세 보기

GLM 4.7

Z.ai

2 credits/1K

Z.ai's GLM 4.7 served through the OpenRouter aggregator - a flagship text model with enhanced programming capabilities and more stable multi-step reasoning/execution, showing clear gains on complex agent tasks. 200K context, billed at OpenRouter's actual usage cost.

빠름Ultra
상세 보기

GPT-6 Astra

OpenAI

5 credits/1K

OpenAI's most capable model (GPT-6 generation, September 2026), built for the hardest end-to-end work: complex reasoning, coding, computer use, research and document creation. 1M token context window, 128K max output tokens, cached inputs at a 90% discount. Prompts above 272K input tokens are billed at 2x input / 1.5x output for the whole request.

빠름Ultra
상세 보기

GPT-5.6 Luna

OpenAI

1 credits/1K

The fast, low-cost tier of OpenAI's GPT-5.6 family (GA July 2026). Luna is built for high-throughput workloads at $0.20/$1.20 per million tokens, with a 1M token context window, 128K max output tokens, and cached inputs at a 90% discount. Prompts above 272K input tokens are billed at 2x input / 1.5x output for the whole request.

빠름Ultra
상세 보기

GPT-5.6 Sol

OpenAI

5 credits/1K

The flagship tier of OpenAI's GPT-5.6 family (GA July 2026). Sol delivers the strongest reasoning, coding, and multimodal performance of the generation with a 1M token context window, 128K max output tokens, and cached inputs at a 90% discount. Prompts above 272K input tokens are billed at 2x input / 1.5x output for the whole request.

빠름Ultra
상세 보기

GPT-5.6 Terra

OpenAI

3 credits/1K

The balanced tier of OpenAI's GPT-5.6 family (GA July 2026). Terra offers near-flagship quality at $2/$12 per million tokens (40% of Sol), with a 1M token context window, 128K max output tokens, and cached inputs at a 90% discount. Prompts above 272K input tokens are billed at 2x input / 1.5x output for the whole request.

빠름Ultra
상세 보기

GPT-5.5

OpenAI

5 credits/1K

OpenAI's newest flagship model with a 1.05M token context window and 128K max output tokens. Supports cached inputs at 10× discount and improved reasoning, coding, and multimodal performance over the GPT-5.4 series.

빠름Ultra
상세 보기

GPT-5.4

OpenAI

5 credits/1K

OpenAI's newest flagship model with 1M context window and 128K output tokens. Delivers top-tier reasoning across all domains with adjustable reasoning effort levels from none to xhigh.

빠름Ultra
상세 보기

GPT-5.4 Mini

OpenAI

2 credits/1K

Fast and cost-efficient variant of GPT-5.4 with 400K context window and 128K output tokens. Excellent balance of performance and affordability for everyday tasks.

빠름Ultra
상세 보기

GPT-5.4 Nano

OpenAI

1 credits/1K

Ultra-lightweight and fastest GPT-5.4 variant with 400K context and 128K output. Designed for high-throughput, low-latency applications at minimal cost. Supports MCP for tool integration.

빠름Ultra
상세 보기

GPT-5.2 Pro

OpenAI

176 credits/1K

Maximum-capability tier of the GPT-5.2 family (dated snapshot). Thinks longer with more compute to deliver the most reliable answers on the hardest reasoning and coding tasks.

빠름Ultra
상세 보기

GPT-5.2

OpenAI

4 credits/1K

OpenAI's latest and most advanced GPT model. Delivers state-of-the-art performance across reasoning, coding, and creative tasks with enhanced capabilities.

빠름Ultra
상세 보기

GPT-5.2 Codex

OpenAI

15 credits/1K

Latest Codex generation built on the GPT-5.2 base. Brings GPT-5.2's stronger reasoning to agentic coding harnesses, served via the OpenAI Responses API.

빠름Ultra
상세 보기

GPT-5.1 (2025-11-13)

OpenAI

3 credits/1K

Dated snapshot of GPT-5.1 for reproducible results. Supports cached input tokens for cost savings on repeated context. Ideal for production deployments requiring model version pinning.

빠름Ultra
상세 보기

GPT-5.1 Codex

OpenAI

10 credits/1K

GPT-5.1 variant optimized for agentic coding in Codex-style harnesses. Tuned for long tool-use loops, code editing, and repository-scale tasks, and served via the OpenAI Responses API.

빠름Ultra
상세 보기

GPT-5.1 Codex Max

OpenAI

10 credits/1K

OpenAI's long-horizon agentic coding flagship in the GPT-5.1 Codex family. Supports compaction to keep very long coding sessions within context, and is served via the OpenAI Responses API.

빠름Ultra
상세 보기

GPT-5.1 Codex Mini

OpenAI

2 credits/1K

Cheaper, faster member of the GPT-5.1 Codex family for lighter coding tasks. Keeps the agentic-coding tuning and Responses API interface at roughly one fifth of the Codex price.

빠름Ultra
상세 보기

GPT-5

OpenAI

3 credits/1K

OpenAI's latest flagship model. Delivers exceptional performance across reasoning, coding, and creative tasks with a massive 1M token context window and 32K output tokens. Supports vision, function calling, and JSON mode.

빠름Ultra
상세 보기

GPT-5 Mini

OpenAI

1 credits/1K

Fast and efficient variant of GPT-5. Delivers strong performance across reasoning, coding, and creative tasks with a 1M token context window and 32K output tokens, at a fraction of the cost of GPT-5.

빠름Ultra
상세 보기

GPT-5 Nano

OpenAI

1 credits/1K

Ultra-fast and lightweight variant of GPT-5. Designed for high-throughput, low-latency applications with a 1M token context window and 32K output tokens at minimal cost.

빠름Ultra
상세 보기

GPT-5 Pro

OpenAI

125 credits/1K

OpenAI's maximum-capability GPT-5 tier. Uses more compute to think longer and deliver the most reliable answers on the hardest reasoning, coding, and analysis tasks.

빠름Ultra
상세 보기

GPT-4.1

OpenAI

3 credits/1K

OpenAI's most capable model for coding and instruction following. Features a 1M token context window, 32K output tokens, and major improvements in coding, complex prompts, and long-context tasks. 20% cheaper than GPT-4o on output.

빠름Ultra
상세 보기

GPT-4.1 Mini

OpenAI

1 credits/1K

A significant leap in small model performance. Matches or exceeds GPT-4o in intelligence while reducing latency by nearly half and cost by 83%. Ideal balance of speed, quality, and affordability.

빠름Ultra
상세 보기

GPT-4.1 Nano

OpenAI

1 credits/1K

OpenAI's fastest and cheapest model. Optimized for classification, autocompletion, and low-latency tasks. Ultra-affordable at $0.10/1M input tokens.

빠름Ultra
상세 보기

GPT-4o

OpenAI

3 credits/1K

OpenAI's flagship multimodal model. Industry-leading performance in reasoning, coding, and creative tasks with native vision capabilities and structured output support.

빠름Ultra
상세 보기

GPT-4o Mini

OpenAI

1 credits/1K

Cost-effective, fast model with strong performance. Best for high-volume tasks where speed and cost matter more than absolute capability.

빠름Ultra
상세 보기

GPT Audio

OpenAI

12 credits/1K

Full-size multimodal model with native audio input/output via chat completions. Handles speech in and speech out with strong general intelligence behind it.

빠름Ultra
상세 보기

GPT Audio Mini

OpenAI

1 credits/1K

Lightweight multimodal model with native audio input/output capabilities. Optimized for voice-based interactions and audio processing tasks.

빠름Ultra
상세 보기

GPT-OSS 120B

OpenAI

0.2 credits/1K

OpenAI's open-weight gpt-oss-120b, a 117B-parameter Mixture-of-Experts model (5.1B active per forward pass) built for high-reasoning, agentic and general-purpose production use - served through OpenRouter only; it is not available on the direct OpenAI API. 131K-token context, billed at OpenRouter's actual usage cost.

빠름Ultra
상세 보기

Grok 4.6

xAI

7 credits/1K

xAI's flagship Grok model. Strongest reasoning and general capability in the Grok family, with a 500K-token context window.

빠름Ultra
상세 보기

Grok 4.5

xAI

7 credits/1K

xAI's coding-focused Grok model with a 500K-token context window - same token rates as Grok 4.6 with cheaper cache reads.

빠름Ultra
상세 보기

Grok 4.3

xAI

3 credits/1K

Cost-efficient Grok model with a 1M-token context window - under half the price of the 4.5/4.6 tier.

빠름Ultra
상세 보기

Grok 4.3 (OpenRouter)

xAI

3 credits/1K

xAI's Grok 4.3 served through the OpenRouter aggregator - 1M-token context, vision input, billed at OpenRouter's actual usage cost. Prefer the direct grok-4.3 route unless you need the OpenRouter path.

빠름Ultra
상세 보기

Grok 4.20 Multi-Agent

xAI

3 credits/1K

Grok 4.20 Multi-Agent beta (0309 snapshot) - deep-research style multi-agent orchestration with a 1M-token context window. Served through the Responses API only.

빠름Ultra
상세 보기

Grok 4.20 Non-Reasoning

xAI

3 credits/1K

Grok 4.20 non-reasoning variant (0309 snapshot) - fast responses with a 1M-token context window at the same rates as the reasoning variant.

빠름Ultra
상세 보기

Grok 4.20 Reasoning

xAI

3 credits/1K

Grok 4.20 reasoning variant (0309 snapshot) - chain-of-thought reasoning with a 1M-token context window. Reasoning tokens are billed additively on top of completion tokens.

빠름Ultra
상세 보기

Grok Build 0.1

xAI

3 credits/1K

xAI's agentic-coding Grok model with a 256K-token context window - the lowest rates in the Grok family.

빠름Ultra
상세 보기

Kimi K3

Moonshot AI

17 credits/1K

Moonshot AI's Kimi K3 served through the OpenRouter aggregator - a 2.8T-parameter open-weight multimodal reasoning model with a 1M-token context and image/video input, built for complex coding, knowledge work, and long-horizon agentic workflows. Billed at OpenRouter's actual usage cost.

빠름Ultra
상세 보기

Kimi K2.7 Code

Moonshot AI

4 credits/1K

Moonshot AI's Kimi K2.7 Code served through the OpenRouter aggregator - a coding-focused member of the Kimi K2 family built to complete end-to-end programming tasks reliably over long contexts, with a native multimodal mixture-of-experts architecture, 256K context, and image input. Billed at OpenRouter's actual usage cost.

빠름Ultra
상세 보기

Llama 3.3 70B Instruct

Meta

0.4 credits/1K

Meta's Llama 3.3 70B instruction-tuned multilingual model served through OpenRouter - text in/text out with 131K-token context, a strong open-weight general-purpose chat model. This id replaces the retired direct llama-3.3-70b-versatile route. Billed at OpenRouter's actual usage cost.

빠름Ultra
상세 보기

Llama 3.1 8B Instruct

Meta

0.1 credits/1K

Meta's Llama 3.1 8B instruction-tuned model served through OpenRouter - a fast, efficient small open-weight model with 131K-token context for high-volume, low-cost workloads. This id replaces the retired direct llama-3.1-8b-instant route. Billed at OpenRouter's actual usage cost.

빠름Ultra
상세 보기

MiniMax M3

MiniMax

1.4 credits/1K

MiniMax's M3 multimodal foundation model served through OpenRouter - text, image and video input, 1M-token context with up to 512K output tokens, suited for long-horizon agentic work and coding. Billed at OpenRouter's actual usage cost.

빠름Ultra
상세 보기

MiniMax M2.7 (OpenRouter)

MiniMax

1.4 credits/1K

MiniMax's M2.7 served through OpenRouter - an agentic large language model built for autonomous, real-world productivity with 204K-token context and multi-agent tool use. This id replaces the retired direct MiniMax-M2.7 route. Billed at OpenRouter's actual usage cost.

빠름Ultra
상세 보기

MiniMax M2.7

MiniMax

1 credits/1K

MiniMax's flagship M2-series language model, served through an OpenAI-compatible API. Strong multilingual capability (notably Chinese and English) at a very low price point ($0.30/$1.20 per million tokens) with prompt cache reads at $0.06/M.

빠름Ultra
상세 보기

Mistral Large 3 (2512)

Mistral AI

1.9 credits/1K

Mistral Large 3 (2512) served through the OpenRouter aggregator - Mistral's most capable model to date, a sparse mixture-of-experts with 41B active parameters (675B total) released under Apache 2.0, with 256K context and image/file input. Billed at OpenRouter's actual usage cost.

빠름Ultra
상세 보기

Muse Spark 1.3

Meta

5 credits/1K

Meta's Muse Spark 1.3 over the direct Meta API - a multimodal reasoning model for agentic work that accepts text, images, video, audio and PDF documents and returns text, with a 1M-token context window and cached-input pricing. Answers are generated after reasoning tokens, so keep max_tokens at 2,048 or more.

빠름Ultra
상세 보기

Muse Spark 1.3 Contributor

Meta

0.3 credits/1K

The discounted Contributor tier of Meta's Muse Spark 1.3 - same model, about 1/12 of the standard price, on the condition that prompts and responses may be used by Meta to train its models. Do not send personal, confidential or customer data. Shares an upstream limit of 100 requests / 3M tokens per minute.

빠름Ultra
상세 보기

Muse Spark 1.1

Meta

5 credits/1K

Meta's Muse Spark 1.1 served through the OpenRouter aggregator - a multimodal reasoning model built for agentic tasks that accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context. Billed at OpenRouter's actual usage cost.

빠름Ultra
상세 보기

OpenAI o4-mini

OpenAI

2 credits/1K

Fast, cost-effective reasoning model optimized for coding and STEM tasks. Provides strong reasoning at a fraction of the cost of larger reasoning models.

빠름Ultra
상세 보기

o3

OpenAI

9 credits/1K

OpenAI's frontier reasoning model. Uses extended chain-of-thought to solve complex problems in science, coding, and math with high accuracy at a competitive price.

빠름Ultra
상세 보기

o3 Pro

OpenAI

93 credits/1K

Highest-effort version of o3. Thinks longer with more compute to deliver the most reliable answers on the hardest science, math, and coding problems.

빠름Ultra
상세 보기

OpenAI o3-mini

OpenAI

2 credits/1K

Efficient reasoning model that delivers strong performance at lower cost. Ideal for tasks requiring reasoning without the overhead of larger models.

빠름Ultra
상세 보기

o1 Pro

OpenAI

697 credits/1K

Legacy pro-tier reasoning model. Uses more compute than o1 for the most reliable answers on hard problems. Predecessor to o3-pro, which offers better value.

빠름Ultra
상세 보기

OpenAI o1

OpenAI

15 credits/1K

OpenAI's most advanced reasoning model. Uses extended thinking time to solve complex problems in science, coding, and math with exceptional accuracy.

빠름Ultra
상세 보기

Qwen3.8 27B

Alibaba Qwen

3 credits/1K

Qwen3.8 27B served through the OpenRouter aggregator - an open-weight dense vision-language model for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking mode and a 256K-token context. Billed at OpenRouter's actual usage cost.

빠름Ultra
상세 보기

Qwen3.8 2.4T-A95B

Alibaba Qwen

7 credits/1K

Qwen3.8 2.4T-A95B served through the OpenRouter aggregator - the open-weight sparse mixture-of-experts variant of Qwen3.8 Max, with 95B active parameters out of 2.4T total. Text-only reasoning model with a 1M-token context and up to 262K output tokens, billed at OpenRouter's actual usage cost.

빠름Ultra
상세 보기

Qwen3.8 Max

Alibaba Qwen

7 credits/1K

Alibaba's Qwen3.8 Max served through the OpenRouter aggregator - the flagship of the Qwen3.8 series and GA successor to Qwen3.8 Max Preview. A multimodal reasoning model for complex reasoning, visual understanding, and agentic work with a 1M-token context, billed at OpenRouter's actual usage cost.

빠름Ultra
상세 보기

Qwen3.7 Flash

Alibaba Qwen

0.1 credits/1K

Alibaba's Qwen3.7 Flash served through the OpenRouter aggregator - a low-cost vision-language reasoning model for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition and spatial understanding. 1M-token context, billed at OpenRouter's actual usage cost.

빠름Ultra
상세 보기

Qwen3 Max

Alibaba Qwen

4 credits/1K

Alibaba's Qwen3 Max served through the OpenRouter aggregator - an updated release built on the Qwen3 series with major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage. Text-only with a 256K-token context, billed at OpenRouter's actual usage cost.

빠름Ultra
상세 보기

Solar Pro 4

Upstage

0.1 credits/1K

Upstage's Solar Pro 4 served through the OpenRouter aggregator - a cost-efficient LLM with a 524K-token context window, built for long-horizon tasks and agentic workflows with strengths in office productivity and document-intensive work. Billed at OpenRouter's actual usage cost.

빠름Ultra
상세 보기

Text Embedding 3 Large

OpenAI

0.24 credits/1K

OpenAI's most capable text embedding model. Generates high-quality vector representations for semantic search, clustering, and similarity tasks where retrieval accuracy matters most.

빠름Ultra
상세 보기

Text Embedding 3 Small

OpenAI

0.04 credits/1K

OpenAI's cost-efficient text embedding model. Generates vector representations for semantic search, clustering, and similarity tasks at very low cost.

빠름Ultra
상세 보기

가격 단위: 크레딧/토큰 기준입니다. 예: 1,000 토큰 입력, 500 토큰 출력 시 gpt-4o-mini는 0.3 + 0.6 = 0.9 크레딧

GPT-5 / GPT-5.2 / O-Series 주의사항

GPT-5, GPT-5.2, o1, o3 등 추론(Reasoning) 모델은 일반 모델과 파라미터가 다릅니다:

  • max_tokens max_completion_tokens 사용
  • temperature, top_p 지원 안 함
  • 새 파라미터: reasoning_effort (minimal/low/medium/high)
G

OpenAI (GPT)

GPT-4o / GPT-4.1 (일반 모델)

curl -X POST https://api.core.today/llm/openai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Hello!"}],
    "max_tokens": 1000,
    "temperature": 0.7
  }'

GPT-5 / O-Series (추론 모델)

curl -X POST https://api.core.today/llm/openai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
    "model": "gpt-5",
    "messages": [{"role": "user", "content": "Explain quantum computing"}],
    "max_completion_tokens": 16000,
    "reasoning_effort": "medium"
  }'

GPT-5 전용 파라미터

  • max_completion_tokens - 최대 출력 토큰 (max_tokens 대신 사용)
  • reasoning_effort - 추론 수준: minimal, low, medium, high

Codex 모델 (코드 특화)

Codex 모델은 Responses API 전용입니다

gpt-5.1-codex, gpt-5.1-codex-mini 등 Codex 모델은 /v1/chat/completions를 지원하지 않습니다. 대신 /v1/responses 엔드포인트를 사용해야 합니다.

게이트웨이는 경로를 그대로 프록시하므로, 클라이언트에서 엔드포인트 경로만 변경하면 됩니다: /llm/openai/v1/responses → OpenAI /v1/responses

# Codex 모델: /v1/responses 엔드포인트 사용
curl -X POST https://api.core.today/llm/openai/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
    "model": "gpt-5.1-codex",
    "instructions": "You are a helpful coding assistant.",
    "input": "Write a Python function to merge two sorted lists",
    "max_output_tokens": 16000
  }'
Responses API 주요 차이점:
  • messages input (문자열 또는 메시지 배열)
  • 시스템 프롬프트: instructions 파라미터 사용
  • 출력 토큰 제한: max_output_tokens 사용
  • 스트리밍 시 이벤트 형식: response.output_text.delta
대상 모델: gpt-5.1-codex, gpt-5.1-codex-mini, gpt-5.2-codex 등 모델 이름에 "codex"가 포함된 모델
모델InputOutput
모델InputOutput
모델InputOutput
모델InputOutput
C

Anthropic (Claude)

curl -X POST https://api.core.today/llm/anthropic/v1/messages \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
    "model": "claude-sonnet-4",
    "max_tokens": 1024,
    "messages": [
      {"role": "user", "content": "Explain quantum computing simply."}
    ]
  }'

Claude Opus 4.7 주의사항

claude-opus-4-7 temperature 파라미터가 deprecated 되었습니다. 요청 바디에 포함하면 Anthropic이 400: `temperature` is deprecated for this model 으로 거절합니다 — 이 모델 호출 시 temperature 필드를 제거하세요.

모델InputOutput
모델InputOutput
모델InputOutput
G

Google (Gemini)

curl -X POST "https://api.core.today/llm/gemini/v1beta/models/gemini-2.5-pro:generateContent" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
    "contents": [
      {
        "parts": [{"text": "Write a haiku about programming"}]
      }
    ]
  }'
모델InputOutput
참고: gemini-2.5-pro는 200,000 토큰 초과 시 longcontext 가격 적용 (Input: 0.0050, Output: 0.0300)
기타: gemini-embedding-001 (임베딩 전용, Input 0.0003) | gemini-3-pro-preview-longcontext (Input 0.0080, Output 0.0360)
+

추가 프로바이더 (OpenAI 호환) — xAI Grok · DeepSeek · Meta · OpenRouter

요청·응답 형식은 OpenAI chat completions와 동일하고 게이트웨이 경로 프리픽스만 다릅니다. OpenAI SDK의 base_url을 아래 주소로 바꾸면 그대로 동작합니다.

프로바이더base_url모델 ID 예시
xAI Grokhttps://api.core.today/llm/grok/v1grok-4.6, grok-4.3, grok-build-0.1
DeepSeekhttps://api.core.today/llm/deepseek/v1deepseek-v4-flash, deepseek-v4-pro
Metahttps://api.core.today/llm/meta/v1muse-spark-1.3, muse-spark-1.3-contributor
OpenRouterhttps://api.core.today/llm/openrouter/v1qwen/qwen3.8-max, moonshotai/kimi-k3, openai/gpt-oss-120b
curl -X POST https://api.core.today/llm/openrouter/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
    "model": "qwen/qwen3.8-max",
    "messages": [{"role": "user", "content": "Explain MoE routing in two sentences"}]
  }'
모델InputOutput
모델InputOutput
모델InputOutput
Meta Muse Spark 안내: muse-spark-1.3-contributor프롬프트와 응답이 Meta 모델 학습에 제공되는 조건의 할인 티어입니다 (표준 대비 약 1/12 단가). 개인정보·기밀·고객 데이터가 포함된 요청에는 표준 muse-spark-1.3을 쓰세요. 두 모델 모두 추론 모델이라 답변이 reasoning 토큰 뒤에 생성됩니다 — 실측으로 "세 단어로 인사"에도 약 700 completion 토큰(그중 reasoning 687)이 쓰였고, max_tokens 400에서는 content: null로 끝났습니다. max_tokens는 2,048 이상을 권장하고, reasoning 토큰도 출력 단가로 과금됩니다. 컨텍스트 1M, 이미지·영상·오디오·PDF 입력 지원, 캐시 입력 할인.
OpenRouter 과금 안내: 모델 ID는 vendor/model 형식이며 허용 목록(allowlist) 18종만 호출 가능합니다. 위 표의 단가는 표시·예약 추정용이고, 실제 청구는 OpenRouter 응답의 usage.cost(확정 USD) × 1,858 크레딧/USD 실비 정산입니다 — OpenRouter 측 가격이 바뀌어도 청구는 항상 실비 기준입니다. Claude/GPT/Gemini는 직접 라우트(/llm/anthropic 등)가 애그리게이터 수수료 없이 더 저렴합니다.
참고: Grok 전 모델은 프롬프트가 200K 토큰에 도달하면 요청 전체가 2배 단가(longcontext 티어)로 과금됩니다 — 자동 적용이며 -longcontext ID를 직접 호출하면 400으로 거부됩니다. grok-4.20-multi-agent-0309Responses API 전용(/llm/grok/v1/responses, chat/completions는 400)이며 서브에이전트 토큰까지 합산 청구되어 한 마디 질문에도 회당 약 4K 토큰(≈9크레딧)이 정산됩니다 — 짧은 대화용이 아니라 리서치급 질문에 쓰세요. DeepSeek은 시간대와 무관하게 표준(피크) 단가로 과금됩니다. Groq/MiniMax 직접 연동(llama-3.3-70b-versatile, MiniMax-M2.7)은 제공하지 않으며 같은 모델을 OpenRouter ID(meta-llama/llama-3.3-70b-instruct, minimax/minimax-m2.7)로 사용하세요.

스트리밍 응답

실시간으로 응답을 받으려면 stream: true를 추가하세요:

curl -X POST https://api.core.today/llm/openai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
    "model": "gpt-5",
    "messages": [{"role": "user", "content": "Tell me a long story"}],
    "stream": true
  }'

프롬프트 캐싱 (Claude 자동 비용 절감)

Claude 모델을 호출할 때 매번 똑같이 반복되는 앞부분(시스템 프롬프트·대화 히스토리)을 게이트웨이가 자동으로 캐시합니다. 두 번째 요청부터 그 부분은 다시 계산되지 않고 캐시에서 읽혀, 입력 토큰가의 10% 가격으로 청구됩니다. 별도 설정이나 코드 변경은 필요 없습니다.

캐시되는 부분조건
시스템 프롬프트 (system)항상 — 같은 시스템 프롬프트를 쓰는 모든 요청
대화 히스토리 (직전 답변까지)멀티턴 대화일 때 — 이전 assistant 답변이 포함된 요청
새로 보낸 마지막 질문캐시하지 않음 — 매번 바뀌므로 정상 가격

같은 시스템 프롬프트를 재사용하거나 대화를 계속 이어가면 캐시가 잘 적중합니다:

1번째 요청:  [긴 시스템 프롬프트] + "질문 A"   → 시스템 프롬프트를 캐시에 저장
2번째 요청:  [긴 시스템 프롬프트] + "질문 B"   → 캐시에서 읽음 (10% 가격)
3번째 요청:  [긴 시스템 프롬프트] + "질문 C"   → 캐시에서 읽음
캐시 적중 확인: 응답의 usage.cache_read_input_tokens가 0보다 크면 캐시가 적중한 것입니다. (cache_creation_input_tokens는 캐시에 새로 저장된 토큰)
참고
  • Claude(Anthropic) 요청에만 적용됩니다. OpenAI는 서버가 자동 캐싱, Gemini는 별도 방식입니다.
  • 앞부분이 최소 캐시 길이(대략 1,024토큰)에 못 미치면 캐시가 적용되지 않습니다 — 이때는 추가 비용도 없습니다.
  • 캐시는 마지막 사용 후 약 5분간 유지되며, 그 안에 다시 쓰이면 자동 연장됩니다.
  • 요청에 직접 cache_control을 넣은 경우, 그 설정을 그대로 사용합니다(게이트웨이가 손대지 않음).

비용 계산 예시

1,000 토큰 입력, 500 토큰 출력 기준:

모델계산총 비용
gpt-4o-mini0.3 + 0.60.9 크레딧
gpt-52.5 + 10.012.5 크레딧
claude-3-haiku0.5 + 1.251.75 크레딧
claude-sonnet-46.0 + 15.021.0 크레딧
gemini-2.0-flash0.2 + 0.40.6 크레딧