Skip to main content
Core.Today
|
DeepSeekFastHigh

DeepSeek Flash (V4.1)

DeepSeek's fast, low-cost model - currently DeepSeek V4.1 Flash, with text and image input. Thinking mode is on by default, with a 1M-token context window, up to 384K output tokens, and cache reads at ~98% off. $0.30 input / $1.20 output per million tokens.

557/2,230credits
input / output ยท per 1M tokens
Best-in-class price/performance
Text and image input
1M context window, up to 384K output tokens
Thinking mode on by default (switchable per request)
OpenAI-compatible API (chat completions)
Context cache hits billed at ~98% discount

Run it right now

Test this model instantly in the Console Playground โ€” no code required

Sign in to try

Use with AI Assistant

Copy usage instructions for Claude, ChatGPT, or other AI

Model Specifications

Context Window
1.0M
tokens
Max Output
393K
tokens
Training Cutoff
2025-05
Compatible SDK
OpenAI

Capabilities

Vision
Function Calling
Streaming
JSON Mode
System Prompt

Token Pricing (per 1M tokens)

Token TypeCreditsUSD Equivalent
Input Tokens557$0.37
Output Tokens2,230$1.49
Cached Tokens11.148$0.01

* 1,500 credits โ‰ˆ $1 (actual charges may vary based on usage)

Quick Start

curl -X POST "https://api.core.today/llm/deepseek/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
  "model": "deepseek-flash",
  "messages": [
    {
      "role": "user",
      "content": "Explain vector databases in 3 sentences."
    }
  ],
  "max_tokens": 1024
}'

Parameters

ParameterTypeRequiredDefaultDescription
modelstringYes-Model ID, e.g. "deepseek-flash".
messagesarrayYes-Chat messages in OpenAI format (system/user/assistant roles).
temperaturenumberNo0.7Sampling temperature (0-2).
max_tokensintegerNo2048Maximum completion tokens.
streambooleanNofalseStream the response as server-sent events.

Examples

Chat completion

OpenAI-compatible chat completion request

curl -X POST "https://api.core.today/llm/deepseek/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
  "model": "deepseek-flash",
  "messages": [
    {
      "role": "user",
      "content": "Explain vector databases in 3 sentences."
    }
  ],
  "max_tokens": 1024
}'

Tips & Best Practices

1Works with the OpenAI SDK - set base_url to https://ai.api.core.today/llm/deepseek/v1
2Thinking mode is on by default - disable it per request for latency-sensitive calls
3Cache hits are billed at ~98% off - reuse long system prompts and shared prefixes
4Billed at DeepSeek's standard (peak) rate at all hours - off-peak discounts are not passed through
5Anthropic-format endpoint is not supported through the gateway - use the OpenAI format
6Billed at DeepSeek's standard (peak) rate at all hours โ€” DeepSeek's off-peak discount is not passed through

Use Cases

High-volume chat and summarization at low cost
Long-document analysis with the 1M context
Migration target for the legacy deepseek-v4-flash id and the retired deepseek-chat / deepseek-reasoner