# DeepSeek Flash (V4.1) - Core.Today AI API > DeepSeek's fast, low-cost model - currently DeepSeek V4.1 Flash, with text and image input. Thinking mode is on by default, with a 1M-token context window, up to 384K output tokens, and cache reads at ~98% off. $0.30 input / $1.20 output per million tokens. - **Provider**: DeepSeek - **Model ID**: deepseek-flash - **Category**: LLM - **Credits**: 1.4 per 1K tokens (avg) - **Speed**: Fast - **Quality**: High ## Model Specifications - **Context Window**: 1.0M tokens - **Max Output**: 393K tokens - **Training Cutoff**: 2025-05 - **Supported Formats**: text, json, markdown - **Compatible SDK**: OpenAI ### Capabilities - Vision (image input) - Function Calling - Streaming - JSON Mode - System Prompt ### Token Pricing (per 1M tokens) - **Input Tokens**: 557.4 credits ($0.37) - **Output Tokens**: 2,229.6 credits ($1.49) - **Cached Tokens**: 11.148 credits ($0.01) ## Features - Best-in-class price/performance - Text and image input - 1M context window, up to 384K output tokens - Thinking mode on by default (switchable per request) - OpenAI-compatible API (chat completions) - Context cache hits billed at ~98% discount ## Use Cases - High-volume chat and summarization at low cost - Long-document analysis with the 1M context - Migration target for the legacy deepseek-v4-flash id and the retired deepseek-chat / deepseek-reasoner ## API Endpoint Base URL: https://api.core.today Endpoint: POST /llm/deepseek/v1/chat/completions ## Authentication Header: Authorization: Bearer YOUR_API_KEY Note: LLM endpoints use OpenAI-compatible format with Authorization Bearer token. ## Input Parameters ### Required - **model**: string - Model ID, e.g. "deepseek-flash". - **messages**: array - Chat messages in OpenAI format (system/user/assistant roles). ### Optional - **temperature**: number (default: 0.7) - Sampling temperature (0-2). - **max_tokens**: integer (default: 2048) - Maximum completion tokens. - **stream**: boolean (default: false) - Stream the response as server-sent events. ## Examples ### Chat completion OpenAI-compatible chat completion request ```json { "model": "deepseek-flash", "messages": [ { "role": "user", "content": "Explain vector databases in 3 sentences." } ], "max_tokens": 1024 } ``` ## Response Format ```json { "id": "chatcmpl-abc123", "object": "chat.completion", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Response text here" }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 100, "completion_tokens": 50, "total_tokens": 150 } } ``` ## Tips - Works with the OpenAI SDK - set base_url to https://ai.api.core.today/llm/deepseek/v1 - Thinking mode is on by default - disable it per request for latency-sensitive calls - Cache hits are billed at ~98% off - reuse long system prompts and shared prefixes - Billed at DeepSeek's standard (peak) rate at all hours - off-peak discounts are not passed through - Anthropic-format endpoint is not supported through the gateway - use the OpenAI format - Billed at DeepSeek's standard (peak) rate at all hours — DeepSeek's off-peak discount is not passed through ## Documentation https://api-docs.deepseek.com/