Skip to main content
Core.Today
|
GoogleFastHigh

Gemini 3.6 Flash

Google's Gemini 3.6 Flash (released 2026-07-21), an earlier release in the Gemini 3.x Flash line - newer Flash generations (3.7 and 3.8) are available. A 1M-token context window with 65,536 max output tokens, multimodal input, thinking and tool use, priced at $1.50 input / $7.50 output per million tokens with cached inputs at $0.15/M.

2,787/13,935credits
input / output ยท per 1M tokens
Gemini 3.6 Flash generation (released 2026-07-21)
1,048,576 token context window
65,536 max output tokens
Multimodal input: text, image, audio, video, PDF
Cached input pricing (90% discount)
Function calling, structured outputs, thinking, search grounding, code execution

Run it right now

Test this model instantly in the Console Playground โ€” no code required

Sign in to try

Use with AI Assistant

Copy usage instructions for Claude, ChatGPT, or other AI

Model Specifications

Context Window
1.0M
tokens
Max Output
66K
tokens
Training Cutoff
2025-01
Compatible SDK
OpenAI, Google AI

Capabilities

Vision
Function Calling
Streaming
JSON Mode
System Prompt

Token Pricing (per 1M tokens)

Token TypeCreditsUSD Equivalent
Input Tokens2,787$1.86
Output Tokens13,935$9.29
Cached Tokens278.7$0.19

* 1,500 credits โ‰ˆ $1 (actual charges may vary based on usage)

Quick Start

curl -X POST "https://api.core.today/llm/gemini/v1beta/openai/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
  "model": "gemini-3.6-flash",
  "messages": [
    {
      "role": "user",
      "content": "Explain the difference between ECS and EKS briefly."
    }
  ],
  "stream": true,
  "max_tokens": 800
}'

Parameters

ParameterTypeRequiredDefaultDescription
messagesarrayYes-Array of message objects (OpenAI format). Supports text, image, audio, video, and PDF inputs.
temperaturefloatNo1Sampling temperature (0-2). Lower values produce more deterministic outputs.
max_tokensintegerNo-Maximum output tokens. Max: 65,536. Context window (input + output): 1,048,576 tokens.
response_formatobjectNo-Output format constraint. Use `{ type: 'json_object' }` for structured JSON output.
streambooleanNofalseEnable Server-Sent Events streaming.

Examples

Production Chat

High-throughput assistant responses with streaming

curl -X POST "https://api.core.today/llm/gemini/v1beta/openai/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
  "model": "gemini-3.6-flash",
  "messages": [
    {
      "role": "user",
      "content": "Explain the difference between ECS and EKS briefly."
    }
  ],
  "stream": true,
  "max_tokens": 800
}'

Document Extraction

Pull structured fields out of a long document

curl -X POST "https://api.core.today/llm/gemini/v1beta/openai/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
  "model": "gemini-3.6-flash",
  "messages": [
    {
      "role": "user",
      "content": "Summarize the attached contract PDF and list every deadline it mentions."
    }
  ],
  "response_format": {
    "type": "json_object"
  },
  "max_tokens": 2000
}'

Tips & Best Practices

1Same price as Gemini 3.8 Flash ($1.50/$7.50) - for new projects prefer the newest generation, gemini-3.8-flash; pin gemini-3.6-flash only if you need this version's behavior
2Use cached inputs ($0.15/M, 90% off) for repeated context
3Max output is 65,536 tokens - chunk very long generations
4Step down to Gemini 3.5 Flash Lite ($0.30/$2.50) for simple high-volume tasks
5Priority inference (serviceTier: priority) bills token charges at 1.8x; flex and batch discounts are not passed through and settle at the standard rate

Use Cases

High-volume production chat workloads
Multimodal understanding (image, video, audio, PDF)
Long document and codebase analysis
Data extraction and summarization at scale
Repeated context with cached inputs (RAG)