Skip to main content
Core.Today
|
GoogleFastStandard

Gemini 3.5 Flash Lite

Google's lightweight Gemini 3.5 model (released 2026-07-21) for high-throughput, budget-conscious workloads. $0.30/$2.50 per million tokens with cached input at $0.03/M. Successor to Gemini 3.1 Flash Lite, which Google retires on 2027-05-07.

557/4,645credits
input / output ยท per 1M tokens
Lightweight, low-latency Gemini 3.5 tier
1,048,576 token context window
65,536 max output tokens
Multimodal input: text, image, audio, video, PDF
Cached input pricing (90% discount)
Function calling, structured outputs, thinking, search grounding

Run it right now

Test this model instantly in the Console Playground โ€” no code required

Sign in to try

Use with AI Assistant

Copy usage instructions for Claude, ChatGPT, or other AI

Model Specifications

Context Window
1.0M
tokens
Max Output
66K
tokens
Training Cutoff
January 2025
Compatible SDK
OpenAI, Google AI

Capabilities

Vision
Function Calling
Streaming
JSON Mode
System Prompt

Token Pricing (per 1M tokens)

Token TypeCreditsUSD Equivalent
Input Tokens557$0.37
Output Tokens4,645$3.10
Cached Tokens55.74$0.04

* 1,500 credits โ‰ˆ $1 (actual charges may vary based on usage)

Quick Start

curl -X POST "https://api.core.today/llm/gemini/v1beta/openai/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
  "model": "gemini-3.5-flash-lite",
  "messages": [
    {
      "role": "system",
      "content": "Classify the following text as: spam, not_spam. Respond with only the label."
    },
    {
      "role": "user",
      "content": "Congratulations! You have been selected for a special prize. Click here to claim now!"
    }
  ],
  "max_tokens": 50,
  "temperature": 0
}'

Parameters

ParameterTypeRequiredDefaultDescription
messagesarrayYes-Array of message objects (OpenAI format). Supports text, image, audio, video, and PDF inputs.
temperaturefloatNo1Sampling temperature (0-2). Lower values produce more deterministic outputs.
max_tokensintegerNo-Maximum output tokens. Max: 65,536. Context window (input + output): 1,048,576 tokens.
response_formatobjectNo-Output format constraint. Use `{ type: 'json_object' }` for structured JSON output.
streambooleanNofalseEnable Server-Sent Events streaming.

Examples

Quick Classification

Lightweight text classification with Flash Lite

curl -X POST "https://api.core.today/llm/gemini/v1beta/openai/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
  "model": "gemini-3.5-flash-lite",
  "messages": [
    {
      "role": "system",
      "content": "Classify the following text as: spam, not_spam. Respond with only the label."
    },
    {
      "role": "user",
      "content": "Congratulations! You have been selected for a special prize. Click here to claim now!"
    }
  ],
  "max_tokens": 50,
  "temperature": 0
}'

Tips & Best Practices

1$0.30/$2.50 per 1M tokens - output costs more than 3.1 Flash Lite ($1.50), so compare on your real output length
2Gemini 3.1 Flash Lite is retired on 2027-05-07 - migrate to this model before then
3Use cached input tokens ($0.03/M, 90% off) for repeated context
4Ideal for high-volume classification and routing tasks
5Priority inference (serviceTier: priority) bills token charges at 1.8x; flex and batch discounts are not passed through and settle at the standard rate

Use Cases

High-volume text processing
Real-time chat applications
Quick classification and routing
Lightweight data extraction
Migration target for Gemini 3.1 Flash Lite