# Gemini 3.7 Flash - Core.Today AI API > Google's Gemini 3.7 Flash (released 2026-08-13), an earlier release in the Gemini 3.x Flash line - newer Flash generations (3.8) are available. A 1M-token context window with 65,536 max output tokens, multimodal input, thinking and tool use, priced at $1.50 input / $7.50 output per million tokens with cached inputs at $0.15/M. - **Provider**: Google - **Model ID**: gemini-3.7-flash - **Category**: LLM - **Credits**: 1 per request - **Speed**: Fast - **Quality**: High ## Model Specifications - **Context Window**: 1.0M tokens - **Max Output**: 66K tokens - **Training Cutoff**: 2025-01 - **Supported Formats**: text, json, markdown - **Compatible SDK**: OpenAI, Google AI ### Capabilities - Vision (image input) - Function Calling - Streaming - JSON Mode - System Prompt ### Token Pricing (per 1M tokens) - **Input Tokens**: 2,787 credits ($1.86) - **Output Tokens**: 13,935 credits ($9.29) - **Cached Tokens**: 278.7 credits ($0.19) ## Features - Gemini 3.7 Flash generation (released 2026-08-13) - 1,048,576 token context window - 65,536 max output tokens - Multimodal input: text, image, audio, video, PDF - Cached input pricing (90% discount) - Function calling, structured outputs, thinking, search grounding, code execution ## Use Cases - High-volume production chat workloads - Multimodal understanding (image, video, audio, PDF) - Long document and codebase analysis - Data extraction and summarization at scale - Repeated context with cached inputs (RAG) ## API Endpoint Base URL: https://api.core.today Endpoint: POST /llm/gemini/v1beta/openai/chat/completions ## Authentication Header: Authorization: Bearer YOUR_API_KEY Note: LLM endpoints use OpenAI-compatible format with Authorization Bearer token. ## Input Parameters ### Required - **messages**: array - Array of message objects (OpenAI format). Supports text, image, audio, video, and PDF inputs. ### Optional - **temperature**: float (default: 1) - Sampling temperature (0-2). Lower values produce more deterministic outputs. - **max_tokens**: integer - Maximum output tokens. Max: 65,536. Context window (input + output): 1,048,576 tokens. - **response_format**: object - Output format constraint. Use `{ type: 'json_object' }` for structured JSON output. - **stream**: boolean (default: false) - Enable Server-Sent Events streaming. ## Examples ### Production Chat High-throughput assistant responses with streaming ```json { "model": "gemini-3.7-flash", "messages": [ { "role": "user", "content": "Explain the difference between ECS and EKS briefly." } ], "stream": true, "max_tokens": 800 } ``` ### Document Extraction Pull structured fields out of a long document ```json { "model": "gemini-3.7-flash", "messages": [ { "role": "user", "content": "Summarize the attached contract PDF and list every deadline it mentions." } ], "response_format": { "type": "json_object" }, "max_tokens": 2000 } ``` ## Response Format ```json { "id": "chatcmpl-abc123", "object": "chat.completion", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Response text here" }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 100, "completion_tokens": 50, "total_tokens": 150 } } ``` ## Tips - Same price as Gemini 3.8 Flash ($1.50/$7.50) - for new projects prefer the newest generation, gemini-3.8-flash; pin gemini-3.7-flash only if you need this version's behavior - Use cached inputs ($0.15/M, 90% off) for repeated context - Max output is 65,536 tokens - chunk very long generations - Step down to Gemini 3.5 Flash Lite ($0.30/$2.50) for simple high-volume tasks - Priority inference (serviceTier: priority) bills token charges at 1.8x; flex and batch discounts are not passed through and settle at the standard rate ## Documentation https://ai.google.dev/gemini-api/docs