Skip to main content
Core.Today
|
Z.aiFastHigh

GLM 5.3 Flash

Z.ai's GLM 5.3 Flash served through the OpenRouter aggregator (listed 2026-08-26) - a low-cost native multimodal model (text, image, video input) for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention keeps long-context behavior accurate across a 1.31M-token window. Billed at OpenRouter's actual usage cost.

279/929credits
input / output ยท per 1M tokens
1.31M-token context window (1M via the top provider), very large output budget
Very low cost - budget-tier input and output rates
Native multimodal input (text, image, video) with reasoning (reasoning_effort + parallel tool calls)
OpenAI-compatible API (chat completions) via /llm/openrouter/v1
Billed at OpenRouter's actual usage.cost - listed rates are estimates

Run it right now

Test this model instantly in the Console Playground โ€” no code required

Sign in to try

Use with AI Assistant

Copy usage instructions for Claude, ChatGPT, or other AI

Model Specifications

Context Window
1.3M
tokens
Max Output
944K
tokens
Training Cutoff
Not published
Compatible SDK
OpenAI

Capabilities

Vision
Function Calling
Streaming
JSON Mode
System Prompt

Token Pricing (per 1M tokens)

Token TypeCreditsUSD Equivalent
Input Tokens279$0.19
Output Tokens929$0.62
Cached Tokens55.74$0.04

* 1,500 credits โ‰ˆ $1 (actual charges may vary based on usage)

Quick Start

curl -X POST "https://api.core.today/llm/openrouter/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
  "model": "z-ai/glm-5.3-flash",
  "messages": [
    {
      "role": "user",
      "content": "Summarize the attached report in 5 bullets."
    }
  ],
  "max_tokens": 1024
}'

Parameters

ParameterTypeRequiredDefaultDescription
modelstringYes-Model ID in vendor/model form, e.g. "z-ai/glm-5.3-flash".
messagesarrayYes-Chat messages in OpenAI format (system/user/assistant roles).
temperaturenumberNo0.7Sampling temperature (0-2).
max_tokensintegerNo2048Maximum completion tokens.
streambooleanNofalseStream the response as server-sent events.

Examples

Chat completion

OpenAI-compatible chat completion request via OpenRouter

curl -X POST "https://api.core.today/llm/openrouter/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer cdt_your_api_key" \
  -d '{
  "model": "z-ai/glm-5.3-flash",
  "messages": [
    {
      "role": "user",
      "content": "Summarize the attached report in 5 bullets."
    }
  ],
  "max_tokens": 1024
}'

Tips & Best Practices

1Works with the OpenAI SDK - set base_url to https://ai.api.core.today/llm/openrouter/v1
2Charged at OpenRouter's usage.cost x 1,858 credits/USD - the listed rates are estimates and may lag OpenRouter price changes
3Accepts image and video input in OpenAI content-part format; use reasoning_effort to trade latency for depth, and step up to z-ai/glm-5.3 for the hardest engineering tasks

Use Cases

High-volume coding assistance and agent loops
Screenshot, image and video understanding in agents
Cheap long-context analysis over very long inputs