thairouter
All models

GLM 5.3 Flash

by Z.ai Added Sep 9, 2026
Liveopen-weightsreasoninglong-context
thairouter/glm-5.3-flash
Try now Open in playground Get an API keyCode samples

Try now

Streams straight from ThaiRouter with this model's preset and the thinking effort you pick below. Sign in, confirm billing once, and go.

฿5/฿15 per 1M tokens · billed from your credits · nothing storedPlayground

Overview

Fast reasoning model with a very long context window. Strong multilingual quality, good at instruction following, coding and structured output. Streams reasoning before the final answer.

Context
393.2K
tokens
Max output
131.1K
tokens per reply
Input
฿5
per 1M tokens
Output
฿15
per 1M tokens

Live stats

Measured from real traffic through ThaiRouter: tiles are the last 7 days, charts the last 30. Aggregate only — no prompts are stored.

Latency p50
Latency p95
Throughput
Success rate
Tokens per day
Mean latency

Benchmarks

Reference scores published by Z.ai · GLM-4.6, not measured on ThaiRouter. Live latency and throughput are in the section above.

reasoning
AIME 202593.9%
GPQA Diamond82.9%
coding
LiveCodeBench v682.8%
agentic
SWE-bench Verified68%

API

OpenAI-compatible. Point your client at https://api.thairouter.ai/v1 and send this id.

Thinking effort

Default max. Cannot be switched off; pick low for fast replies. Thinking tokens are billed as output.

curl https://api.thairouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THAIROUTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "thairouter/glm-5.3-flash",
    "messages": [{"role": "user", "content": "สวัสดี"}],
    "reasoning_effort": "max"
  }'

More models