Z.ai
GLM-5.3-Flash
Specs
- Input price
- $0.15/M
- Cached input price
- $0.03/M
- Output price
- $0.50/M
- Context window
- 1M tokens
Z.ai's first natively multimodal GLM-5 model - 320B total, 18B active, MIT license, native vision, and a 1M-token context window at flash-tier pricing.
Capabilities
- 320B total / 18B active MoE under a clean MIT license
- Native vision understanding from pre-training (30T-token multimodal corpus)
- Hybrid sparse-plus-linear attention with Manifold-Constrained Hyper-Connections
- reasoning_effort control (low, high, max) from flash-tier to max-effort runs
Best for
- Multimodal document pipelines: scanned records, charts, and screenshots
- High-volume agentic coding at flash-tier cost (Terminal-Bench 2.1 at 84.3)
- Long-context repo-level work on a budget
Limitations to keep in mind
- Vendor-reported benchmarks on new harnesses - treat as directional until third-party reruns
- Vision path is not a diagnostic or DICOM tool
- 1M-token context not yet verified against third-party runners
HIPAA-compliant hosting
GLM-5.3-Flash is available under our HIPAA-compliant design-partner program, with a signed Business Associate Agreement (BAA), encryption in transit and at rest, and access controls. We are currently onboarding design partners, with general availability coming soon. Run the model behind a unified OpenRouter-style API and swap to another model with a single parameter.
Pricing in context
Open source models like GLM-5.3-Flash can be served efficiently on optimized inference infrastructure, with savings passed through to you. Exact savings depend on the model and your volume, but open source inference is typically a fraction of the per-token cost of closed-source frontier models on Azure OpenAI or AWS Bedrock - without cloud egress lock-in or minimum commitments.
Design partner program
Run GLM-5.3-Flash under HIPAA
Join the waitlist to be prioritized. We'll reach out with a qualification call and early access.