z-ai/glm-5.3-flash
1048K context|$0.15/M input tokens|$0.5/M output tokens|$0.035/M cache read
GLM-5.3-Flash is a native multimodal mixture-of-experts model for coding, agentic workflows, reasoning, tool use, and visual understanding.
Input / M tokens
$0.15
Output / M tokens
$0.5
Cache read
$0.035
Context
1048K
Sample code and API for GLM 5.3 Flash
You can also use NEAR AI Cloud with OpenAI's client API:
Verify GLM 5.3 Flash
This model is hosted in a GPU TEE. TEE enables strong privacy with end-to-end encryption for data in transmission and in use. TEE protects the entire lifecycle of the model inference and gives you a proof of the execution. It allows you to verify that the given a model and an input, it produces a specific output both onchain and offchain.
Check out the docs for more details.Get the attestation report.
Retrieve the response signature using the chat completion ID. Signatures are available for 5 minutes.