Qwen/Qwen3.6-35B-A3B-FP8
262K context|$0.17/M input tokens|$1.1/M output tokens|$0.056/M cache read
Qwen 3.6 35B is a fast mixture-of-experts language model with ~3B active parameters per token. Strong at reasoning, coding, and multilingual tasks with 32K context window.
Input / M tokens
$0.17
Output / M tokens
$1.1
Cache read
$0.056
Context
262K
Sample code and API for Qwen 3.6 35B A3B FP8
You can also use NEAR AI Cloud with OpenAI's client API:
Verify Qwen 3.6 35B A3B FP8
This model is hosted in a GPU TEE. TEE enables strong privacy with end-to-end encryption for data in transmission and in use. TEE protects the entire lifecycle of the model inference and gives you a proof of the execution. It allows you to verify that the given a model and an input, it produces a specific output both onchain and offchain.
Check out the docs for more details.Get the attestation report.
Retrieve the response signature using the chat completion ID. Signatures are available for 5 minutes.