deepseek-ai/DeepSeek-V4-Flash
1048K context|$0.17/M input tokens|$0.35/M output tokens|$0.035/M cache read
DeepSeek V4 Flash — large mixture-of-experts language model from DeepSeek, FP8-quantized. Served on H200 with TP=4 and EAGLE speculative decoding in a TDX-confidential inference CVM.
Input / M tokens
$0.17
Output / M tokens
$0.35
Cache read
$0.035
Context
1048K
Sample code and API for DeepSeek V4 Flash
You can also use NEAR AI Cloud with OpenAI's client API:
Verify DeepSeek V4 Flash
This model is hosted in a GPU TEE. TEE enables strong privacy with end-to-end encryption for data in transmission and in use. TEE protects the entire lifecycle of the model inference and gives you a proof of the execution. It allows you to verify that the given a model and an input, it produces a specific output both onchain and offchain.
Check out the docs for more details.Get the attestation report.
Retrieve the response signature using the chat completion ID. Signatures are available for 5 minutes.