DeepSeek-V4.1-Flash

NEW

Key Features

A dedicated vision encoder integrates image embeddings with text from pretraining onward.
The model supports contexts of up to one million tokens.
The backbone activates 8B parameters during prefill and 16B during decoding.
CSA2 combines cross-layer cache sharing, index reuse, and hierarchical sparse indexing.
The architecture includes FP4 main KV caching for a smaller memory footprint.
An integer reasoning-effort control ranges from 1 to 100.
DSpark supports speculative decoding with confidence-scheduled verification.
Engram adds sparsely accessed conditional memory through token-based lookup.

Its 552B-parameter backbone uses a causal encoder-decoder arrangement, activating 8B parameters during prefill and 16B during decoding. Compressed Sparse Attention 2 shares cache structures and sparse indices across layers. Additional mechanisms include FP4 KV caching, Engram conditional memory, and DSpark speculative decoding.


The model is useful for developers building agents that must reason over large documents, images, and extensive interaction histories. Its public checkpoint enables infrastructure-level experimentation, while a continuously adjustable reasoning setting trades computation for accuracy. Hosting remains a substantial multi-accelerator infrastructure task despite sparse activation.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!