The 8B model adapts a Ministral-3-8B-Instruct backbone into a bidirectional encoder for full-sequence retrieval, while the 1B models are produced through structured pruning and distillation from larger teacher checkpoints. The collection includes BF16 variants plus an NVFP4 Blackwell-optimized model that quantizes weights and activations for high-throughput serving.
Nemotron 3 Embed is useful for teams building production retrieval pipelines where better embeddings reduce repeated searches, irrelevant context, and downstream agent token cost. NVIDIA reports the 8B model ranks first on RTEB, strong MMTEB and LongEmbed results, and optimized NIM serving support for enterprise retrieval workloads.


