NVIDIA Nemotron 3 Embed

NEW

Key Features

Provides open embedding models for RAG, agentic retrieval, code retrieval, and memory systems.
Includes an 8B BF16 flagship model for maximum retrieval quality.
Includes 1B BF16 and 1B NVFP4 variants for efficient production deployment.
Ranks the 8B model first on RTEB in the cited release.
Supports 32k context windows for long-input retrieval.
Uses bidirectional adaptation of instruction-model backbones for embedding generation.
Compresses smaller variants through structured pruning and teacher distillation.
Offers an optimized NVIDIA NIM serving path for production-scale retrieval systems.

The 8B model adapts a Ministral-3-8B-Instruct backbone into a bidirectional encoder for full-sequence retrieval, while the 1B models are produced through structured pruning and distillation from larger teacher checkpoints. The collection includes BF16 variants plus an NVFP4 Blackwell-optimized model that quantizes weights and activations for high-throughput serving.


Nemotron 3 Embed is useful for teams building production retrieval pipelines where better embeddings reduce repeated searches, irrelevant context, and downstream agent token cost. NVIDIA reports the 8B model ranks first on RTEB, strong MMTEB and LongEmbed results, and optimized NIM serving support for enterprise retrieval workloads.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner
Zero to AI Engineer Program

Zero to AI Engineer

Skip the degree. Learn real-world AI skills used by AI researchers and engineers. Get certified in 8 weeks or less. No experience required.

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!