High-throughput agents
Nano
AvailableThe smallest member of the family, designed to pair strong reasoning with inference efficiency. NVIDIA describes a hybrid Mamba-Transformer mixture-of-experts architecture.
- Total parameters
- 31.6B
- Active parameters
- 3.2B / 3.6B with embeddings
- Context
- Up to 1M