NVIDIA ENTERPRISE
Model Configuration Guide
Generate optimized deployment configurations for your LLM inference needs
Model
Llama 4 Scout (109B)
Llama 4 Maverick (400B)
Llama 3.3 70B
DeepSeek R1 (671B)
DeepSeek R1 Distill 70B
DeepSeek V3 (685B)
Qwen3 8B
Qwen3 70B
Qwen3 235B MoE
Hardware
H100 80GB
H200 141GB
B200 192GB
GB200 NVL72
Framework
TensorRT-LLM
vLLM
SGLang
Dynamo
Optimization Scenario
Maximum Throughput
Minimum Latency
Balanced
Generate Configuration Guide