Skip to content
MUSEBOARD

GPU / L40S

NVIDIA L40S 48GB

Universal datacenter GPU for inference and fine-tuning

Spec sheet
NVIDIA
Vendor
NVIDIA
Architecture
Ada Lovelace
Memory
48 GB GDDR6 ECC
Bandwidth
864 GB/s
Power
350 W
Class
Datacenter
Form factor
PCIe card
Interconnect
PCIe 4.0
Released
2023
Workloads
Inference · Fine-tuning
Model fit / Inference
8K ctx · batch 1 · ≥10% headroom
GPUs required per model and precision for NVIDIA L40S 48GB
ModelFP16INT8INT4
Llama 3.2 1B3.0 GB1.6 GB0.95 GB
Llama 3.2 3B8.0 GB4.5 GB2.7 GB
Llama 3.1 8B18.7 GB9.9 GB5.5 GB
Llama 3.1 70B158.0 GB80.3 GB41.5 GB
Llama 3.3 70B158.0 GB80.3 GB41.5 GB
Llama 3.1 405B24×897.2 GB16×450.7 GB227.5 GB
Qwen2.5 1.5B3.6 GB1.9 GB1.1 GB
Qwen2.5 3B7.1 GB3.7 GB2.0 GB
Qwen2.5 7B17.2 GB8.9 GB4.7 GB
Qwen2.5 14B34.0 GB17.8 GB9.7 GB
Qwen2.5 32B73.6 GB37.9 GB20.0 GB
Qwen2.5 72B162.6 GB82.7 GB42.7 GB
Mistral 7B17.0 GB9.0 GB5.1 GB
Mistral NeMo 12B28.2 GB14.8 GB8.1 GB
Mistral Small 24B53.3 GB27.3 GB14.3 GB
Mixtral 8x7B103.8 GB52.4 GB26.8 GB
Mixtral 8x22B312.1 GB157.0 GB79.4 GB
Gemma 2 9B23.1 GB13.0 GB7.9 GB
Gemma 2 27B62.9 GB33.0 GB18.0 GB
Phi-3 Mini 3.8B11.6 GB7.4 GB5.3 GB
Phi-3 Medium 14B32.5 GB17.1 GB9.4 GB
DeepSeek R1 Distill Qwen 32B74.3 GB38.2 GB20.2 GB
DeepSeek R1 Distill Llama 70B158.0 GB80.3 GB41.5 GB
DeepSeek V3 671B40×1,477 GB24×738.7 GB16×369.6 GB