Skip to main content

NVIDIA L20 (Ada Inference Accelerator)

Product Overview

The NVIDIA L20 is an Ada Lovelace architecture data center inference card positioned between the L4 and L40S. With 48GB GDDR6 ECC memory, 239 TFLOPS FP8 sparse compute, and a 275W TDP, it is a mainstream PCIe accelerator for generative AI and LLM inference, ranking alongside the H20 as a popular choice for LLM inference in China.

Based on the full AD102 die (11,776 CUDA cores), it supports MIG, vGPU, and structured sparsity, covering AI inference, 3D graphics, and video processing.

Core Specifications

ParameterValue
ArchitectureAda Lovelace (AD102)
ProcessTSMC 4N
CUDA Cores11,776
Tensor Cores368 (4th Gen)
RT Cores92 (3rd Gen)
Memory48 GB GDDR6 ECC
Memory Bandwidth864 GB/s
L2 Cache96 MB
FP3259.8 TFLOPS
FP16 Tensor119.5 TFLOPS (dense) / 239 TFLOPS (sparse)
BF16 Tensor119.5 TFLOPS (dense) / 239 TFLOPS (sparse)
FP8 Tensor119.5 TFLOPS (dense) / 239 TFLOPS (sparse)
INT8 Tensor119.5 TOPS (dense) / 239 TOPS (sparse)
TDP275 W
InterfacePCIe Gen4 ×16 (64 GB/s)
Form Factor2-slot FHFL, passive cooling
Video Engines3× NVENC (incl. AV1) + 3× NVDEC + 4× NVJPEG
MIGSupported (1g/2g/4g instances)
Launch2023-Q4
Price$4,500 (market price, no official MSRP)

L20 vs L4 vs L40S Comparison

MetricL20L4L40S
ArchitectureAda AD102Ada AD104Ada AD102
CUDA Cores11,7767,68018,176
Memory48GB GDDR6 ECC24GB GDDR648GB GDDR6 ECC
Bandwidth864 GB/s300 GB/s864 GB/s
FP8 Tensor (sparse)239 TFLOPS485 TFLOPS733 TFLOPS
FP3259.8 TFLOPS30.3 TFLOPS91.6 TFLOPS
TDP275W72W350W
PositioningLLM inferenceLow-power inferenceAll-round inference + rendering

The L20 matches the L40S on memory (48GB) and bandwidth, but its FP8 sparse compute is only 33% of the L40S; it draws 75W less and costs roughly half — a cost-effective LLM inference choice.

Use Cases

  • LLM / generative AI inference (quantized deployment of 70B-class models)
  • ✅ Multi-instance inference (MIG partitioning, multi-tenant sharing)
  • ✅ Enterprise vGPU virtualization (vPC/vWS)
  • ✅ 3D rendering and video transcoding (NVENC AV1)
  • ✅ Scientific computing / data science (FP32 + Tensor)
  • ❌ Large-scale training (use H100/B200)
  • ❌ Ultra-low-power edge deployment (use L4/T4)

Vendor Information

ParameterValue
VendorNVIDIA
Target marketData center AI inference, generative AI
Price$4,000-$6,000 (market price, no official MSRP)