NVIDIA L20 (Ada Inference Accelerator)
Product Overview
The NVIDIA L20 is an Ada Lovelace architecture data center inference card positioned between the L4 and L40S. With 48GB GDDR6 ECC memory, 239 TFLOPS FP8 sparse compute, and a 275W TDP, it is a mainstream PCIe accelerator for generative AI and LLM inference, ranking alongside the H20 as a popular choice for LLM inference in China.
Based on the full AD102 die (11,776 CUDA cores), it supports MIG, vGPU, and structured sparsity, covering AI inference, 3D graphics, and video processing.
Core Specifications
| Parameter | Value |
|---|---|
| Architecture | Ada Lovelace (AD102) |
| Process | TSMC 4N |
| CUDA Cores | 11,776 |
| Tensor Cores | 368 (4th Gen) |
| RT Cores | 92 (3rd Gen) |
| Memory | 48 GB GDDR6 ECC |
| Memory Bandwidth | 864 GB/s |
| L2 Cache | 96 MB |
| FP32 | 59.8 TFLOPS |
| FP16 Tensor | 119.5 TFLOPS (dense) / 239 TFLOPS (sparse) |
| BF16 Tensor | 119.5 TFLOPS (dense) / 239 TFLOPS (sparse) |
| FP8 Tensor | 119.5 TFLOPS (dense) / 239 TFLOPS (sparse) |
| INT8 Tensor | 119.5 TOPS (dense) / 239 TOPS (sparse) |
| TDP | 275 W |
| Interface | PCIe Gen4 ×16 (64 GB/s) |
| Form Factor | 2-slot FHFL, passive cooling |
| Video Engines | 3× NVENC (incl. AV1) + 3× NVDEC + 4× NVJPEG |
| MIG | Supported (1g/2g/4g instances) |
| Launch | 2023-Q4 |
| Price | $4,500 (market price, no official MSRP) |
L20 vs L4 vs L40S Comparison
| Metric | L20 | L4 | L40S |
|---|---|---|---|
| Architecture | Ada AD102 | Ada AD104 | Ada AD102 |
| CUDA Cores | 11,776 | 7,680 | 18,176 |
| Memory | 48GB GDDR6 ECC | 24GB GDDR6 | 48GB GDDR6 ECC |
| Bandwidth | 864 GB/s | 300 GB/s | 864 GB/s |
| FP8 Tensor (sparse) | 239 TFLOPS | 485 TFLOPS | 733 TFLOPS |
| FP32 | 59.8 TFLOPS | 30.3 TFLOPS | 91.6 TFLOPS |
| TDP | 275W | 72W | 350W |
| Positioning | LLM inference | Low-power inference | All-round inference + rendering |
The L20 matches the L40S on memory (48GB) and bandwidth, but its FP8 sparse compute is only 33% of the L40S; it draws 75W less and costs roughly half — a cost-effective LLM inference choice.
Use Cases
- ✅ LLM / generative AI inference (quantized deployment of 70B-class models)
- ✅ Multi-instance inference (MIG partitioning, multi-tenant sharing)
- ✅ Enterprise vGPU virtualization (vPC/vWS)
- ✅ 3D rendering and video transcoding (NVENC AV1)
- ✅ Scientific computing / data science (FP32 + Tensor)
- ❌ Large-scale training (use H100/B200)
- ❌ Ultra-low-power edge deployment (use L4/T4)
Vendor Information
| Parameter | Value |
|---|---|
| Vendor | NVIDIA |
| Target market | Data center AI inference, generative AI |
| Price | $4,000-$6,000 (market price, no official MSRP) |
Related Cards
- NVIDIA L4 - Low-power inference
- NVIDIA L40S - All-round inference + rendering
- NVIDIA H20 - China-specific large-memory inference
- NVIDIA A100 - Previous-generation data center training