Skip to main content

Kunlun R300

Kunlun R300 is an OAM accelerator module (OCP-OAI standard) built on the 2nd-generation Kunlun XPU-R chip by Kunlunxin (a Baidu company). Same silicon as the R200 full-height PCIe card but in module form: R300 modules populate the R480-X8 server baseboard (8 modules per board) for large-scale data-center training and inference clusters.


Core Specifications

SpecValue
ArchitectureKunlun Gen-2 XPU-R (proprietary architecture, on-chip video codec)
Process7nm
FP16128 TFLOPS
INT8256 TOPS
FP3232 TFLOPS
Memory32 GB GDDR6
Memory bandwidth512 GB/s
TDP150 W
Form factorOAM accelerator module (OCP-OAI standard, UBB baseboard)
Chip-to-chip200 GB/s (dual-ring topology inside R480-X8)
Launch2021-08 (with Kunlun Gen-2)
PriceUndisclosed

📌 Form-factor clarification (2026-09 cross-validation): The R300 is not a new chip — it is the OAM module variant of Kunlun Gen 2. Kunlunxin's product matrix lists the R200 as the PCIe card, the R300 as the OAM module line, and the R480-X8 as the 8-module server baseboard; compute specifications are identical across all three (FP16 128 / INT8 256 / FP32 32 / 32GB GDDR6 / 512GB/s / 150W).


Highlights

  • OCP-OAI standard: UBB server baseboard carrying 8 OAM modules (R480-X8) delivers roughly 1 PFLOPS FP16 per server
  • Integrated codec: on-chip video encode/decode enables decode-plus-AI-inference pipelines without host data shuttling
  • Hardware virtualization: per-chip virtualization improves cloud resource utilization
  • 10K-card cluster building block: a foundational compute unit of Baidu Cloud's 10K-card Kunlun cluster (lit up in 2024 — China's first at-scale cluster on proprietary silicon)

Positioning

Kunlun product lineage: Kunlun Gen 1 (K100/K200 PCIe cards) → Kunlun Gen 2 (R200 PCIe card / R300 OAM module / R480-X8 server) → Kunlun Gen 3 → P800 (2024, third-gen training card). The R300 serves Baidu Cloud and China Mobile-class carrier AI data centers (Kunlun-based servers took the top share in China Mobile's 2025-2026 inference-type AI server tender, CUDA-ecosystem segment).


Use Cases

  • Data-center-scale training and inference (R480-X8 8-module servers)
  • Search, recommendation, and video-analytics inference for Baidu's core businesses
  • Carrier and xinchuang AI data-center tenders
  • ❌ Single-card developer setups (choose the R200 PCIe form factor)

References