Skip to main content

Cambricon MLU690 (Chinese AI Training/Inference Chip)

Product Overview

Cambricon MLU690 is Cambricon's flagship chip for the data center AI training and inference market, in mass production in early 2026. Using a dual-die chiplet package and the SMIC 5nm-class (N+2) process, its 700+ TFLOPS FP16 compute represents the highest paper specification among Chinese AI chips; single-card compute is roughly 70% of the NVIDIA H100, reaching 80-90% of the H100 in pure inference scenarios.

Strategic positioning: ByteDance is the primary customer (over 500,000 units of the MLU580/590/690 combined purchased), and orders have been secured from leading internet companies such as Alibaba and Baidu. Against the backdrop of US chip export controls on China, the MLU690 is a key domestic alternative in China's large-model training/inference market.

📌 Supernode deployment (WAIC 2026): the MLU690 is in mass production, accompanied by the rollout of a 256-card scale supernode cluster solution already deployed in multiple intelligent computing centers — marking the move of Chinese chips from "single-card compute" toward "system-level compute".

Core Specifications

ParameterValue
ArchitectureMLUarch (new-generation MLU architecture)
ProcessTSMC 4nm (dual-die chiplet, Cambricon 4.0 architecture); an SMIC N+2 domestic version is also reported
PackagingDual-die chiplet
Transistor CountNot disclosed
Memory196 GB HBM3
Memory Bandwidth3.35 TB/s
FP16/BF16700+ TFLOPS
INT42,800+ TOPS (web reports; INT8 not disclosed)
FP32Not disclosed (estimated ~175 TFLOPS)
TDP400–500 W (web consensus 400W+)
InterconnectMLU-Link 890 Gbps / card
PCIePCIe 5.0
Announced2025
Mass Production2026-Q1
Unit Price¥140,000 - 150,000

MLU690 vs MLU590 vs H100

MetricMLU690MLU590H100 SXMGap (vs H100)
ProcessTSMC 4nm (dual die)7nm chipletTSMC 4nmOn par
Memory196GB HBM396GB HBM2e80GB HBM3+145%
Memory Bandwidth3.35 TB/s~2 TB/s3.35 TB/sOn par
FP16700+ TFLOPS345 TFLOPS989 TFLOPS~71%
INT82,800+ TOPS1,380 TOPS1,979 TOPS+41%
TDP~500W~350W700W-29%
InterconnectMLU-Link 890GbpsMLU-LinkNVLink 4 900GB/sNearly on par
Software ecosystemNeuWareNeuWareCUDAGap remains

Key insight: the MLU690's memory capacity is 2.45x that of the H100, and its INT8 inference compute surpasses the H100 by 41%. Its shortcomings mainly come from FP16 training compute and interconnect efficiency in distributed training beyond a thousand cards.

Process and Supply Chain

  • Foundry: after TSMC halted supply due to the Entity List, production shifted entirely to SMIC N+2 (5nm-class). The MLU690 is one of the largest customers for SMIC N+2
  • Yield: Sigmaintell estimates SMIC's 5nm yield at about 34% in 2025Q1, with 40%+ possible by Q4
  • Capacity booking: N+2 capacity is locked in through 2027
  • Next generation: tentatively the MLU790 / MLU690 successor, expected to ship in 2026Q4 and ramp in 2027

Smart Driving: MLU610

Cambricon has also launched the MLU610 chip for smart driving scenarios, which has won orders from BYD, Li Auto, and others, expanding into the edge/automotive AI market.

NeuWare Software Stack

LayerToolDescription
AI frameworksPyTorch-CambriconPyTorch adaptation
TensorFlow-CambriconTensorFlow adaptation
CompilerNeuWare CCnvcc-like
RuntimeNeuWare RuntimeCUDA Runtime-like
Math libraryNeuBLAScuBLAS-like
Deep learning libraryNeuDNNcuDNN-like
Inference engineMagicMindAmong the world's first commercial MLIR graph-compiler inference frameworks
Communication libraryMLU-LinkMulti-card interconnect

⚠️ Ecosystem limitations: NeuWare unifies cloud, edge, and endpoint with a single codebase — the only vendor in China doing so. However, distributed training beyond a thousand cards suffers relatively high losses (MLU-Link 890Gbps is close to NVLink 4's 900GB/s, but engineering maturity still lags); single-server 8-card and ~100-card small clusters perform excellently.

Vendor Information

ItemDetails
CompanyCambricon Technologies Corporation Limited
Stock Code688256.SH (STAR Market)
FounderChen Tianshi (Institute of Computing Technology, Chinese Academy of Sciences)
Founded2016-03
IPO2020-07-20 (first AI chip stock on the STAR Market)
2026Q1 Revenue¥2.885 billion (+160% YoY)
Market CapOver RMB 1 trillion (June 2026)
HeadquartersHaidian District, Beijing
Official Websitehttps://www.cambricon.com
Key CustomersByteDance (largest customer, ~65% of purchases), Alibaba, Baidu, China Mobile