Cambricon MLU690 (Chinese AI Training/Inference Chip)
Product Overview
Cambricon MLU690 is Cambricon's flagship chip for the data center AI training and inference market, in mass production in early 2026. Using a dual-die chiplet package and the SMIC 5nm-class (N+2) process, its 700+ TFLOPS FP16 compute represents the highest paper specification among Chinese AI chips; single-card compute is roughly 70% of the NVIDIA H100, reaching 80-90% of the H100 in pure inference scenarios.
Strategic positioning: ByteDance is the primary customer (over 500,000 units of the MLU580/590/690 combined purchased), and orders have been secured from leading internet companies such as Alibaba and Baidu. Against the backdrop of US chip export controls on China, the MLU690 is a key domestic alternative in China's large-model training/inference market.
📌 Supernode deployment (WAIC 2026): the MLU690 is in mass production, accompanied by the rollout of a 256-card scale supernode cluster solution already deployed in multiple intelligent computing centers — marking the move of Chinese chips from "single-card compute" toward "system-level compute".
Core Specifications
| Parameter | Value |
|---|---|
| Architecture | MLUarch (new-generation MLU architecture) |
| Process | TSMC 4nm (dual-die chiplet, Cambricon 4.0 architecture); an SMIC N+2 domestic version is also reported |
| Packaging | Dual-die chiplet |
| Transistor Count | Not disclosed |
| Memory | 196 GB HBM3 |
| Memory Bandwidth | 3.35 TB/s |
| FP16/BF16 | 700+ TFLOPS |
| INT4 | 2,800+ TOPS (web reports; INT8 not disclosed) |
| FP32 | Not disclosed (estimated ~175 TFLOPS) |
| TDP | 400–500 W (web consensus 400W+) |
| Interconnect | MLU-Link 890 Gbps / card |
| PCIe | PCIe 5.0 |
| Announced | 2025 |
| Mass Production | 2026-Q1 |
| Unit Price | ¥140,000 - 150,000 |
MLU690 vs MLU590 vs H100
| Metric | MLU690 | MLU590 | H100 SXM | Gap (vs H100) |
|---|---|---|---|---|
| Process | TSMC 4nm (dual die) | 7nm chiplet | TSMC 4nm | On par |
| Memory | 196GB HBM3 | 96GB HBM2e | 80GB HBM3 | +145% |
| Memory Bandwidth | 3.35 TB/s | ~2 TB/s | 3.35 TB/s | On par |
| FP16 | 700+ TFLOPS | 345 TFLOPS | 989 TFLOPS | ~71% |
| INT8 | 2,800+ TOPS | 1,380 TOPS | 1,979 TOPS | +41% |
| TDP | ~500W | ~350W | 700W | -29% |
| Interconnect | MLU-Link 890Gbps | MLU-Link | NVLink 4 900GB/s | Nearly on par |
| Software ecosystem | NeuWare | NeuWare | CUDA | Gap remains |
Key insight: the MLU690's memory capacity is 2.45x that of the H100, and its INT8 inference compute surpasses the H100 by 41%. Its shortcomings mainly come from FP16 training compute and interconnect efficiency in distributed training beyond a thousand cards.
Process and Supply Chain
- Foundry: after TSMC halted supply due to the Entity List, production shifted entirely to SMIC N+2 (5nm-class). The MLU690 is one of the largest customers for SMIC N+2
- Yield: Sigmaintell estimates SMIC's 5nm yield at about 34% in 2025Q1, with 40%+ possible by Q4
- Capacity booking: N+2 capacity is locked in through 2027
- Next generation: tentatively the MLU790 / MLU690 successor, expected to ship in 2026Q4 and ramp in 2027
Smart Driving: MLU610
Cambricon has also launched the MLU610 chip for smart driving scenarios, which has won orders from BYD, Li Auto, and others, expanding into the edge/automotive AI market.
NeuWare Software Stack
| Layer | Tool | Description |
|---|---|---|
| AI frameworks | PyTorch-Cambricon | PyTorch adaptation |
| TensorFlow-Cambricon | TensorFlow adaptation | |
| Compiler | NeuWare CC | nvcc-like |
| Runtime | NeuWare Runtime | CUDA Runtime-like |
| Math library | NeuBLAS | cuBLAS-like |
| Deep learning library | NeuDNN | cuDNN-like |
| Inference engine | MagicMind | Among the world's first commercial MLIR graph-compiler inference frameworks |
| Communication library | MLU-Link | Multi-card interconnect |
⚠️ Ecosystem limitations: NeuWare unifies cloud, edge, and endpoint with a single codebase — the only vendor in China doing so. However, distributed training beyond a thousand cards suffers relatively high losses (MLU-Link 890Gbps is close to NVLink 4's 900GB/s, but engineering maturity still lags); single-server 8-card and ~100-card small clusters perform excellently.
Vendor Information
| Item | Details |
|---|---|
| Company | Cambricon Technologies Corporation Limited |
| Stock Code | 688256.SH (STAR Market) |
| Founder | Chen Tianshi (Institute of Computing Technology, Chinese Academy of Sciences) |
| Founded | 2016-03 |
| IPO | 2020-07-20 (first AI chip stock on the STAR Market) |
| 2026Q1 Revenue | ¥2.885 billion (+160% YoY) |
| Market Cap | Over RMB 1 trillion (June 2026) |
| Headquarters | Haidian District, Beijing |
| Official Website | https://www.cambricon.com |
| Key Customers | ByteDance (largest customer, ~65% of purchases), Alibaba, Baidu, China Mobile |
Related Products
- Huawei Ascend 910C - Chinese full-stack training
- Moore Threads MTT S5000 - Chinese full-function GPU
- NVIDIA H100 - International flagship (subject to export controls)
- NVIDIA H200 - H100 upgrade
- Full comparison table