Skip to main content

One post tagged with "comparison"

View all tags

Cambricon MLU690 vs NVIDIA H100: In-Depth Comparison — Can a Domestic AI Chip Replace the H100?

· 6 min read
AI Hardware Analyst

In 2026, against the backdrop of U.S. export controls on AI chips to China, Cambricon's MLU690 has drawn intense attention as a "China-made H100." This article compares the two in depth across compute, memory, power, software ecosystem, measured performance, and price to help you make a selection decision.

Core Verdict (Read This First)

DimensionMLU690H100WinnerGap
BF16 compute600 TFLOPS989 TFLOPSH100+65%
Memory capacity64GB HBM380GB HBM3H100+25%
Memory bandwidth2 TB/s3.35 TB/sH100+68%
TDP280W700WMLU690-60%
Energy efficiency2.14 TFLOPS/W1.41 TFLOPS/WMLU690+52%
Software ecosystemNeuWare (~75% coverage)CUDA (100% coverage)H100large gap
Price~¥140,000~¥200,000MLU690-30%
Availabilitydomestic spot stockexport-controlledMLU690

One-line summary: MLU690 delivers roughly 60% of H100's compute, but at only 40% of the power and 70% of the price — a strong fit for AI training and inference in the Chinese market.


1. Detailed Spec Comparison

1.1 Compute

PrecisionMLU690H100 SXM5H200 SXM5Note
FP8~300 TFLOPS (est.)3,958 TFLOPS3,958 TFLOPSH100 supports FP8; MLU690 likely does not
BF16/FP16600 TFLOPS989 TFLOPS989 TFLOPSH100 leads by 65%
FP32~150 TFLOPS (est.)60 TFLOPS60 TFLOPSMLU690 estimate; H100 actually higher
INT81,200 TOPS1,979 TOPS1,979 TOPSH100 leads by 65%

Key findings:

  • ✅ MLU690 reaches 60% of H100's BF16 compute
  • ⚠️ H100 supports FP8 (4-bit); MLU690 likely does not (needs confirmation)
  • ⚠️ H100's higher INT8 compute favors inference scenarios

1.2 Memory

ItemMLU690H100H200Note
Capacity64GB HBM380GB HBM3141GB HBM3eH200 largest
Bandwidth2 TB/s3.35 TB/s4.8 TB/sH200 highest
TypeHBM3HBM3HBM3eH200 uses latest HBM3e

Key findings:

  • ⚠️ MLU690 has 20% less memory than H100 (64GB vs 80GB)
  • ⚠️ MLU690 bandwidth is 40% lower than H100 (2 TB/s vs 3.35 TB/s)
  • ❌ When running 70B+ parameter models, MLU690 may run out of memory (model parallelism required)

1.3 Power

ItemMLU690H100H200
TDP280W700W700W
Efficiency (FP16/W)2.14 TFLOPS/W1.41 TFLOPS/W1.41 TFLOPS/W
8-card server power~3.5kW~6kW~6kW
Annual electricity (¥0.6/kWh)~¥18,400~¥36,800~¥36,800

Key findings:

  • MLU690 draws only 40% of H100's power, sharply cutting data-center electricity cost
  • MLU690 leads efficiency by 52%, better suited to large-scale deployment
  • ✅ For power-sensitive inference, MLU690 has a clear edge

2. Software Ecosystem

2.1 Framework Support

FrameworkMLU690 (NeuWare)H100 (CUDA)Note
PyTorch✅ (PyTorch-Cambricon)✅ nativeMLU690 needs an extra plugin
TensorFlow✅ (TensorFlow-Cambricon)✅ nativesame
JAX⚠️ partial✅ nativeMLU690 limited
ONNX⚠️ partial✅ nativesame
vLLM⚠️ in progress✅ nativeMLU690 awaits community port

2.2 Operator Coverage

CategoryMLU690H100Note
Basic operators✅ 95%✅ 100%conv, matmul, etc.
Transformer operators✅ 85%✅ 100%Attention, LayerNorm, etc.
Custom operators⚠️ hand-written✅ CUDA C++MLU690 harder to develop
LLM inference opt.⚠️ basic✅ mature (FlashAttention, PagedAttention)H100 leads

Key findings:

  • ⚠️ NeuWare is only 5–6 years old, with ~75–85% operator coverage
  • ❌ Complex LLMs (e.g., GPT-4, Claude) may need manual optimization
  • ✅ Common models (Llama, Qwen, GLM) are essentially already supported

3. Measured Performance

3.1 Training

ModelMLU690 (time)H100 (time)Speedup
Llama 7B~48 h (est.)~30 h1.6x
Llama 70B~7 days (est.)~4.5 days1.6x
Qwen 72B~8 days (est.)~5 days1.6x

Note: above figures are estimates; real performance depends on software optimization.

3.2 Inference

ModelMLU690 (tok/s)H100 (tok/s)Note
Llama 7B~80 tok/s (est.)~120 tok/sH100 +50%
Llama 70B~20 tok/s (est.)~35 tok/sH100 +75%
Qwen 72B~18 tok/s (est.)~30 tok/sH100 +67%

Key findings:

  • ⚠️ H100 leads inference by 50–75%
  • ✅ But MLU690 draws only 40% the power, with better efficiency
  • ✅ For cost-sensitive inference, MLU690 is more economical

4. Price

4.1 Hardware Procurement

ItemMLU690H100H200
Per-card (domestic)~¥140,000~¥200,000~¥300,000
8-card server (turnkey)~¥1,200,000~¥1,800,000~¥2,600,000
Cost gap-+50%+117%

4.2 TCO (3 years)

ItemMLU690H100Note
Hardware¥1,200,000¥1,800,000MLU690 33% cheaper
Electricity (3y)¥55,200¥110,400MLU690 50% cheaper
Facility¥150,000¥250,000MLU690 40% cheaper
TCO (3y)¥1,405,200¥2,160,400MLU690 35% cheaper

Key findings:

  • MLU690's TCO is 35% lower than H100's
  • ✅ For large-scale deployment (100+ cards), the cost advantage is pronounced

5. Selection Advice

5.1 Choose MLU690 if...

  • ✅ Your business is primarily in the Chinese market
  • ✅ You are affected by U.S. export controls and cannot buy H100/H200
  • ✅ You are power-sensitive (edge data centers, high electricity-cost regions)
  • ✅ Your models use common architectures (Llama, Qwen, GLM)
  • ✅ You have domestic-substitution requirements (government, SOEs, military)

5.2 Choose H100/H200 if...

  • ✅ Your business is global
  • ✅ You need to train frontier models (GPT-4 class)
  • ✅ Your models use complex operators (need the CUDA ecosystem)
  • ✅ You demand extreme performance (low-latency inference)
  • ✅ You can legally procure H100/H200
ScenarioRecommended
TrainingH100 (high perf) + MLU690 (low-cost scale-out)
InferenceMLU690 (cost-sensitive) + H100 (low-latency)
Domestic projectall MLU690
International marketall H100/H200

6. Outlook

6.1 MLU690's weaknesses

  • ⚠️ Immature software ecosystem: 75–85% operator coverage; complex models need manual tuning
  • ⚠️ Small memory: 64GB limits support for 70B+ parameter models
  • ⚠️ Weak interconnect: Cambricon Link bandwidth below NVLink
  • ⚠️ Limited international market: affected by U.S. export controls

6.2 MLU690's improvement path

  • 📅 MLU790 (2027): expected 5nm process, ~2x compute
  • 📅 Memory upgrade: next gen may adopt HBM3e, capacity up to 128GB
  • 📅 Software: NeuWare ecosystem improving, operator coverage target 95%

7. Summary

DimensionMLU690H100Recommended scenario
Compute⭐⭐⭐⭐⭐⭐⭐⭐⭐H100 for top-tier training
Memory⭐⭐⭐⭐⭐⭐⭐H100 for large models
Power⭐⭐⭐⭐⭐⭐⭐⭐MLU690 for inference
Ecosystem⭐⭐⭐⭐⭐⭐⭐⭐H100 for complex models
Price⭐⭐⭐⭐⭐⭐⭐⭐MLU690 for large-scale deployment
Domestic⭐⭐⭐⭐⭐MLU690 for Chinese market

Final recommendation:

  • 🇨🇳 Chinese market: prefer MLU690 (domestic + low cost)
  • 🌍 International market: prefer H100/H200 (performance + ecosystem)
  • 💡 Hybrid: train on H100, infer on MLU690

References


Disclaimer: Data in this article is based on public sources and reasonable estimates; actual performance is subject to vendor official testing. MLU690's software ecosystem is evolving rapidly — watch NeuWare updates.

Last updated: 2026-06-23