Skip to main content

MetaX XiYun C500 (2022)

Product Overview

XiYun C500 is MetaX Integrated Circuit's first training-inference integrated general-purpose GPU released in 2022, based on self-developed XCORE 1.0 architecture, equipped with 64GB HBM2e memory, supporting FP64/FP32/TF32/FP16/BF16/INT8 multi-precision mixed computing, FP16 compute 280 TFLOPS, INT8 compute 560 TOPS, interface supports PCIe Gen5 and MetaXLink multi-card interconnection, is the first product of MetaX's "XiYun" C series.

Positioning: Training-inference integrated GPU, balancing AI training and inference scenarios, performance better than NVIDIA H20 (according to third-party benchmarks).

Core Specifications

ItemParameter
ArchitectureSelf-developed XCORE 1.0 (dozens of core IPs)
ProcessNot disclosed (estimated 7nm)
FP3254 TFLOPS (vector 18 + matrix 36)
TF32140 TFLOPS
FP16280 TFLOPS
BF16280 TFLOPS
INT8560 TOPS
Memory Capacity64 GB HBM2e
Memory BandwidthNot disclosed (estimated ~1.6 TB/s)
TDP350 W (estimated)
InterconnectMetaXLink (7 high-speed interconnect ports, up to 64 cards interconnection)
InterfacePCIe Gen5 + MetaXLink
FP64 Support✅ (scientific computing/meteorological prediction)
Release2022
Mass ProductionSince 2023
Software StackMXMACA (CUDA compatible, migration cost reduced by 90%)

⚠️ Specification Note: Process, TDP, and memory bandwidth not fully disclosed by official sources, subject to MetaX's subsequent official data sheet.

XiYun C Series Product Line

ProductArchitectureMemoryFP16 TFLOPSReleaseStatus
XiYun C500XCORE 1.064GB HBM2e280 TFLOPS2022On sale
XiYun C550XCORE 1.xNot disclosedNot disclosed2024On sale
XiYun C588XCORE 1.xNot disclosedNot disclosed2024+On sale
XiYun C600XCORE 1.5144GB HBM3eFP8 1000 TFLOPS2025Risk production

Comparison with NVIDIA H20

MetricXiYun C500NVIDIA H20Difference
FP16280 TFLOPS~300 TFLOPS-7% (close)
INT8560 TOPS~600 TOPS-7% (close)
Memory64GB HBM2e96GB HBM3-33%
InterconnectMetaXLinkNVLinkTBD
EcosystemMXMACA (CUDA compatible)CUDAH20 mature
Price~¥38,900/card~¥200,000/cardC500 80% cheaper

Third-party benchmark: According to public information, XiYun C500 series training-inference integrated GPU performancebetter than H20.

MXMACA Software Ecosystem

LayerToolDescription
Software StackMXMACAMetaX unified computing architecture
AI FrameworkPyTorchNative support
DistributedDeepSpeedDistributed training
CUDA CompatibilityAutomatic migration toolCode migration cost reduced by 90%+
Large ModelSupports domestic thousand-card clusterFull parameter training verified

Application Scenarios

  • Domestic large model training (280 TFLOPS FP16, 64GB memory)
  • AI inference as a service (560 TOPS INT8)
  • Scientific computing (FP64 double-precision support)
  • Meteorological prediction (HPC traditional scenarios)
  • Domestic AI computing center (cost-performance advantage)
  • CUDA migration scenarios (90%+ migration cost reduction)
  • FP8 inference (no direct FP8 format support)
  • Ultra-large-scale clusters (MetaXLink TBD vs NVLink)

References