Skip to main content

Moore Threads MTT S4000 (2023)

Product Overview

MTT S4000 is Moore Threads' large model AI computing accelerator released in December 2023, based on self-developed Qiyuan GPU architecture (third-generation MUSA core architecture), equipped with 48GB GDDR6 memory (bandwidth 768 GB/s), FP32 compute 25 TFLOPS, TF32 compute 50 TFLOPS, INT8 compute 200 TOPS, customized optimization for training, fine-tuning and inference of 100-billion-parameter large language models, combined with advanced graphics rendering capabilities, video encoding/decoding capabilities and ultra-high-definition 8K HDR display output.

Positioning: Full-function meta-computing card (training+inference integration + graphics rendering), core component of KUAE AI computing center solution.

Core Specifications

ItemParameter
ArchitectureSelf-developed Qiyuan GPU (third-generation MUSA core architecture)
ProcessNot disclosed (estimated 7nm/6nm)
FP3225 TFLOPS
TF3250 TFLOPS
INT8200 TOPS
FP16/BF16Supported (specific values not disclosed)
Memory Capacity48 GB GDDR6
Memory Bandwidth768 GB/s
TDP450 W
InterconnectMTLink (x8 Serdes, up to 56Gbps PAM4)
InterfacePCIe 5.0 x16, 4× DisplayPort
PowerCPU 8-pin × 1
ReleaseDecember 2023
Mass ProductionSince 2024
Software StackMUSA software stack (CUDA compatible)

MUSA Architecture Evolution

ArchitectureCoreRepresentative ProductRelease
First-generation MUSAChunxiaoMTT S80/S70 (consumer)2022
Second-generation MUSAQuyuan (improved)MTT S30002023
Third-generation MUSAQiyuan GPUMTT S40002023.12

Comparison with MTT S3000

MetricMTT S3000MTT S4000Improvement
ArchitectureSecond-generation MUSAThird-generation MUSA (Qiyuan GPU)New generation
MemoryNot disclosed48GB GDDR6Larger
BandwidthNot disclosed768 GB/sHigher
FP32Not disclosed25 TFLOPSValue disclosed
TDPNot disclosed450WData center grade
Release20232023.12Same period improvement

KUAE AI Computing Center Solution

MTT S4000 is the core component of Moore Threads' KUAE AI computing center solution:

  • 100-billion-parameter large model training, fine-tuning, inference full-stack support
  • MTLink multi-card high-speed interconnect (x8 Serdes, 56Gbps PAM4)
  • MUSA software stack fully supports PyTorch/DeepSpeed and other mainstream frameworks
  • CUDA compatibility layer, reducing model migration costs

Application Scenarios

  • 100-billion-parameter large model training (customized optimization)
  • Large model inference as a service (INT8 200 TOPS)
  • Graphics rendering + AI hybrid workloads (full-function GPU)
  • Video encoding/decoding (8K HDR display output)
  • Domestic AI computing center (KUAE solution)
  • Ultra-high FP16 training compute (25 TFLOPS FP32 lower than H100)
  • Ultra-large-scale clusters (MTLink TBD vs NVLink)

Product Matrix

SeriesPositioningRepresentative Product
MTT S SeriesServer GPU (data center)S3000, S4000, S5000
MTT S Series (Consumer)Desktop GPUS80, S70
KUAEAI computing center solutionS4000 + MTLink + MUSA software stack

References