Skip to main content

Enflame CloudBlaze S60

Enflame's third-generation AI inference accelerator, based on self-developed GCU320 (Suisi 320) chip, released in March 2024, typical single-card power consumption ~300W, targeting large-scale data center deployment.


Core Specifications

SpecificationValue
ArchitectureGCU320 (Suisi 320)
Process7 nm (estimated)
TDP300 W (typical)
Memory48 GB HBM2e (estimated)
Memory Bandwidth1.6 TB/s (estimated)
FP32 Compute25 TFLOPS (estimated)
FP16 / BF16 Compute100 TFLOPS (estimated)
INT8 Compute200 TOPS (estimated)
InterfacePCIe 5.0 x16, full-height full-length
Release2024-03
PriceNot disclosed (estimated ¥35,000)

Technical Highlights

  • Third-generation inference card: Based on Suisi 320 (GCU320) chip, Enflame's third-generation AI inference product
  • Large model inference optimization: Optimized for LLAMA, GPT and other large language model inference scenarios
  • Search/ad/recommendation support: Supports high-concurrency inference scenarios for search, advertising, and recommendation systems
  • Easy migration: Broad model coverage, strong usability, supports smooth migration from NVIDIA GPU
  • High-density deployment: Typical power consumption 300W, suitable for large-scale data center deployment

Product Positioning

CloudBlaze S60 is Enflame's new-generation AI inference accelerator for large-scale data center deployment, targeting NVIDIA L4/L40. As the third-generation product, S60 has significant improvements in compute, memory, bandwidth and other aspects.


Application Scenarios

  • Large language model inference (LLAMA, GPT, ChatGLM, etc.)
  • Search, advertising, recommendation system inference
  • Computer vision inference (CV)
  • Natural language processing inference (NLP)
  • Data center large-scale inference deployment

Reference Price

ChannelPriceDescription
Official pricingNot disclosedReleased in 2024, estimated ¥30,000–40,000/card
Channel estimate≈ ¥35,000Estimated based on L4 pricing ratio


References