Skip to main content

Iluvatar TianGai 100 (BI-V100)

Product Overview

TianGai 100 (model BI-V100) is Iluvatar's first domestic fully self-developed general-purpose GPU training accelerator officially released in March 2021, using 7nm process, equipped with 32GB HBM2 memory, board-level power consumption 250W, supporting FP32/FP16/INT8 mixed-precision training, compatible with international mainstream GPU general computing models, supporting PyTorch/TensorFlow and other mainstream AI frameworks, is the founding work of Iluvatar's "TianGai" training product series.

Strategic Significance: Achieved a major breakthrough from 0 to 1 for domestic general-purpose GPU products, filling the gap in domestic cloud training GPUs.

Core Specifications

ItemParameter
ArchitectureIluvatar self-developed general-purpose GPU architecture
Process7nm (estimated TSMC)
FP16128 TFLOPS
INT8256 TOPS
INT3232 TFLOPS
Memory Capacity32 GB HBM2
Memory BandwidthNot disclosed (estimated ~1 TB/s)
TDP250 W (board-level power consumption)
InterfacePCIe Gen4.0 x16
Inter-chip Interconnect64 GB/s bidirectional bandwidth
CoolingPassive cooling
DimensionsFull-length full-height dual-slot PCIe card
ReleaseMarch 2021
Mass ProductionSince 2021
Software StackIluvatar computing software stack (PyTorch/TensorFlow compatible)

⚠️ Specification Note: Memory bandwidth not fully disclosed by official sources, subject to Iluvatar's subsequent official data sheet.

TianGai Series Product Line

ProductReleaseFP16 TFLOPSMemoryStatus
TianGai 100 (BI-V100)2021128 TFLOPS32GB HBM2On sale
TianGai 150 (BI-V150)2023Not disclosed (estimated higher)Not disclosedOn sale
TongYang TY10002024+Not disclosedNot disclosedNext generation

Software Ecosystem

LayerToolDescription
AI FrameworkPyTorch / TensorFlowNative compatibility
Programming LanguageCUDA C++ / OpenCLSupports mainstream programming models
Operator LibraryIluvatar computing software stackNative operators + custom operators
Cluster TrainingSupports distributed trainingMulti-card interconnect

Application Scenarios

  • Domestic large model training (below 100 billion parameters)
  • AI framework migration (CUDA programming model compatible)
  • Government/state-owned enterprise AI projects (supply chain security)
  • Scientific computing (FP32/INT32 support)
  • Ultra-high compute requirements (FP16 128 TFLOPS lower than H100)
  • Emerging FP8 precision (FP8 not supported)

Comparison with ZhiKai 100 (MR100)

MetricTianGai 100 (BI-V100)ZhiKai 100 (MR100)Difference
PositioningTraining (Training)Inference (Inference)Different scenarios
FP16128 TFLOPS96 TFLOPSTianGai stronger
INT8256 TOPS192 TOPSTianGai stronger
Video DecodingNot supported128-channel 1080PZhiKai exclusive
TDP250WEstimated 250-300WSimilar

References