Skip to main content

3 posts tagged with "Domestic Substitution"

AI chip domestic substitution process

View all tags

Domestic Big Three 2026 H2: Localization Rate Crosses 40% Toward 60%, Ascend 960 Roadmap, MLU690 and S5000 Ecosystems Ramp Up

· 6 min read
Industry Research Team

In 2026, China's AI chip market landscape has shifted from "NVIDIA unipolar dominance" to "overseas vendors leading, domestic multi-route catch-up." According to industry research, China's overall AI accelerator market was ~4M units in 2025, of which 1.65M were domestic, with share first breaking 40%; as products iterate and fabs follow up, the localization rate is expected to rise to 60%-70% by 2027. This article focuses on the latest H2 2026 progress of Huawei Ascend, Cambricon, and Moore Threads — the domestic "Big Three."


1. Huawei Ascend: 950 Capacity Fully Booked, 960 Roadmap Unveiled

Ascend's core advantage is "architecture + full-stack ecosystem synergy," with ~800K units shipped in 2025, capturing 50% of the total domestic vendor share. The product iteration cadence is clear:

TimeProductNote
2025 Q1Ascend 910CMain transitional model
2026 Q1Ascend 950PRInference flagship
2026 Q4 (planned)Ascend 950DTTraining flagship, drives domestic HBM iteration
2027-2028Ascend 960 / 970Roadmap products

950 series capacity has entered a "fully booked" state: 950PR entered mass production in April 2026; June monthly capacity jumped to 500K-600K units (nearly 10x MoM), with a full-year target of 1.2M units at 100% certainty; ByteDance locked in 350K units for $5.6B, while Tencent / Alibaba / Baidu combined locked in 400K units.

Ascend 960 roadmap specs (per roadmap disclosure):

MetricAscend 960
ArchitectureAscend 6th gen (Da Vinci v6)
FP8 compute~4 PFLOPS
Memory288GB
Memory bandwidth9.6 TB/s
Super-nodeAtlas 960 SuperPoD, 15,488 cards, Lingqu optical-electrical converged bus
Debut2027 Q4 (roadmap)

The previous-gen Ascend 384 super-node has cumulatively shipped over 750 sets, deployed across 20+ industries including internet, operators, finance, education, and healthcare — Huawei calls it "the only domestic super-node that has trained a SOTA model."


2. Cambricon MLU690: H2 Mass Production, Entering ByteDance Bidding Window

Cambricon is the core domestic compute leader in the absence of an Ascend IPO, with the technology gap continuously narrowing:

  • Siyuan 590 (7nm): Performance equivalent to 80% of A100, already supports DeepSeek, continuously adapting to mainstream large models like Qwen 3 and GLM
  • Siyuan 690 series: Will enter mass production in H2 2026, expected to achieve order scale-up during ByteDance's H2 bidding window
  • Revenue certainty: Equity incentive targets show >100% revenue growth for the next 3 years: 2026 revenue target 13.5B RMB, 2027 27B RMB, 2028 60B RMB

Cambricon fully benefits from the industry dividend of "domestic CSP capex + full adaptation of domestic large models and domestic chips," making it the most direct elasticity play on rising localization rate.


3. Moore Threads MTT S5000: Full-Function GPU + Ecosystem Breakthrough

Moore Threads takes a differentiated "full-function GPU" route, with the flagship MTT S5000 based on the 4th-gen "Pinghu" MUSA architecture:

MetricMTT S5000
Dense AI compute1000 TFLOPS
Memory80GB
Memory bandwidth1.6 TB/s
Inter-card interconnect784 GB/s
PrecisionFP8 to FP64 full precision (training + inference)
SecurityFirst batch to pass national "Safe and Reliable Evaluation" (Level I)

Its engineering capability is verified: the Kuae (KUAE) intelligent computing cluster based on S5000 achieves 95% training linear scaling efficiency, with compute efficiency loss within 5% at ten-thousand-card scale; supports checkpoint-resume training with effective training time ratio >90%; and has trained a MoE-236B base model with >25 trillion tokens of corpus from scratch.

The ecosystem is Moore Threads' deepest moat: MUSA has achieved 100% core math library compatibility, 3000+ PyTorch operator compatibility, covers 55 categories of core AI operators, has official vLLM and SGLang support, Day-0 adaptation of mainstream models, and 800K+ developers. Its PD heterogeneous-disaggregation solution achieves equivalent replacement of international high-end GPUs at a 2:1 ratio with S5000, significantly reducing inference cost.

The 5th-gen "Huagang" architecture (released 2025-12) supports FP4 to FP64 full precision, with 50% higher compute density and 10x better energy efficiency than the previous gen, supporting 100K+ card clusters; cumulative R&D investment in the "Huashan" (train-infer integrated) and "Lushan" (graphics rendering) new chips based on this architecture exceeds 900M RMB.


4. Software Ecosystem Decides: Day-0 Adaptation Becomes Routine

Beyond hardware, software ecosystem realization is the watershed for domestic compute in 2026:

  • Huawei's CANN heterogeneous computing architecture and MindSeries suite are fully open-sourced, with the community incubating 67 projects, 12.44M+ lines of code, and 3,500+ monthly active developers
  • The "release-and-adapt" closed loop between domestic large models and domestic chips has basically formed: Tencent Hunyuan T3 (295B), DeepSeek-V4, and GLM-5.2 all completed Day-0 adaptation
  • 2026 is regarded as the "first year of domestic super-nodes"; Huatai Securities estimates China's super-node architecture market will reach 341.4B RMB by 2028, with a 2026-2028 CAGR of 194%

5. Industry Judgment: From "Can It Be Built" to "Can It Be Used Well"

The domestic Big Three are converging along three paths:

  1. Huawei: Locks government/enterprise and internet big customers with super-node system-level capability + full-stack software
  2. Cambricon: Impacts the revenue inflection point by narrowing the training-side gap + scaling up via big-customer bidding
  3. Moore Threads: Covers cloud-edge-end full scenarios with full-function GPU generality + mature CUDA-compatible ecosystem

The common shortcoming of all three remains advanced process and HBM supply — precisely the core link of overseas controls. But as domestic HBM iterates and fabs follow up, a realistic path to 60%-70% localization by 2027 exists.

References


This article is compiled from public industry research, broker views, and corporate announcements as of August 2026. Some shipment and market-share figures are third-party estimates, not officially confirmed data.

Huawei Ascend 950 Series Capacity & Orders Deep Dive: 950PR Monthly Capacity Jumps 10×, ByteDance Locks In 350k Units for $5.6B

· 4 min read
Industry Research Team

The Ascend 950 series (950PR inference / 950DT training) has become the core supply of domestic AI compute. Per multiple brokerages and industry research, 950 series capacity is 100% booked with scarce spot supply; the full-year 1.2M-unit target is "100% certain," with expectations of an upward revision to 1.5M. This article summarizes capacity and order data as of July 2026.

1. Capacity pace: ~10× MoM jump in June

Time950PR monthly capacityNotes
May 202650k-60k unitsNear full production
June 2026500k-600k units~10× MoM; SMIC, Hua Hong tier-1 suppliers on overtime
Q3 2026 (est.)700k-800k unitsPer month
Full-year 2026 target1.2M unitsUpward revision to 1.5M expected

Supply chain delivery is tight: high-speed backplanes and liquid-cooling connectors' lead time stretched from 2 weeks to 6-8 weeks; orders are booked into 2027.

2. Order structure: top cloud providers + operators + overseas

CustomerLocked volumeAmount / Notes
ByteDance350k 950PR$5.6B, concentrated delivery from Q3 2026
Tencent / Alibaba / Baidu~250k 950PR + 150k 950DTCombined ~400k units
Three major operators200k+ unitsCentralized procurement, for intelligent compute centers and AI private networks
OverseasSouth Korea 2,000 units, Malaysia 3,000 servers, Russia ten-thousand-card clusterFrom pilot to commercial

3. Shipment forecast: firmly #1 domestic

Per CCA (Kezhi) Consulting estimates:

Metric20252026 (forecast)
Huawei Ascend total shipments812k cards1.026M cards
Of which 950PR~800k units
Of which 950DT~100k-200k units

Huawei has completed the product transition from the 910 series to the 950 series. The internet industry has become Ascend's largest application market; competitive advantage is extending from single-hardware performance to software ecosystem and system capabilities.

4. Going overseas: formal South Korea entry in Q4

Per Korean media ETNews, Huawei plans Q4 2026 to formally enter the South Korean market with the Ascend series and Atlas 950 SuperPod:

  • Local distributor agreements signed; two channel partners including SK Shieldus selected
  • Main products: 950PR (mass-produced and delivered since April) and 950DT (launched Q4)
  • Official line: 950PR inference performance is 2.87× that of H20, priced at about 1/4 of it

5. WAIC 2026: 1024-card live debut confirmed

At WAIC 2026 (July 17-20, Shanghai), Huawei's Atlas 950 SuperPoD live hardware made its first public appearance — a 16 compute-cabinet, 1,024 Ascend-card scale — and won the conference's top honor, the SAIL Award:

  • Core metrics: total compute 1 EFLOPS FP8 / 2 EFLOPS FP4, 256 TB globally unified memory addressing, Lingqu 2.0 interconnect, 3 μs ultra-low RTT latency
  • Full configuration: 128 compute cabinets + 32 interconnect cabinets = 160 cabinets, ~1000㎡ footprint, carrying 8,192 Ascend 950DT, planned Q4 2026 launch
  • Commercial foundation: previous-gen 384 SuperNode has cumulatively shipped 750+ units, deployed in 20+ industries
  • Software ecosystem: CANN fully open-sourced end of 2025; community incubated 67 projects, 12.44M+ lines of code, 3,500+ monthly active developers

WAIC's debut confirmed the 950 series' "SuperNode-first" product logic: beyond single-card compute, system-level effective compute (interconnect bandwidth + unified memory + low latency) is the key dimension for domestic compute to benchmark against international flagships.

Ascend roadmap recap

ProductPositioningKey metrics (official roadmap)
950PRInference1 PFLOPS (FP8) / 2 PFLOPS (FP4), 2 TB/s interconnect
950DTTrainingSuperNode core, launched Q4
960Train/inference2 PFLOPS (FP8) / 4 PFLOPS
970Next-genIn planning

Industry interpretation

  1. Domestic substitution moves from inference to training: 950PR (inference) ramps first, 950DT (training) follows in Q4, combined with the Atlas 950 SuperPoD ten-thousand-card interconnect — domestic compute now has the complete "training substitution" puzzle for the first time.
  2. Capacity is the biggest variable: order certainty is extremely high, but SMIC/Hua Hong advanced-process capacity, HBM supply, and advanced packaging remain ramp bottlenecks — the root of "scarce spot supply."
  3. Going overseas opens a second growth curve: bulk procurement from South Korea, Malaysia, Russia, and Latin America marks domestic compute's shift from "internal circulation" to "external circulation."

References


Data in this article is based on official and major brokerage research; capacity/orders are dynamic figures and will be continuously updated.

China AI Chip Landscape 2025: Ascend, Cambricon, Hygon — Who Will Dominate?

· 5 min read
Industry Research Team

Escalating U.S. export controls are forcing China's AI chip industry to accelerate self-reliance. By 2025, the discussion around domestic Chinese AI chips has shifted from "are they usable?" to "which one should I choose?"

This article systematically reviews the major players, core products, and actual deployment status of domestic AI chips, helping developers and procurement decision-makers understand the competitive landscape.