Skip to main content

8 posts tagged with "Moore Threads"

Moore Threads MTT series GPUs and domestic GPU progress

View all tags

摩尔线程双线作战:MTT S5000 智算卡与"庐山"图形 GPU 的AI+游戏并举

· 6 min read
Industry Research Team

当多数国产 AI 芯片公司all in 大模型算力时,摩尔线程依然坚持"两条腿走路":一手是 MTT S5000 智算卡撑起的夸娥万卡集群,一手是瞄准 3A 游戏的"庐山"图形 GPU。9 月初的半年度业绩说明会给出了两条线的最新坐标——这条"AI+游戏"并举的路线,正在从权宜之计变成差异化壁垒。

一、MTT S5000:平湖架构的训推一体旗舰​

MTT S5000 基于第四代 MUSA"平湖"架构,搭载 PH100 芯片、4096 个自研 MUSA Core,2025 年 2 月公开参数。核心规格如下:

项目参数
架构MUSA 第四代(平湖),PH100 芯片
显存80 GB(颗粒类型官方未公开)
显存带宽1.6 TB/s
FP8 稠密算力1000 TFLOPS(液冷版)/ 920 TFLOPS(风冷版)
BF16/FP16400 TFLOPS
FP32100 TFLOPS
INT8800 TOPS
互联MTLink 784 GB/s(8 卡节点内全互联)
TDP500 W(市场口径)
形态OAM / PCIe,液冷 + 风冷双版本
单价~$4000–6000(OAM)

几点澄清值得注意:其一,"单卡 AI 稠密算力 1000 TFLOPS"特指 FP8 精度(液冷版),并非 FP16;其二,80GB 显存的颗粒类型官方从未公布,市场流传的"GDDR6X"说法与 1.6 TB/s 带宽自相矛盾,本站已订正为"类型未公开";其三,S5000 配备硬件级 FP8 Tensor Core,支持 FP8→FP64 全精度,并已基于 FP8 完成 DeepSeek-V4 的 Day-0 适配。加上《安全可靠测评结果公告(2026 年第 2 号)》的通过——AI 训推芯片首次纳入该体系——S5000 同时拿到了信创市场的入场券。

二、万卡落地:夸娥集群与京东云合作​

S5000 的成色,最终要靠集群说话。基于夸娥(KUAE)万卡智算集群的实测数据:

指标实测值
集群算力10 EFLOPS
Dense 大模型训练 MFU60%
MoE 大模型训练 MFU40%
训练线性扩展效率95%(64 卡→1024 卡保持 90%+)
有效训练时间占比>90%
第三方对齐智源 RoboBrain 2.5 与 H100 集群 loss 差异仅 0.62%

95% 的线性扩展效率背后是 ACE 异步通信引擎——把通信任务从计算核心卸载,实现计算与通信零冲突并行。商业侧,2026 年 3 月签下 6.6 亿元夸娥集群订单;软件侧,MiniMax M3、智谱 GLM-5.2、阿里 Qwen3.5、美团万亿参数 LongCat-2.0 相继完成 Day-0 适配。最关键的落地信号来自 9 月 19 日:京东云与摩尔线程合作的万卡级集群已部署上线服务——从"能跑通"到"规模化服务外部客户",国产 GPU 集群跨过了商用验证的分水岭。随后的数贸会(9 月 23–27 日,杭州)上,摩尔线程集中展示了 AI 智算卡、万卡级集群,以及"模型训练工厂、词元生产工厂、智能体生产工厂"三层能力,产品覆盖云-边-端全场景。

三、"庐山"GPU:从兼容显卡到渲染范式革命​

图形产品线上,9 月 3 日半年度业绩说明会披露了新一代"庐山"GPU 的指标——基于花港架构,预计 2026 年底推出:

  • 3A 游戏渲染性能提升 15 倍;
  • 光线追踪性能提升 50 倍;
  • AI 性能提升 64 倍。

比倍数更重要的是两项渲染技术的组合:一是硬件光线追踪,原生支持 DirectX Raytracing,让国产 GPU 第一次在光追管线层面与主流标准对话;二是AI 生成式渲染 AGR,基于全自研 MTAGR 1.0,把渲染范式从"逐像素计算"推向"生成"——用 AI 直接生成画面内容,而非暴力算出每个像素。前者补齐了传统管线的硬件短板,后者则是在渲染范式切换窗口期的卡位:当英伟达 DLSS、AMD FSR 都在走向 AI 化时,AGR 让摩尔线程有机会在下一代渲染标准里占据一席,而不是永远追赶。

四、游戏反哺 AI:双架构复用的独特路线​

把两条产品线放在一起看,摩尔线程的打法浮现出独特逻辑:图形与 AI 共享同一套 MUSA 架构底座——SIMT 核心、张量单元、光追核心、统一软件栈全部复用。游戏 GPU 的价值因此不只是营收 diversification:海量游戏场景是最好的架构压力测试与现金流来源,图形侧打磨的调度、显存与光速能力,反向反哺 AI 卡的通用计算质量;而 AI 业务积攒的先进制程流片经验,又拉高图形芯片的天花板。

已登陆科创板的摩尔线程,如今可以用上市平台的融资能力同时养两条产品线。这与纯 AI 芯片公司"单点押注大模型"的路线形成鲜明对比——在国产算力需求旺盛但波动性尚存的市场里,双线作战既是对冲,也是壁垒。

小结​

摩尔线程的 2026 年验证了"全功能 GPU"路线的可行性:MTT S5000 以 FP8 稠密 1000 TFLOPS、80GB 显存撑起夸娥万卡集群,并通过京东云万卡级集群部署完成了规模化商用验证;"庐山"GPU 则以硬件光追与 AI 生成式渲染 AGR 完成了从"国产兼容显卡"到"渲染范式探索者"的技术跃迁。图形+AI 双架构复用的独特路线,让摩尔线程在国产 AI 芯片阵营中占据了一个难以复制的生态位。2026 年底"庐山"如期推出与否,将是这条双线故事的下一次大考。

摩尔线程"庐山"GPU年底亮相:3A游戏渲染提升15倍、光追性能提升50倍

· 5 min read
Industry Research Team

从"国产兼容显卡"到"硬件光追+AI生成式渲染",摩尔线程正在完成一次技术范式层面的跃迁——而AI算力与游戏GPU的双线并进,让这次跃迁有了更厚实的商业底座。

"庐山"来了:15倍渲染、50倍光追、64倍AI​

2026年9月3日,摩尔线程在半年度业绩说明会上介绍了下一代高性能图形渲染芯片规划:基于**"花港"架构的新一代GPU——"庐山"**,预计实现三大性能跃升:

  • 3A游戏渲染性能提升15倍
  • 光线追踪性能提升50倍
  • AI性能提升64倍

官方计划于2026年底推出"庐山"芯片(上市时间以官方发布为准)。这三个数字放在一起,勾勒出的不是一次常规的代际迭代,而是一次"重写底座"式的升级:光追50倍意味着硬件光追从"支持"走向"实用",AI性能64倍则预示着AI能力将从图形附庸升级为核心卖点。

两大核心渲染技术​

"庐山"的性能跃升背后,是两项核心渲染技术支撑:

其一是实时光线追踪。"庐山"基于花港架构设计了硬件光线追踪加速引擎,并支持DirectX Raytracing(DXR)。此前国产GPU普遍通过软件手段或有限硬件支持来兼容光追,而专用加速引擎+DXR支持的组合,意味着"庐山"将具备与主流游戏引擎光追管线对接的能力——这是3A游戏生态的门票。

其二是AI生成式渲染AGR。摩尔线程推出全自研的MTAGR 1.0技术,推动渲染范式"从计算走向生成"。传统光栅化渲染是逐像素计算的暴力美学,而生成式渲染借助AI模型直接生成画面细节,用更少的算力换更高的画质与帧率。NVIDIA的DLSS、AMD的FSR已验证了AI超分路径的价值,MTAGR 1.0则是国产GPU在渲染范式变革上的主动跟进——且选择直接押注"生成"而非仅仅"超分"。

双线并进:游戏GPU与AI算力​

"庐山"的故事不止于游戏。摩尔线程2026年的商业化进展呈现出游戏与AI算力两条腿走路的格局:

时间事件意义
此前登陆科创板上市打开资本通道,支撑高研发投入
9月3日半年度业绩说明会公布"庐山"规划图形GPU技术路线图明确
9月19日(报道)京东云与摩尔线程合作万卡级集群部署上线国产GPU万卡集群规模化落地
9月23-27日数贸会(杭州)展示AI智算卡、万卡级集群展示"模型训练工厂/词元生产工厂/智能体生产工厂"

在杭州数贸会上,摩尔线程展示的产品覆盖云-边-端全场景:AI智算卡、万卡级集群,以及面向大模型生命周期的"模型训练工厂""词元生产工厂""智能体生产工厂"。这套叙事与京东云万卡集群的落地相互印证——摩尔线程的AI算力业务已经从"卖卡"进入"卖集群、卖工厂"的深水区。

技术跃迁的深层逻辑​

回顾国产GPU的发展路径,第一代产品的核心命题是"兼容"——让Linux桌面、办公软件、视频播放跑起来;而"庐山"的命题变成了"范式"——硬件光追对接主流3A管线、AGR参与全球渲染范式变革、AI性能面向大模型时代。从跟随式兼容到范式级参与,这是国产GPU厂商第一次在图形技术方向上与国际一线站在同一条起跑线附近。

AI性能提升64倍的目标也值得放在行业背景下解读:图形与AI本就同源,硬件光追引擎与AI算力单元共享大量基础能力。"庐山"把AI性能提升作为核心指标,说明花港架构在设计之初就为图形+AI双重负载做了统一规划。

小结​

"庐山"GPU承载着摩尔线程的三重跨越:技术上,从软件兼容走向硬件光追与MTAGR 1.0生成式渲染的范式创新;产品上,从单一显卡走向云-边-端全栈布局;商业上,从科创板上市到京东云万卡集群落地,AI算力与游戏GPU双线并进。3A渲染15倍、光追50倍、AI性能64倍的目标能否兑现,要等2026年底产品揭晓(以官方发布为准),但方向已经清晰:国产GPU不再满足于"能用",而是开始追求"领先"。

JD Cloud's 100,000-Card Cluster Bets on Moore Threads: Domestic GPUs Enter a Top AI Cloud's Core Compute Base for the First Time

· 6 min read
Industry Research Team

This article is based on official announcements from the 2026 JD Global Technology Explorer Conference (September 9) and public market information; order values and revenue forecasts are brokerage/media estimates, not company announcements.

On September 9, at the 2026 JD Global Technology Explorer Conference, JD Cloud announced a milestone decision: partnering with Moore Threads to build a 100,000-card full-function GPU cluster, creating hyperscale domestic intelligent computing infrastructure.

Two keywords in this sentence deserve amplification: "full-function GPU" and "100,000-card-class core cluster." The former means it must run not just inference but the front lines of large model training, inference, and embodied AI; the latter means a domestic GPU has, for the first time, been placed at the core of a top AI cloud provider's compute base — not a pilot, not an adaptation, not an all-in-one appliance, but 100,000 cards.

1. Partnership Details: From 10,000 to 100,000 Cards — How Big Is the Order?​

  • Prior foundation: JD Cloud has already built a domestic 10,000-card cluster with Moore Threads and other partners; this move is a magnitude leap from 10,000 to 100,000 cards
  • Order scale: According to market and brokerage information, GPU modules are supplied exclusively by Moore Threads, with a total order value of about RMB 20-30 billion; revenue can be recognized as early as next year according to delivery cadence. Some brokerages have raised Moore Threads' revenue expectation for next year to RMB 15-20 billion (this estimate is not a company announcement; refer to official disclosures)
  • Partnership depth: Full-stack coordination from chips and cloud platform to model training, supporting the iteration of JD's JoyAI model family and forming a closed loop of "data, training, simulation, deployment"
  • Openness: Compute is open to all industries, focused on large model training, inference, and embodied AI

Moore Threads founder Zhang Jianzhong was direct on stage: "The Scaling Law still holds — 100,000-card clusters are an inevitable trend." JD Cloud President Cao Peng positioned domestic compute as a core pillar of JD's physical AI strategy.

2. Why Moore Threads? Three Calculations Behind the Procurement Logic​

Tech companies buy cards with no sentiment involved. JD's choice of Moore Threads comes down to three calculations that all add up:

1. The stability calculation: Moore Threads has commercially deployed thousand-card and 10,000-card large clusters under a single network, and has achieved breakthroughs in core training scenarios such as foundation models, embodied brains, and world models — the engineering validation of a 10,000-card cluster is the prerequisite for 100,000 cards; this is not a cold start.

2. The integration calculation: Full-function GPUs can plug into existing IT systems and cloud platform scheduling, and the MUSA software stack's adaptation to large model frameworks has already passed training-grade workloads.

3. The ROI calculation: This is the most critical shift. Buyers of domestic GPUs used to be mostly "policy-friendly" projects; JD writing a 100,000-card cluster into its capital expenditure means the product's return on investment can now stand up to investor scrutiny — the buyer structure shifting from "daring to use" to "rushing to use" is the hallmark of commercial maturity for domestic GPUs.

3. S6000: Next-Gen Chip Taped Out and Back, with a Dual-Supply Safeguard​

According to market information, Moore Threads' next-generation chip, the S6000, has successfully come back from the fab and been distributed to vendors for testing, with ample FAB and memory supply guarantees. Note that the S6000 has not been officially released and its specifications are not public; this article makes no speculation. The on-sale flagship MTT S5000 has on-site data of 400 TFLOPS FP16 and 80GB of memory (the 1.6TB/s bandwidth is HBM-class).

Regarding HBM supply constraints, Moore Threads' response strategy is reportedly "next-generation product iteration + multi-source supply chain safeguards" — until domestic HBM capacity ramp-up is complete, this is the same problem every domestic GPU vendor must solve.

4. "100,000 Cards" Is Not One Company's Game: The Domestic GPU Cluster Landscape​

It is worth widening the view — 100,000 cards is now a collective goal for domestic compute:

Player100,000-Card MoveCompute Base
JD CloudAnnounced co-built 100,000-card cluster on September 9Moore Threads full-function GPU
Sugon 8000Released at WAIC in July, completed in Zhengzhou; a fully domestic 100,000-card AI compute clusterHygon DCU (of the Shensuan BW1000 family)
HuaweiAtlas 950 SuperPoD super node (8,192 cards); 100,000-card super node in 2027Ascend 950/960

Two chip routes (full-function GPU vs DCU vs NPU) and two paths (commercial cloud vs national supercomputing) point to the same validation question: can domestic compute reliably run real training workloads at 100,000-card scale. It is worth emphasizing that there is currently no public third-party benchmark comparison for the Moore Threads x JD cluster, and no completion timetable — going from "usable" to "running well" at 100,000 cards still requires engineering validation.

5. The Capital Markets Perspective​

Moore Threads listed on the STAR Market in December 2025: an IPO price of RMB 114.28, closing at RMB 600.5 on the first day; 2025 revenue of RMB 1.505 billion (up 243.37% year over year), with a net loss of about RMB 1.001 billion. If the brokerage-raised revenue expectation of RMB 15-20 billion materializes, 2027 will be the key inflection point from "high-growth loss-making" to "profitable at scale" — which would also become the first complete answer sheet for the domestic GPU business model.

Summary​

From debuting with DeepSeek all-in-one appliances in 2025 to entering a 100,000-card core cluster in 2026, domestic GPUs completed the cognitive leap from "usable" to "commercially viable at scale" in under two years. The value of JD Cloud's order lies not in its amount but in its significance as a sample: when top internet buyers begin procuring domestic GPUs on ROI logic, substitution is no longer a policy narrative but a commercial fact.

(Order values, revenue forecasts, and supply chain status come from public market information and brokerage research estimates; refer to company announcements; chip specification data is available in the on-site full comparison table.)

国产 AI 芯片半年报大检阅:寒武纪净赚 23 亿、壁仞营收暴涨 20 倍、燧原登板,"抢芯大战"白热化

· 6 min read
Industry Research Team

8 月底至 9 月初,国产 AI 芯片厂商半年报密集披露,加上燧原科技 9 月 2 日启动 IPO 申购,被称为"国产四小龙"的摩尔线程、沐曦、壁仞、燧原全部完成上市,加上早已在科创板的寒武纪——国产 AI 芯片的资本市场拼图就此补齐。更重要的是,财报数字第一次集体印证了一件事:国产算力正从"能用"跨向"好用",商业化拐点已现。


1. 半年报成绩单:增长是主旋律,盈利是分水岭​

厂商2026H1 营收同比盈利状态技术路线
寒武纪59.96 亿元+108.1%归母净利 23.11 亿自研 MLU 指令集(DSA)
摩尔线程17.36 亿元+147.4%净亏 1156 万(收窄)全功能 GPU(兼容 CUDA 路线)
壁仞科技12.36 亿元+1997.6%未盈利通用 GPU
沐曦13.24 亿元+44.7%净利 6.12 亿(首次扭亏)通用 GPU
燧原11.20 亿元+279.1%未盈利DSA 专用架构(TopsRider)

几个值得注意的细节:

  • 沐曦率先跨过盈利线:8 月 31 日披露的半年报显示净利润 6.12 亿元、同比扭亏(上年同期亏损 1.86 亿);不过扣非净利润仍为 -4900 万(亏损收窄 75.8%)——含金量仍在爬坡。
  • 壁仞低基数暴增:近 20 倍的同比增速来自上年同期极低的收入基数,但 12.36 亿的绝对体量已与燧原、沐曦同量级,第二梯队座次重新洗牌。
  • 研发强度惊人:沐曦研发费用占营收 39.7%、摩尔线程 44.3%、壁仞高达 65%——高研发投入是全员未完全盈利的根本原因,也是未来竞争力的来源。
  • 摩尔线程毛利率承压:从 78% 降至 57%,成本增速(+245%)远超营收增速(+147%)——大规模量产期的品控与爬坡成本开始显现。

2. 燧原登板:四小龙资本拼图补齐​

燧原科技 9 月 2 日启动公开申购,发行 4303 万股新股(约占上市后总股本 10%),募资目标 60 亿元,采用 DSA 专用架构 + 自研 TopsRider 软件平台(不兼容 CUDA)。至此:

  • 科创板:寒武纪(2020 年)、摩尔线程、沐曦、壁仞(2026 年)
  • 燧原:2026 年 9 月完成上市

国产 AI 芯片第一梯队全部进入公开市场,融资通道打开后,研发投入的"军备竞赛"将进一步升级。

3. 需求端:百万卡缺口,订单排到三年后​

财报爆发的另一面,是需求端的极度饥渴:

  • 产能缺口:行业调研显示,2026 年国产 AI 芯片需求规模约 400 万颗,实际交付约 300 万颗,存在百万级缺口;
  • 订单周期:内蒙古乌兰察布远景星河基地(规划支撑百万卡级并行算力)表示"在手订单已排到三年后";
  • 供给创新:算力紧缺催生"集装箱式算力中心"——20/40 英尺标准集装箱为载体,插电接水即可在 24 小时内完成部署,把传统"一年工期"压缩到"一天上线";
  • 市场空间:IDC 数据显示 2025 年中国 AI 加速卡出货约 400 万片,国产占约 41%(165 万片);CIC 预测中国 AI 加速器市场 2028 年将超万亿元,其中国产方案份额约 90%。

4. 两种路线的两种赌注​

市场研究机构预测 2026 年中国高端 AI 芯片市场中,国产方案份额将接近 90%,其中华为约 62%、寒武纪约 14%,剩余由摩尔线程、沐曦、壁仞等共同占据。而寒武纪与摩尔线程这对"双龙头",正在押注两种截然不同的未来:

寒武纪:用现金买确定性。 半年报资产负债表上,存货 82.48 亿 + 预付款 29.14 亿,两项合计 111.6 亿、占总资产 61%——在实体清单限制下"产能即订单",提前锁定未来 12-18 个月的晶圆产能,代价是经营现金流净额同比下滑 66%,且上半年已计提存货跌价损失 3.97 亿。底气来自订单确定性:57 亿元股权激励计划的考核目标是 2026 年营收不低于 135 亿、2026-2028 年累计不低于 1000 亿。

摩尔线程:用亏损买时间。 全功能 GPU 的"大而全"路线(AI 计算 + 图形渲染 + 物理仿真 + 视频编解码)前期投入巨大,需要在现金流耗尽之前跨过规模化的门槛。上半年亏损已收窄至千万级,IPO 募资到位后,时间窗口正在变宽。

5. 软件生态:被忽视的胜负手​

需求外溢(英伟达 H20 系列在中国市场遇冷)推动 DeepSeek、智谱 GLM、月之暗面 Kimi 等头部模型纷纷适配国产加速器(昇腾、沐曦、摩尔线程等),这给了国产芯片难得的"真实负载打磨机会"。但正如行业共识:芯片可以三年一代,生态需要十年之功。寒武纪自研指令集的封闭高效、摩尔线程 MUSA 对 CUDA 生态的兼容路线、燧原 DSA 的专用极致——哪条路能沉淀出真正的开发者生态,才是决定 2030 年格局的关键变量。


相关链接​

参考资料​


本文基于上市公司半年报、Pandaily 与央视财经等公开报道整理。财务数据为公司披露口径,市场份额为研究机构预测,不构成任何投资建议。

Domestic Big Three 2026 H2: Localization Rate Crosses 40% Toward 60%, Ascend 960 Roadmap, MLU690 and S5000 Ecosystems Ramp Up

· 6 min read
Industry Research Team

In 2026, China's AI chip market landscape has shifted from "NVIDIA unipolar dominance" to "overseas vendors leading, domestic multi-route catch-up." According to industry research, China's overall AI accelerator market was ~4M units in 2025, of which 1.65M were domestic, with share first breaking 40%; as products iterate and fabs follow up, the localization rate is expected to rise to 60%-70% by 2027. This article focuses on the latest H2 2026 progress of Huawei Ascend, Cambricon, and Moore Threads — the domestic "Big Three."


1. Huawei Ascend: 950 Capacity Fully Booked, 960 Roadmap Unveiled​

Ascend's core advantage is "architecture + full-stack ecosystem synergy," with ~800K units shipped in 2025, capturing 50% of the total domestic vendor share. The product iteration cadence is clear:

TimeProductNote
2025 Q1Ascend 910CMain transitional model
2026 Q1Ascend 950PRInference flagship
2026 Q4 (planned)Ascend 950DTTraining flagship, drives domestic HBM iteration
2027-2028Ascend 960 / 970Roadmap products

950 series capacity has entered a "fully booked" state: 950PR entered mass production in April 2026; June monthly capacity jumped to 500K-600K units (nearly 10x MoM), with a full-year target of 1.2M units at 100% certainty; ByteDance locked in 350K units for $5.6B, while Tencent / Alibaba / Baidu combined locked in 400K units.

Ascend 960 roadmap specs (per roadmap disclosure):

MetricAscend 960
ArchitectureAscend 6th gen (Da Vinci v6)
FP8 compute~4 PFLOPS
Memory288GB
Memory bandwidth9.6 TB/s
Super-nodeAtlas 960 SuperPoD, 15,488 cards, Lingqu optical-electrical converged bus
Debut2027 Q4 (roadmap)

The previous-gen Ascend 384 super-node has cumulatively shipped over 750 sets, deployed across 20+ industries including internet, operators, finance, education, and healthcare — Huawei calls it "the only domestic super-node that has trained a SOTA model."


2. Cambricon MLU690: H2 Mass Production, Entering ByteDance Bidding Window​

Cambricon is the core domestic compute leader in the absence of an Ascend IPO, with the technology gap continuously narrowing:

  • Siyuan 590 (7nm): Performance equivalent to 80% of A100, already supports DeepSeek, continuously adapting to mainstream large models like Qwen 3 and GLM
  • Siyuan 690 series: Will enter mass production in H2 2026, expected to achieve order scale-up during ByteDance's H2 bidding window
  • Revenue certainty: Equity incentive targets show >100% revenue growth for the next 3 years: 2026 revenue target 13.5B RMB, 2027 27B RMB, 2028 60B RMB

Cambricon fully benefits from the industry dividend of "domestic CSP capex + full adaptation of domestic large models and domestic chips," making it the most direct elasticity play on rising localization rate.


3. Moore Threads MTT S5000: Full-Function GPU + Ecosystem Breakthrough​

Moore Threads takes a differentiated "full-function GPU" route, with the flagship MTT S5000 based on the 4th-gen "Pinghu" MUSA architecture:

MetricMTT S5000
Dense AI compute1000 TFLOPS
Memory80GB
Memory bandwidth1.6 TB/s
Inter-card interconnect784 GB/s
PrecisionFP8 to FP64 full precision (training + inference)
SecurityFirst batch to pass national "Safe and Reliable Evaluation" (Level I)

Its engineering capability is verified: the Kuae (KUAE) intelligent computing cluster based on S5000 achieves 95% training linear scaling efficiency, with compute efficiency loss within 5% at ten-thousand-card scale; supports checkpoint-resume training with effective training time ratio >90%; and has trained a MoE-236B base model with >25 trillion tokens of corpus from scratch.

The ecosystem is Moore Threads' deepest moat: MUSA has achieved 100% core math library compatibility, 3000+ PyTorch operator compatibility, covers 55 categories of core AI operators, has official vLLM and SGLang support, Day-0 adaptation of mainstream models, and 800K+ developers. Its PD heterogeneous-disaggregation solution achieves equivalent replacement of international high-end GPUs at a 2:1 ratio with S5000, significantly reducing inference cost.

The 5th-gen "Huagang" architecture (released 2025-12) supports FP4 to FP64 full precision, with 50% higher compute density and 10x better energy efficiency than the previous gen, supporting 100K+ card clusters; cumulative R&D investment in the "Huashan" (train-infer integrated) and "Lushan" (graphics rendering) new chips based on this architecture exceeds 900M RMB.


4. Software Ecosystem Decides: Day-0 Adaptation Becomes Routine​

Beyond hardware, software ecosystem realization is the watershed for domestic compute in 2026:

  • Huawei's CANN heterogeneous computing architecture and MindSeries suite are fully open-sourced, with the community incubating 67 projects, 12.44M+ lines of code, and 3,500+ monthly active developers
  • The "release-and-adapt" closed loop between domestic large models and domestic chips has basically formed: Tencent Hunyuan T3 (295B), DeepSeek-V4, and GLM-5.2 all completed Day-0 adaptation
  • 2026 is regarded as the "first year of domestic super-nodes"; Huatai Securities estimates China's super-node architecture market will reach 341.4B RMB by 2028, with a 2026-2028 CAGR of 194%

5. Industry Judgment: From "Can It Be Built" to "Can It Be Used Well"​

The domestic Big Three are converging along three paths:

  1. Huawei: Locks government/enterprise and internet big customers with super-node system-level capability + full-stack software
  2. Cambricon: Impacts the revenue inflection point by narrowing the training-side gap + scaling up via big-customer bidding
  3. Moore Threads: Covers cloud-edge-end full scenarios with full-function GPU generality + mature CUDA-compatible ecosystem

The common shortcoming of all three remains advanced process and HBM supply — precisely the core link of overseas controls. But as domestic HBM iterates and fabs follow up, a realistic path to 60%-70% localization by 2027 exists.

References​


This article is compiled from public industry research, broker views, and corporate announcements as of August 2026. Some shipment and market-share figures are third-party estimates, not officially confirmed data.

WAIC 2026 Recap: Huawei Atlas 950 SuperPoD Live Hardware Wins SAIL Grand Award, Domestic Compute Enters the "System-Level" Showdown

· 5 min read
Industry Research Team

The 2026 World Artificial Intelligence Conference (WAIC) was held July 17-20, 2026 at the Shanghai World Expo Center, themed "Intelligent Partners, Creating the Future Together." Over 1,100 companies showcased 3,000+ exhibits, with 300+ products debuting globally. For the compute-card industry, this concentrated review of domestic compute sent a clear signal: the competitive main line is shifting from "single-chip peak compute" to "SuperNode system-level effective compute."

1. Huawei Atlas 950 SuperPoD: live debut, wins SAIL grand award​

Huawei's Atlas 950 SuperPoD live hardware made its first public appearance at WAIC 2026, on-site carrying 16 compute cabinets with 1,024 Ascend cards total. With three system-level innovations — "ultra-wide bandwidth, ultra-low latency, unified memory addressing" — it stood out from hundreds of domestic and international entries to win the conference's top honor, the SAIL (Super AI Leader) Award.

Core parameters (confirmed on-site at WAIC)​

MetricAtlas 950 SuperPoD
Exhibited scale16 compute cabinets / 1,024 Ascend cards
Max interconnect scale8,192 Ascend NPU cards fully interconnected (full config)
Interconnect protocolHuawei in-house "Lingqu" (UnifiedBus) 2.0
Total compute1 EFLOPS FP8 / 2 EFLOPS FP4 (1,024 cards); full 8,192-card ~8 EFLOPS FP8
Unified memory256 TB globally unified memory address space
Interconnect latency3 μs ultra-low RTT; TB-level NPU interconnect bandwidth
Full config128 compute cabinets + 32 interconnect cabinets = 160 cabinets, ~1000㎡, carrying 8,192 Ascend 950DT
LaunchFull config planned for Q4 2026
CoolingFully liquid-cooled blind-plug architecture

Huawei disclosed for the first time: the previous-gen Ascend 384 SuperNode has cumulatively shipped 750+ units commercially, deployed across 20+ industries including internet, operators, finance, education, healthcare, transportation, and manufacturing, calling it "the only domestic SuperNode that has trained SOTA models."

2. Software ecosystem: CANN fully open-sourced, developers at scale​

Beyond hardware, Huawei highlighted open-source software ecosystem progress:

  • CANN heterogeneous compute architecture and MindSeries base software suite were fully open-sourced end of 2025;
  • The CANN open-source community has incubated 67 projects, 12.44M+ lines of code, with 3,500+ monthly active developers;
  • Huawei has co-developed 7,000+ solutions with 3,000+ industry partners, serving 2,000+ core government/enterprise customers;
  • WAIC showcased 60+ real business scenarios, 20+ benchmark cases, covering the full chain from technology breakthrough to scaled commercial deployment.

3. Domestic chips' Day-0 adaptation becomes routine​

On July 6, 2026, Tencent released the MoE model Hunyuan T3 (295B parameters, 256K context); domestic chips rapidly completed Day-0 adaptation:

VendorChipAdaptation status
Moore ThreadsMTT S5000Completed rapid Hunyuan T3 adaptation (previously adapted DeepSeek-V4, GLM-5.2)
MetaXXiyun C seriesIn-house MXMACA stack first to full-chain Day-0 adaptation, zero-code deployment

Moore Threads also showcased the MTT C256 SuperNode (first-of-its-kind single-layer Scale-up 256-card full interconnect, sub-microsecond latency) and three AI-factory solutions — "model training factory / token production factory / agent production factory."

4. More domestic compute debut highlights​

Vendor / productHighlight
Orient AlphaChip DF1000World's first "software-defined + near-memory computing" 3D chip, interconnect pitch compressed to sub-micron
ZhongHao XinYing "Xuyu"Fully in-house next-gen TPU-architecture AI-specific chip, with Taize 2.0 server
Enflame × IluvatarDomestic high-performance Matrix SuperNode based on OEX+dOCS architecture, shortlisted for the conference "Excellent AI Leader Award"
Rongming MicroelectronicsAdvancing next-gen VPU, evolving from video processing to "visual-agent compute base"

The domestic AI chip lineup also included Moore Threads, MetaX, Enflame, Houmo, Cixiong, Suaneng, SemiDrive, Phytium, Aixin, Iluvatar, and others.

Industry interpretation: from "can it be built" to "is it used well"​

WAIC 2026 reflects a fundamental shift in the competitive stage of domestic AI chips:

  1. SuperNode becomes the main battlefield: beyond single-chip performance, system-level capabilities — "inter-chip interconnect + cluster scale + cooling" — become the breakthrough key. Huawei Lingqu and Enflame/Iluvatar OEX are both pushing here. Huatai Securities defines 2026 as the "first year of domestic SuperNodes," estimating China's SuperNode architecture market could reach ¥341.4B by 2028, with 2026-2028 CAGR of 194%.
  2. Software ecosystem delivers: Day-0 adaptation has gone from slogan to routine; the "launch-and-adapt" closed loop between domestic large models (DeepSeek-V4, GLM-5.2, Hunyuan T3) and domestic chips is essentially formed.
  3. Demand-side endorsement: China Mobile earlier released its 2026-2027 AI SuperNode centralized procurement announcement — about 6,208 cards, over ¥2B — accelerating domestic SuperNode scaled commercialization.

References​


This article is compiled from WAIC 2026 (July 17-20) on-site and official disclosures, and will continuously track the 950 SuperNode Q4 launch.

Domestic GPU IPO Wave: The "Four Little Dragons" Assemble on Capital Markets, Moore Threads MTT S5000 Benchmarks Against H100

· 5 min read
Industry Research Team

From December 2025 to July 2026 — just half a year — at least 6 AI chip companies have listed or are about to list on capital markets. Together with already-listed Cambricon, Hygon, and Iluvatar, the domestic GPU corps' total market cap is approaching ¥2 trillion. This marks the critical climb from domestic GPUs being "usable" to "useful."

1. The "Four Little Dragons" assemble on capital markets​

CompanyListing statusRaise / issue priceSponsor
Moore ThreadsListed (STAR Market sh688795, 2025-12-05)Issue price ¥114.28, raised ¥8BCITIC Securities
MetaXIPO accepted (2026-06-30)¥3.904B (total investment ¥5B)Huatai United
EnflamePassed review (2026-06-15)¥6B—
BirenHKEX / sprinting——

Already-listed camp: Cambricon (sh688256, STAR Market 2020-07-20), Hygon, Iluvatar (HKEX). Moore Threads turned a book profit of ¥29.35M in Q1; MetaX narrowed losses 57.7% and gave a 2026 breakeven timeline.

2. Moore Threads MTT S5000: benchmarking against H100​

Moore Threads announced its flagship AI train+inference GPU MTT S5000 successfully completed full-pipeline adaptation validation of Zhipu's new-generation large model GLM-5 — measured performance "breaks the domestic compute ceiling":

MetricMTT S5000
Architecture4th-gen "Pinghu" architecture
FP8 compute1 PFLOPS (1,000 TFLOPS)
Memory bandwidth1.6 TB/s
PositioningFull-function train+inference GPU, benchmarks against NVIDIA H100
ProductionMass-produced; clusters online supporting trillion-parameter training

Deployment validation: jointly completed full-pipeline training of embodied-brain model RoboBrain 2.5 with BAAI; partnered with SiliconFlow for high-performance DeepSeek-V3 inference, single-card speed near international top products. IPO funds go to three directions: next-gen AI train+inference chip, next-gen graphics chip, next-gen AI SoC chip.

WAIC 2026 new progress: Moore Threads showcased the MTT C256 SuperNode (first-of-its-kind single-layer Scale-up 256-card full interconnect, sub-microsecond latency) and three AI-factory solutions — "model training factory / token production factory / agent production factory"; the company pre-announced H1 2026 revenue of ¥1.65B-1.75B, up 135%-149% YoY.

3. Cambricon: dual flagships MLU590/690​

ChipProcessComputeMemoryCustomer / status
MLU590 (思元590)7nm ChipletINT8 512 TOPS / FP16 345 TFLOPS96 GB HBM2eByteDance inference mainstay, ~80% of A100 overall, mass shipments early 2026
MLU690 (思元690)5nm-class (SMIC N+2)FP16 700+ TFLOPS / INT8 2800+ TOPS196 GB HBM3 (3.35 TB/s)Dual-die packaging, MLU-Link 890 Gbps; ~70% of H100 (80-90% pure inference); ByteDance largest customer, mass production early 2026

Cambricon is the only domestic AI chip vendor with a "unified edge-cloud architecture" — one MLU instruction set spans 思元 220 (edge) → 370 (border) → 590/690 (cloud), with one NeuWare toolchain across compute tiers.

Capital and performance double explosion: Cambricon's total market cap exceeded ¥1 trillion on June 30, 2026, becoming the STAR Market's first "trillion-yuan stock," up 75%+ YTD. On performance, Q1 2026 revenue ¥2.885B (+160% YoY), deducted net profit ¥934M; full-year 2025 revenue ¥6.497B (+453% YoY), net profit attributable to parent ¥2.059B, ending long-term losses. ByteDance has cumulatively deployed over 100k 思元 590/690, its largest customer.

4. DeepSeek-V4 effect: changing the expectation coordinate system​

On April 24, 2026, DeepSeek released the trillion-parameter flagship DeepSeek-V4. Unlike a year earlier when V3's launch sparked debate over "can domestic chips even run large models," this time multiple domestic chips — Huawei Ascend, Cambricon, Hygon, MetaX, Moore Threads, Kunlun, T-Head, Iluvatar — completed adaptation on launch day.

The evaluation coordinate system is shifting: from "what percentage of NVIDIA's same-generation product performance" to "can it carry the real workloads of top-tier large models."

Industry interpretation​

  1. Capital ammunition in place: dense IPOs provide ample funding for domestic GPU R&D iteration and capacity expansion, moving from "technology breakthrough" to "commercial virtuous cycle."
  2. Train+inference becomes the mainstream route: Moore Threads takes the full-function GPU route (graphics+AI+general compute), differentiating from Huawei Ascend's "AI-focused."
  3. Software ecosystem is the decider: Day-0 adaptation and the maturity of unified software stacks (MUSA / NeuWare / MXMACA) are replacing raw peak compute as the core yardstick of domestic GPU "usability."

References​


This article continuously tracks the domestic GPU listing process and product iteration.

China's Domestic AI Chip Triopoly (2026): Ascend, Cambricon, Moore Threads — Who Is the "China H100"?

· 7 min read
AI Hardware Analyst

Against the backdrop of U.S. export controls, China's AI chip market is forming a "three-way standoff." This article compares the technical routes, product specs, software ecosystems, and commercial progress of the three major domestic AI chip vendors: Huawei Ascend, Cambricon MLU, and Moore Threads MTT.


Key Points​

  • Huawei Ascend: leader in domestic AI training chips; Ascend 950 in mass production; most mature software ecosystem
  • Cambricon MLU690: the "China H100," compute close to H200, clear efficiency advantage
  • Moore Threads MTT S5000: full-function GPU route; achieved Day-0 support for Qwen3.5 and GLM-5.2 in June 2026
  • Shared challenge: affected by U.S. export controls, primarily aimed at the Chinese market, limited internationally

I. Vendor Overview​

VendorFoundedFounderListed2025 RevenueMain Customers
Huawei Ascend2018 (division)Ren Zhengfeiprivate (wholly owned by Huawei)~¥20B (est.)Chinese gov, SOEs, military
Cambricon2016Chen Tianshi (CAS)2020-07 (STAR Market 688256)~¥5.2BByteDance, Alibaba, Baidu
Moore Threads2020Zhang Jianzhong (ex-NVIDIA China)2023-12 (STAR Market 688495)~¥1.5B (est.)gov, SOEs, gaming cos.

Strategic Positioning​

VendorTech routeCore strengthMain challenge
Huawei AscendAI-training-specific (Da Vinci)co-optimized HW/SW, carrier channelssanctions, process limits
CambriconAI-training-specific (MLUarch)high efficiency, competitive priceimmature ecosystem
Moore ThreadsFull-function GPU (MUSA)graphics + AI + general compute, Day-0 supportcompute below dedicated AI chips

II. Flagship Product Comparison​

1. Huawei Ascend 950DT (2026 flagship)​

ItemSpec
BF16 compute1,000 TFLOPS
Memory144GB HiZQ 2.0 (in-house HBM)
Memory bandwidth4 TB/s
TDP400W
ProcessN+2 (improved 7nm)
Released2026-04
Mass production2026-Q2
Unit price~¥80,000 (est.)

Strengths:

  • ✅ High large-model inference throughput: 144GB memory friendly to DeepSeek R1 (671B MoE)
  • ✅ Most mature ecosystem: CANN ~85% operator coverage, supports PyTorch, TensorFlow
  • ✅ Strong carrier channel: China Mobile, China Telecom large purchases

Weaknesses:

  • ❌ Process limited: N+2 below TSMC 4nm
  • ❌ Mediocre efficiency: 400W TDP, 2.5 TFLOPS/W

2. Cambricon MLU690 (2026 flagship)​

ItemSpec
BF16 compute600 TFLOPS
Memory64GB HBM3
Memory bandwidth2 TB/s
TDP280W
ProcessTSMC 7nm
Released2025-Q4
Mass production2026-Q1
Unit price~¥140,000 (est.)

Strengths:

  • ✅ Best efficiency: 280W TDP, 2.14 TFLOPS/W (1.5x H100)
  • ✅ Competitive price: ~$20,000, 33% cheaper than H100
  • ✅ Top-tier customer orders: ByteDance, Alibaba, Baidu

Weaknesses:

  • ❌ Small memory: 64GB limits large-model training scale
  • ❌ Immature ecosystem: NeuWare ~75–85% coverage; complex LLMs need manual tuning

3. Moore Threads MTT S5000 (2025 flagship)​

ItemSpec
FP16 compute~1,000 TFLOPS (est.)
Memory80GB GDDR6X
Memory bandwidth1.6 TB/s
TDP~350W
ProcessTSMC 4nm (est.)
Released2025-02
Mass production2025-Q2
Unit price~¥50,000 (est.)

Strengths:

  • ✅ Full-function GPU: graphics + AI + general compute, broader scenarios
  • ✅ Strong Day-0 support: June 2026 Day-0 support for Qwen3.5, GLM-5.2, MiniMax M3
  • ✅ Lowest price: ~¥50,000, high cost-performance

Weaknesses:

  • ❌ Compute below dedicated AI chips: FP16 ~50% of H100
  • ❌ Low memory bandwidth: 1.6 TB/s (48% of H100), limits large-model training

III. Compute Comparison (BF16/FP16)​

ChipBF16 computeMemoryBandwidthTDPEfficiency
Huawei Ascend 950DT1,000 TFLOPS144GB4 TB/s400W2.5 TFLOPS/W
Cambricon MLU690600 TFLOPS64GB2 TB/s280W2.14 TFLOPS/W
Moore Threads MTT S5000~1,000 TFLOPS80GB1.6 TB/s~350W~2.86 TFLOPS/W
NVIDIA H100989 TFLOPS80GB3.35 TB/s700W1.41 TFLOPS/W
NVIDIA H200989 TFLOPS141GB4.8 TB/s700W1.41 TFLOPS/W

Key insights:

  1. Ascend 950DT has the highest compute (1,000 TFLOPS) but mediocre efficiency
  2. Cambricon MLU690 has the best efficiency (2.14 TFLOPS/W), TDP only 280W
  3. Moore Threads MTT S5000 wins on full-function versatility but low bandwidth

IV. Software Ecosystem​

VendorStackFramework supportCoverageMaturity
Huawei AscendCANNPyTorch, TensorFlow, MindSpore~85%⭐⭐⭐⭐ (4/5)
CambriconNeuWarePyTorch-Cambricon, TensorFlow-Cambricon~75–85%⭐⭐⭐ (3/5)
Moore ThreadsMUSIFYPyTorch, TensorFlow, ONNX~70%⭐⭐⭐ (3/5)
NVIDIACUDAall~99%⭐⭐⭐⭐⭐ (5/5)

Ecosystem Maturity Assessment​

Huawei Ascend CANN:

  • ✅ Strength: highest operator coverage, supports MindSpore (in-house framework)
  • ❌ Weakness: steep learning curve, incomplete docs

Cambricon NeuWare:

  • ✅ Strength: PyTorch/TensorFlow compatible, low migration cost
  • ❌ Weakness: complex LLMs need manual tuning

Moore Threads MUSIFY:

  • ✅ Strength: strong Day-0 support, ONNX support
  • ❌ Weakness: lowest operator coverage, dual graphics+AI engine complexity

V. Commercial Progress​

Vendor2026 commercial progressMain customersShipments
Huawei AscendAscend 950 mass production; China Mobile large purchaseChina Mobile, China Telecom, gov~100K/yr (est.)
CambriconMLU690 mass production; ByteDance, Alibaba ordersByteDance, Alibaba, Baidu~50K/yr (est.)
Moore ThreadsMTT S5000 mass production; Day-0 Qwen3.5gov, SOEs, gaming cos.~30K/yr (est.)

Latest as of June 2026​

Huawei Ascend:

  • ✅ Ascend 950DT fully ramping
  • ✅ ¥1B procurement agreement with China Mobile

Cambricon:

  • ✅ MLU690 in volume shipment
  • ✅ ByteDance order ~20K units

Moore Threads:

  • ✅ Day-0 support for Qwen3.5, GLM-5.2, MiniMax M3
  • ✅ MTT S5000 2nd-gen released

VI. Selection Advice​

Scenario 1: Trillion-parameter training (GPT-4 class)​

Recommended: Huawei Ascend 950DT

  • ✅ 144GB large memory supports super-large models
  • ✅ Most mature ecosystem (~85% coverage)
  • ✅ Strong carrier channel, Chinese government backing

Alternative: Cambricon MLU690 (high efficiency, but small memory)

Scenario 2: Tens-to-hundreds-of-billions parameter training​

Recommended: Cambricon MLU690

  • ✅ Best efficiency (2.14 TFLOPS/W), low TCO
  • ✅ Competitive price (~$20,000)
  • ✅ Validated by top customers (ByteDance, Alibaba)

Alternative: Huawei Ascend 920 (more compute, mediocre efficiency)

Scenario 3: Cloud AI inference​

Recommended: Huawei Ascend 950PR (inference-specific)

  • ✅ Well-optimized inference throughput
  • ✅ 128GB memory friendly to MoE models
  • ✅ Mature stack, low deployment cost

Alternative: Moore Threads MTT S5000 (full-function GPU, inference + graphics)

Scenario 4: Edge AI / on-device inference​

Recommended: Moore Threads MTT S5000

  • ✅ Full-function GPU, graphics + AI
  • ✅ Lowest price (~¥50,000)
  • ✅ Strong Day-0 support

Alternative: Huawei Ascend 310 (low power, 8W TDP)

Scenario 5: Domestic substitution (gov, SOEs)​

Recommended: Huawei Ascend 950DT

  • ✅ Chinese government first choice, carrier bulk buys
  • ✅ Co-optimized HW/SW, stable performance
  • ✅ Supported by national semiconductor fund

Alternative: Cambricon MLU690 (high efficiency, competitive price)


VII. Future Roadmap​

Vendor2026 H220272028
Huawei Ascend950DT ramp960 (FP8 ~2 PFLOPS)970 (N+3 process)
CambriconMLU690 rampMLU790 (5nm, BF16 ~1,000 TFLOPS)MLU890 (3nm)
Moore ThreadsMTT S5000 2nd-genMTT S6000 (HBM3, FP16 ~1,500 TFLOPS)MTT S7000

VIII. Summary: Who Is the "China H100"?​

DimensionAscend 950DTMLU690MTT S5000
Compute⭐⭐⭐⭐⭐ (5/5)⭐⭐⭐ (3/5)⭐⭐⭐ (3/5)
Memory⭐⭐⭐⭐⭐ (5/5)⭐⭐ (2/5)⭐⭐⭐ (3/5)
Efficiency⭐⭐⭐ (3/5)⭐⭐⭐⭐⭐ (5/5)⭐⭐⭐⭐ (4/5)
Ecosystem⭐⭐⭐⭐ (4/5)⭐⭐⭐ (3/5)⭐⭐⭐ (3/5)
Price⭐⭐⭐ (3/5)⭐⭐⭐⭐ (4/5)⭐⭐⭐⭐⭐ (5/5)
Overall⭐⭐⭐⭐ (4/5)⭐⭐⭐ (3/5)⭐⭐⭐ (3/5)

Final conclusion:

  • Huawei Ascend 950DT is the domestic AI training chip closest to H100, strongest overall
  • Cambricon MLU690 is the most efficient domestic AI chip, lowest TCO
  • Moore Threads MTT S5000 is the cheapest full-function GPU, suited to edge AI and graphics+AI

References​


Disclaimer: Data based on public sources; actual specs per vendor official. MirrorFrog continuously updates domestic AI chip data — corrections welcome.

Changelog: 2026-06-23 initial release