Skip to main content

9 posts tagged with "AI Hardware"

AI hardware industry dynamics beyond chips (HBM, servers, systems)

View all tags

CoWoS 之困:先进封装如何成为 AI 芯片最隐蔽的瓶颈

· 7 min read
Industry Research Team

谈到 AI 芯片瓶颈,人们第一反应是 HBM 或先进制程,但真正的隐形关卡是封装。一台 3nm 光刻机不够可以排队再买,一条 CoWoS 产线从规划到爬坡却要以年计。2026 年 10 月,GlobalFoundries 与台积电这对代工对手签下互供协议,把封装产能的稀缺性暴露无遗。

AI 芯片封装"三件套"​

现代 AI 芯片的封装不再是把芯片装进外壳,而是构造一个微型系统。三个核心技术构件:

  • 2.5D 中介层(interposer):一块硅材质的"转接板",放在 GPU die 与基板之间,上面布满密度远超 PCB 的走线,让 GPU 与 8-12 颗 HBM 之间以极短距离高速通信;
  • 3D 堆叠(hybrid bonding):用铜-铜混合键合把 die 直接面对面堆叠,不再靠凸点,互连密度提升一个数量级,主要用于 HBM 内部堆叠与缓存堆叠;
  • RDL 重布线层:把 die 上的焊盘"重新布线"到合适位置,承担封装内的横向连接与信号引出。

三者组合,决定了 AI 芯片的集成形态。以 AMD MI455X 为例:8 个 N2 制程计算芯粒(XCD)、2 个 I/O die 与 2 个缓存/fabric die(N3),加上 12 颗 HBM4 堆栈,全部通过 CoWoS-L 封装整合在同一个封装内;NVIDIA Rubin R200 则是 2 颗计算 die + 2 颗 I/O die + 8 颗 HBM4 的 MCM 多芯片模块设计。没有先进封装,就没有现代 AI GPU。

CoWoS 技术梯队:S、R、L 的演进​

台积电的 CoWoS(Chip-on-Wafer-on-Substrate)家族按中介层技术分为三代:

方案中介层技术特点典型应用
CoWoS-S硅中介层(Si interposer)成熟、成本最低,面积上限约 3 倍 reticleH100、早期数据中心 GPU
CoWoS-RRDL 中介层用有机重布线层替代硅,面积更大、成本更低,带宽密度略降H200 及部分 Blackwell 配置
CoWoS-L局部硅互连(LSI)+ RDL硅桥局部互连 + 大面积布线,面积上限突破 3.3 倍 reticle,互连密度最高B200/GB200、MI455X、Rubin

演进的方向很明确:封装面积越来越大(从 1 倍 reticle 到 3 倍以上)、集成组件越来越多(GPU + HBM + 缓存 die + I/O die)、互连密度越来越高。CoWoS-L 已成为旗舰 AI 芯片的事实标准,也意味着单颗芯片占用的封装产能成倍增长。

为什么封装产能比晶圆产能更紧​

直觉上,先进制程更难、更稀缺,但封装恰恰是更紧的那一环,原因有四:

  1. 产能基数小:全球先进制程晶圆厂有数十座,而具备大规模 CoWoS 产能的基地屈指可数,且全部集中在台积电体系内;
  2. 扩产周期长:洁净厂房、无尘搬运、精度设备交付周期都以年计,CoWoS 月产能从规划到释放动辄 18-24 个月;
  3. 需求端放大:HBM 堆栈数量从 6 颗涨到 12 颗、die 数量翻倍,单卡消耗的封装面积持续膨胀,产能需求指数级增长;
  4. 良率敏感:中介层越大越贵,一片超大中介层的报废就是整封装报废,有效产能远低于名义产能。

AMD 在 2026 年明确指出,HBM 和先进封装(CoWoS)供给紧张是其增长的主要约束——注意,AMD 的瓶颈不在自家芯片设计,而在拿不到足够的封装配额。

GF-TSMC 互供:一个罕见的市场信号​

2026 年 10 月 8 日,DIGITIMES 报道 GlobalFoundries 与台积电签署多年期协议,金额约 20 亿美元:GF 将在美国纽约州 Malta 工厂为台积电的 CoWoS 生态生产硅中介层。

这个安排罕见在哪里?GF 与台积电在先进代工市场是直接竞争对手,让对手替自己生产关键耗材,只有一种解释——封装产能稀缺到连代工龙头都要向外借力。对台积电,这是快速扩大 interposer 供给的最短路径;对 GF,这是切入 AI 封装供应链的高价值订单。同时,地缘上它把 part of CoWoS 供应链锚在美国本土,呼应美国先进制造政策。

产业配套也在加速:应用材料(Applied Materials)与半导体封装设备龙头 Besi 于 2026 年 10 月 1 日宣布扩大 EPIC(Electronics Packaging Industry Consortium)中心合作,联合开发混合键合等下一代封装工艺;应用材料还与 Intel 扩大了下一代互连扩展与 Foveros 3D 技术的合作。设备商、代工厂、IDM 三方同时下注,印证封装已是 AI 算力的战略资源。

Foveros 与 CoWoS:两条路线之争​

台积电 CoWoS 并非唯一答案。Intel 的 Foveros 是另一条 3D 封装路线,两者定位有所不同:

维度台积电 CoWoS-LIntel Foveros
核心思路2.5D 中介层承载多 die + 局部 3D 互连3D 堆叠为基础逻辑 die 面对面互连
优势生态成熟、HBM 集成经验深厚、大封装面积互连密度与能效更高,垂直距离更短
主要客户NVIDIA、AMD、Broadcom 等绝大多数 AI 芯片厂Intel 自家产品为主,逐步开放代工
当前角色AI GPU 封装事实标准追赶者,重点在下一代互连扩展

短期内 CoWoS 地位难以撼动——AI GPU 的核心需求是把大量 HBM 和 compute die 拼在一张大封装里,CoWoS-L 正为此而生。Foveros 的机会在于 3D 堆叠密度与能效的长期演进,以及与 Intel 先进制程组合的代工吸引力。

对芯片厂商的启示:封装配额就是出货上限​

2026 年的 AI 芯片竞争已经形成一条铁律:你的出货量不取决于你能设计多少芯片,而取决于你能拿到多少封装配额。

  • 对 NVIDIA、AMD:CoWoS 配额直接决定 Blackwell/Rubin 与 MI400 系列的交付节奏,封装扩产进度就是营收曲线;
  • 对国产芯片:在台积电 CoWoS 体系之外建立自主先进封装能力,与自研 HBM 同等重要,是摆脱供应链约束的另一半拼图;
  • 对云厂商与大买家:锁定芯片产能的同时必须锁定配套封装与 HBM,三者缺一即断供。

小结​

AI 芯片的瓶颈正在从"造出更小的晶体管"转向"把更多 die 更好地装在一起"。CoWoS-S/R/L 的技术梯队撑起了从 H100 到 Rubin 的每一次性能跨越,而 GF 与台积电的互供协议则宣告:封装产能已稀缺到让对手合作。下一个五年,谁掌握先进封装,谁就握有 AI 算力版图的划分权。

从 HBM2e 到 HBM4E:一文读懂 AI 显存的四代演进

· 6 min read
Industry Research Team

大模型参数每 18 个月翻一番,GPU 算力每代翻 3-5 倍,而真正卡住 AI 芯片脖子的,往往是那圈贴在 GPU 旁边的"小方块"——HBM 高带宽显存。2026 年,HBM4 进入全面量产,HBM4E 已在路上,本文带你一次读懂四代演进的技术脉络。

HBM 是什么:把内存"竖起来"的 3D 堆栈​

传统 GDDR 显存颗粒是平铺在 PCB 上的,走线长、寄生电感大,带宽天花板低。HBM(High Bandwidth Memory)的思路完全不同:

  • 3D 堆栈:把 8 层甚至 12 层 DRAM 裸片垂直叠在一起,通过**微凸点(micro-bump)**层间互连;
  • TSV 硅通孔:每层裸片打穿硅片形成垂直通路,让堆栈内部各层共享一条超宽的总线;
  • 与 GPU 同封装:HBM 不再焊接在主板上,而是作为独立堆栈放在 GPU 同一个封装内,通过**硅中介层(interposer)**与 GPU 的 die 相连,走线距离从几十毫米缩短到毫米级。

这三件事叠加的结果是:接口位宽可以从 GDDR 的 32bit/颗粒跃升到每堆栈 1024bit 甚至 2048bit,带宽提升十倍以上。代价则是制造难度极高——TSV、堆叠键合、中介层封装每一环都是良率杀手。

四代 HBM 关键参数对比​

代际接口位宽单堆栈带宽单堆栈容量典型层数代表产品
HBM2e1024bit~460 GB/s16-24 GB8 层H200、昇腾 910C
HBM31024bit~819 GB/s24 GB8-12 层H100
HBM3e1024bit~1.15-1.2 TB/s24-36 GB8-12 层B200、MI355X
HBM4/4E2048bit2-2.8+ TB/s36 GB 起12-16 层Rubin R200、MI455X

以站内芯片卡的数据为参照:NVIDIA B200(Blackwell)配备 192GB HBM3e,带宽 8 TB/s;到 Rubin R200,8 颗 HBM4 堆栈带来 288GB 容量与 22 TB/s 带宽,带宽提升 2.75 倍。AMD MI455X 更激进,12 颗 36GB HBM4 堆栈做到 432GB / 23.3 TB/s。SK 海力士的 12 层 HBM4 单堆栈规格为 36GB、2.8 TB/s、11 Gbps 引脚速率,已于 2026 年 7 月量产出货。

为什么 2048bit 接口是分水岭​

前三代 HBM 一直沿用 1024bit 接口,性能提升主要靠拉高引脚速率。这条路在 HBM3e 时代已经接近极限——信号完整性与功耗随速率平方增长,继续提速性价比骤降。

HBM4 做了一个根本性改动:接口位宽翻倍到 2048bit,而引脚速率只温和提升。以 Rubin R200 为例,单堆栈 2048bit 接口配合约 10.8 GT/pin 的引脚速率,即可得到约 2.7-2.8 TB/s 的堆栈带宽(与 SK 海力士单堆栈 2.8 TB/s 的规格吻合),8 颗堆栈合计正是 22 TB/s 的总带宽。SK 海力士单堆栈 2.8 TB/s 的成绩同样是"宽接口 + 适中速率"的结果。

位宽翻倍的意义在于:同样的数据量,需要的引脚速率更低、功耗更低、信号更稳。但它也带来新问题——堆栈与 GPU 之间的中介层走线密度翻倍,封装难度和成本水涨船高。HBM4 的竞争,本质上是封装能力的竞争。

HBM4E、12 层堆栈与 zHBM 展望​

HBM4 的下一步是 HBM4E(Enhanced):

  • 容量继续堆高:16 层堆栈将成为 HBM4E 主流,单堆栈容量迈向 48GB;
  • 速率再上台阶:引脚速率从 HBM4 的 10+ Gbps 提升到 13-15 Gbps 区间;
  • 三星入局:三星的 HBM4E 已通过 NVIDIA 验证,HBM 市场从 SK 海力士一家独大走向三强分食;
  • zHBM 路线:三星还公布了 zHBM 概念——在 HBM 堆栈底部集成逻辑 die(Base Logic 加速器),把部分计算下沉到显存侧,探索"存算一体"式的近存计算,瞄准 HBM4E 之后的时代。

对整机厂而言,HBM4E 意味着单机柜显存总量再上台阶:Rubin 时代的 Vera Rubin NVL72 已有 20.7TB 总显存,Helios 机架(72 颗 MI455X)达 31TB,下一代翻倍几乎是定局。

HBM 定价权:三巨头最强时刻与国产涨价潮​

2026 年是 HBM 厂商十年来话语权最强的时刻:

  • 供不应求:三星、SK 海力士、美光的 HBM 产能已被预订到 2029 年,三大原厂议价能力空前;
  • 成本传导:HBM 占 AI GPU 物料成本的比例从上一代的约 8% 攀升到 20% 以上,显存涨价直接推高整机价格;
  • 国产涨价潮:外部 HBM 紧缺叠加国产需求爆发,昇腾系列报价显著上涨——昇腾 950PR 从年初约 6 万元涨破 8 万元(+30%),昇腾 950DT 指示性报价超过 25 万元,两个月内涨幅最高 50%。

涨价潮的根源不在芯片本身,而在 HBM:谁掌握显存供给,谁就掌握 AI 芯片的出货上限。

两条突围路线:自研 HBM 与定制 HBM​

面对 HBM 瓶颈,产业界走出两条不同的路:

路线一:自研 HBM(华为)。 昇腾 950 系列首次采用华为自研 HBM——950PR 使用成本优先的 HiBL 1.0(128GB、1.6 TB/s),950DT 使用带宽优先的 HiZQ 2.0(144GB、4 TB/s),彻底摆脱对 SK 海力士和三星的依赖。自研 HBM 未必追求绝对性能领先,但解决了供应链"卡脖子"这一生死问题。

路线二:定制 HBM(NVIDIA)。 NVIDIA 推动与显存原厂合作开发 NVHBM——定制化堆栈,把 GPU 需要的接口逻辑直接做进 HBM 底部逻辑层,进一步压缩延迟、提高集成度。定制化意味着 GPU 厂商深度介入显存设计,原厂从"卖颗粒"升级为"卖方案"。

自研是供应链安全的答案,定制是性能极致的答案,两者将长期并存。

小结​

四代 HBM 的演进主线非常清晰:HBM2e 到 HBM3e 在 1024bit 接口内"榨"速率,HBM4 用 2048bit 宽接口换回功耗与信号余量,HBM4E 与 zHBM 则在容量和架构上继续突破。显存已经从 GPU 的配角变成 AI 产业的核心资源——它决定带宽、决定成本、决定出货量,甚至决定各国 AI 算力的自主程度。看懂 HBM,才算真正看懂 AI 芯片的账本。

制程竞赛 2026:N2 量产、18A 蓄力、High-NA EUV 百万晶圆

· 5 min read
Industry Research Team

每一次制程换代都会重新洗牌 AI 算力格局。2026 年,2nm 从路线图走进量产车间:AMD 抢下 N2 AI GPU 首发权,EPYC Venice 成为全球首款 2nm 数据中心 CPU,Intel 的 High-NA EUV 则悄悄跨过百万晶圆门槛。制程竞赛的规则,也在这一年悄悄改变。

制程、能效与晶体管密度的链条​

制程的意义可以用一条链概括:更小的晶体管 → 更高的密度与更低的开关功耗 → 同样封装面积塞进更多计算单元与更大缓存 → 单位 token 的能耗与成本下降。对 AI 芯片而言,这条链的最后一段尤其重要——推理集群的电费和散热已经取代芯片采购价,成为 TCO 的最大头之一,制程每前进一代,同等机柜功率下能跑的 token 数就上一个台阶。

2026 年的 2nm 落地图谱如下:

厂商产品制程状态与要点
AMDInstinct MI455XTSMC N2 计算芯粒 + N3 I/O 与缓存芯粒首款 TSMC N2 AI GPU,3200 亿晶体管,432GB HBM4,2026-07-23 发布
AMDEPYC Venice(Zen 6)TSMC N2全球首款 2nm 数据中心 CPU,256 核,2026 年 5 月量产
GoogleTPU 8t(Trillium 2)TSMC 2nm训练专用 ASIC,双计算 Die,9600 芯片 Pod,121 EFLOPS FP4
IntelFoundry High-NA EUV 产线High-NA EUV已处理超 100 万片晶圆(2026-09-21 当周口径)
三星2nm AI 芯片制程SF2 路线联合台积电推进,目标 2027 年量产,宣称能效比提升 40%(2026-09 口径)

N2 AI GPU 首发为何是 AMD​

按过往节奏,新制程的首发大客户通常是谁?答案是 AMD——而且这不是偶然。MI455X 采用 8 个 N2 计算芯粒(XCD)加 2 个 I/O die 与 2 个缓存/fabric die(N3)的混合封装:把最吃密度的计算芯粒放最先进节点,把成本敏感的 I/O 与缓存留在成熟一档的 N3,用 CoWoS-L 先进封装拼合,既拿到 2nm 的能效红利,又控制了整片成本。

同一家公司连下两城:EPYC Venice 用 N2 做出 8 颗 Zen 6C CCD、每颗 32 核、总计 256 核、1GB L3 缓存的怪兽 CPU,单路内存带宽 1.6TB/s,与 MI455X 组成 Helios 机架(18 颗 Venice + 72 颗 MI455X,31TB HBM4,2.9 FP4 exaFLOPS)。CPU 与 GPU 同代同步上新制程,是 AMD chiplet 战略的复利——分芯粒设计天然适配"最先进节点给计算、次先进节点给 I/O"的分配策略。Google TPU 8t 也选择了同样的 N2 + Broadcom 协作路线,印证了"定制 ASIC 抢新制程"的趋势。

High-NA EUV 百万晶圆的意义​

Intel 在 2026 年 9 月 21 日当周披露,其 Foundry 的 High-NA EUV 产线累计处理晶圆已超过 100 万片。这个数字的意义在于:High-NA EUV——分辨率更高、专为 2nm 及以下节点准备的下一代光刻机——从"实验室样机"正式进入"规模化生产"阶段。百万晶圆意味着工艺窗口验证、良率爬坡和产能利用都跨过了商业化门槛,Intel 18A 与 14A 路线因此有了现实的制造底盘。

对照看三星的动作同样密集:HBM4 已于 2026 年 2 月量产(成为 Rubin 等旗舰 GPU 的 HBM 供应商之一),2nm AI 芯片制程则与台积电形成竞争态势,目标 2027 年量产。制程竞赛与 HBM 竞赛第一次同步开跑。

"十年之约":平台期被拉长​

台积电的投入力度佐证了这个节点的分量:2026 年 CAPEX 顶格到 560 亿美元上限,全年营收增速指引 30% 以上,3nm 产能同时在台湾、亚利桑那与熊本三地扩张。但硬币的另一面是 ASML 与 ZEISS 给出的"十年之约"——下一代光刻(High-NA 之后的技术)可能还需要十年才能成熟。这意味着 2nm/1.4nm 将成为一个超长平台期:没有新的光学分辨率红利可吃,各家的竞争焦点必须转移。

转移的方向已经很清晰。第一是先进封装:MI455X 的 12 颗 HBM4 堆栈、Rubin 的 MCM 多芯片设计,都说明系统级创新正在替代晶体管微缩。第二是 HBM 配套:N2 计算芯粒如果没有 22-23TB/s 的 HBM4 喂数据,算力就是摆设——制程与内存的协同定义能力,比制程数字本身更能决定产品成败。

小结​

2026 年的制程竞赛呈现出三个信号:2nm 实力量产,AMD 借 chiplet 架构连夺 AI GPU 与数据中心 CPU 双首发,Google 定制 ASIC 跟进;Intel 用 High-NA EUV 百万晶圆宣告制造能力回归牌桌;而 ASML 的十年之约意味着制程数字的边际收益递减,先进封装、HBM 协同与能效工程将成为新的竞争主轴。对 AI 算力买家来说,评估一颗芯片时,"几纳米"正在变成一个越来越不完整的答案——封装里装了什么、带宽喂了多少、每瓦能跑多少 token,才是更完整的提问方式。

内存超级周期全面爆发:台积电Q3营收创纪录、三星利润暴涨8倍、HBM4E获NVIDIA认证

· 5 min read
Industry Research Team

财报季密集报喜:台积电 Q3 营收约 NT$1.49 万亿创历史新高,三星 Q3 营业利润同比大增近 8 倍,12 层 HBM4E 已通过 NVIDIA 质量验证——AI 算力的利润池正在向存储环节大规模迁移。

财报季:台积电、三星双双创纪录​

据路透与 DIGITIMES 10 月 8 日至 9 日的报道,财报季传来一连串超预期数据:

  • 台积电:Q3 营收约 NT$1.49 万亿,创历史纪录并超出市场预期,AI 需求为主要驱动力;其中 9 月单月营收约 NT$5118.6 亿,同比增长 54.6%;
  • 三星电子:Q3 营业利润预计超 100 万亿韩元,同比大增近 8 倍,HBM 与 AI 存储需求爆发是核心推手。
公司关键数据
台积电Q3 营收约 NT$1.49 万亿,创纪录并超市场预期;9 月单月约 NT$5118.6 亿,同比 +54.6%
三星电子Q3 营业利润预计超 100 万亿韩元,同比大增近 8 倍,HBM 与 AI 存储需求爆发

代工与存储同步创新高,说明这一轮景气并非单一环节的结构性紧张,而是 AI 算力全链条的需求爆发。

三星:HBM4E 通过验证,下一代路线图发布​

三星在产品层面有两个关键动作:

  • 下一代 12 层 HBM4E 已通过 NVIDIA 及主要超大规模客户的质量验证,为后续放量铺平道路;
  • 发布下一代 AI 存储路线图,包括 zHBM 与 400 层以上 V10 NAND 技术。

HBM4E 通过 NVIDIA 认证,意味着三星在 HBM4 世代的竞争中拿到了进入下一轮的门票。在 AI 存储这个增长最快的细分市场,三星与 SK 海力士的正面交锋还将继续,而 zHBM 与 400 层以上 V10 NAND 的路线图,则表明存储厂商的技术竞赛已经延伸到两个世代之后。

封装与设备:产能稀缺的另一面​

存储之外,产业链上下游同样动作密集:

  • GlobalFoundries 与台积电签署多年期 20 亿美元协议,将在其美国 Malta 工厂为台积电 CoWoS 生态生产硅 interposer——两大代工对手互供,属于罕见安排;
  • Applied Materials 与 Intel 扩大合作,聚焦下一代互连扩展与 Foveros 3D 技术;
  • ASML 与 ZEISS 表示,下一代光刻技术(High-NA 之后)可能还需十年才会到来。

GF 为台积电生产 interposer 尤其值得玩味:在先进封装产能全面紧张的当下,封装环节的稀缺程度已经超过代工环节的竞争——连对手之间也要互相借产能,这正是 CoWoS 类先进封装供不应求的最直观注脚。

内存挤压传导至 AI PC​

DIGITIMES 分析指出,AI PC 正面临内存挤压:在内存厂商的优先级中,服务器 DRAM 需求被放在首位,AI PC 依赖的 LPDDR5X 供给承压。刚刚开启预订、最高配备 128GB 统一内存的 Surface Laptop Ultra,正是这一矛盾的直接体现——AI PC 想要的内存容量越大,与服务器需求的冲突就越尖锐。

美光此前已预测,NAND 与 DRAM 短缺将持续至 2028 年。存储缺货不是短期波动,而是一个跨越数年的结构性周期。

分析:利润池正在迁移​

第一,利润从加速器向存储迁移。"没有分配到 HBM 的 AI GPU 是库存而不是算力"——这句话正在成为行业共识。GPU 再强,没有 HBM 就出不了货,存储环节因此获得了整个 AI 供应链中最强的话语权之一。三星营业利润暴增近 8 倍,就是利润池迁移最直接的证据。

**第二,HBM 定价权回归。**SK 海力士、三星、美光正迎来十年来最强的定价地位。HBM 与 AI 存储的溢价正在重塑存储厂商的盈利结构,也推高了所有 AI 芯片的成本曲线。

**第三,封装产能比代工竞争更稀缺。**GF 与台积电的互供协议说明,当先进封装产能不足时,代工竞争可以让位于产能协作。对芯片厂商而言,封装与 HBM 配额正在成为出货的前提条件。

**第四,对采购方的影响。**HBM 配额、DRAM 供给优先级都在向服务器和 AI 基础设施倾斜,终端设备的内存配置与成本将长期承压。这也与此前的国产芯片涨价潮相互呼应——存储成本上涨正在沿着供应链向所有 AI 算力产品传导。

小结​

台积电 Q3 营收约 NT$1.49 万亿、三星营业利润大增近 8 倍、12 层 HBM4E 通过 NVIDIA 认证、GF 为台积电代工 interposer——这一轮财报季确认了一个事实:内存超级周期已全面爆发,AI 算力的利润池正在加速向存储与封装环节迁移。对芯片厂商而言,HBM 配额就是出货许可证;对采购方而言,锁定存储供给正在成为与锁定 GPU 同等重要的战略动作。

Intel年内第三次涨价:CPU上调10%,AI需求下的产能博弈

· 4 min read
Industry Research Team

2026年10月5日起,Intel CPU正式上调价格10%。这是自2025年以来的第三次涨价——当AI吞掉大量先进制程与封装产能,连最基础的CPU也开始变得紧俏。"比涨价更糟的是没货",戴尔CEO的这句话,或许是这一轮周期最好的注脚。

涨价事实:年内第三次,涨幅10%​

根据2026年9月18日AI日报披露的信息,Intel计划自2026年10月5日起将CPU价格上调10%。值得注意的是,这已是自2025年以来的第三次涨价。涨价的驱动力主要来自两个方面:

  • AI需求外溢:AI服务器的爆发式增长挤占了先进制程晶圆、先进封装(如CoWoS类产能)与HBM之外的配套资源,传统CPU的供应与产能分配同步承压;
  • 产能紧张:半导体结构性短缺并非短期现象,而是供需结构错位的结果。

资本市场:涨价周期中的芯片股​

涨价消息所处的市场环境同样值得关注。同期美股芯片板块整体上涨:NVIDIA收涨2.54%,AMD涨超6%。华尔街对三巨头的上行空间预期如下:

公司华尔街预计上行空间
NVIDIA51.61%
AMD25.78%
Intel16.33%

可以看出,市场对AI纯度更高的公司给出了更高的预期上行空间,Intel的16.33%虽相对保守,但涨价本身也在改善其盈利预期。

供给端的判断:2027年可能更严重​

戴尔CEO迈克尔·戴尔在高盛大会上给出了一个更沉重的判断:半导体结构性短缺短期内难以缓解,2027年可能比2026年更严重。他的原话是:"比涨价更糟的是没货。"

这句话对整个产业链的启示在于:涨价是供需失衡的结果而非原因,如果2027年供给进一步收紧,那么当前的10%涨幅可能只是开始。

Intel的AI芯片线:涨价之外的另一条战线​

在CPU涨价的同期,Intel的AI芯片布局也在推进,主要包括:

  • Gaudi 3 / Gaudi 4:AI加速器主线产品;
  • Jaguar Shores:下一代GPU,接替此前被取消的Falcon Shores路线;
  • Crescent Island:面向推理场景的GPU产品。

在服务器侧,Intel还面临AMD的正面竞争——AMD EPYC "Venice"(Zen 6架构)最高已达256核,x86服务器CPU的核数竞赛与AI服务器配套需求叠加,进一步加剧了先进产能的争夺。

涨价如何传导:整机与云价格​

CPU作为几乎所有服务器的核心部件,其10%的涨幅不会停留在芯片层面,而是沿着"芯片→整机→云服务"的链条逐级传导:

  1. 整机厂商:服务器OEM将面临BOM成本上升,戴尔等厂商已明确表态供给紧张,整机报价上调是大概率事件;
  2. 云厂商:云主机的计算实例价格与裸金属服务器租金,最终会反映上游芯片成本;
  3. AI基础设施:AI服务器同样需要CPU作为宿主与调度核心,GPU集群的配套CPU成本同步上升。

采购方应对建议​

面对"涨价+缺货"的双重压力,采购方可以考虑:

  • 提前锁量:在2027年供给可能进一步收紧的预期下,尽早与厂商或渠道签订量价协议锁定供应;
  • 关注国产替代选项:在x86服务器CPU与AI配套计算领域,国产芯片(如华为昇腾、寒武纪等)正在快速补位,可作为成本与供应风险的对冲选项;
  • 重新评估TCO:涨价环境下,单纯比较芯片单价意义下降,应转向整机级、集群级的总体拥有成本测算。

小结​

Intel年内第三次涨价10%,本质上是一场AI需求驱动下的产能博弈:AI吞噬先进产能,传统CPU被迫让位并提价,而戴尔CEO"2027年可能更严重"的判断,预示这场博弈还远未到终局。对采购方而言,提前锁量、关注国产替代、以TCO视角重新测算成本,是当下最务实的三步棋。

Hot Chips 2026 Full Recap: Rubin, MI455X, Crescent Island Together as AI Compute Delivery Enters the "System-Level" Era

· 7 min read
Industry Research Team

August 23-25, 2026, the 38th Hot Chips (HC38) was held at Stanford's Memorial Auditorium. As the bellwether of global high-performance chip architecture, this conference landed exactly at the most intense moment of the AI compute arms race — the official agenda had 48 entries, including 7 AI accelerators, 6 memory tutorials, 6 CPUs, and 4 each of GPUs and networking. Putting the vendor talks together, one consensus emerged: the unit of AI compute competition has shifted from "single chip" to "whole rack / entire system."


1. Overview: Three Days of Agenda, Almost a Preview of the 2027 AI Rack Market​

Monday (8/24) afternoon's GPU session was the focus, with four talks nearly colliding as the 2027 AI rack market:

  • NVIDIA Rubin GPU ("Driving the Era of Agentic AI"): First chiplet-architecture GPU, 288GB HBM4, ~50 PFLOPS FP4, paired with 88-core Arm-architecture Vera CPU into NVL72 / NVL144 racks, mass production in H2 2026.
  • AMD Instinct MI400 (two talks: architecture + system architecture): Told the "rack-scale" story thoroughly.
  • Intel Crescent Island: A 350W air-cooled card designed for Agentic AI inference.

Tuesday (8/25) afternoon's AI session was almost a parade of "hyperscalers de-NVIDIA-izing": Google's 8th-gen TPU, OpenAI's first custom chip, Microsoft Maia 200, Meta MTIA, and Cerebras wafer-scale rack all appeared together.

Every vendor on stage used the term "Agentic AI" within the first two PPT slides — not a coincidence, but the collective shift in 2026 AI workload design goals.


2. NVIDIA Rubin: One Rack Is a Supercomputer​

What NVIDIA featured at Hot Chips was not a single GPU but the Vera Rubin NVL72 whole cabinet — 72 Rubin GPUs + 36 Vera CPUs, 18 compute trays + 9 NVLink switch trays, about 1.3 million components, nearly 1,300 chips, weighing about 4,000 pounds (~1.8 tons).

The single Rubin GPU specs are equally stunning:

MetricRubin GPUvs Blackwell
Transistors336 billion (TSMC 3nm dual-die)208 billion (+61.5%)
Memory288GB HBM4—
Bandwidth22 TB/s2.8× Blackwell
NVFP4 inference50 PFLOPS5× GB200
Training compute35 PFLOPS3.5×

The most disruptive design is in the compute tray: no cables, no hoses, no fans, all interconnected via the PCB backplane. NVIDIA says assembly time dropped from nearly 2 hours to 5 minutes (20× faster) while improving maintainability.

This time NVIDIA is selling not FLOPS but tokens per megawatt. Citing a SemiAnalysis benchmark based on DeepSeek-v4-PRO (140K+ context, AgentX workload), it claims: versus GB300 NVL72, Vera Rubin NVL72 delivers 10× to up to 30× tokens/MW as interaction intensity rises. A single cabinet provides 3.6 EFLOPS inference compute, whole-cabinet power 190-230kW; long-term capacity target is 1,000 NVL72 cabinets per day.


3. AMD MI455X + Helios: Bigger Memory and Open Interconnect​

AMD's answer is the MI455X + Helios rack going head-to-head with NVIDIA. MI455X uses CDNA 5 architecture, 8 N2-process accelerator dies + N3P-process interconnect die, 256 workgroup processors, 192MB global L2.

MetricMI455Xvs Rubin
Memory432GB HBM4 (12-layer stack)50% higher than Rubin's 288GB
Bandwidth23.3 TB/sSlightly ahead
MXFP4 compute40.26 PFLOPS—
System (Helios 72 cards)2.9 ExaFLOPS FP4 inference—
Price~$5.25M per cabinet—

At the system level, AMD bets on the UALoE (Ultra Accelerator Link over Ethernet) open standard: each GPU provides 3.6 TB/s bidirectional interconnect bandwidth; two 512-port 200G UALoE switch chips in the switch tray total 10.8 TB/s — opening the interconnect protocol to the whole industry while targeting NVLink.

Production cadence: AMD plans to deliver engineering samples and small-batch systems in H2 2026, with large-scale ramp in Q2 2027. Earlier rumors of Helios delay due to cooling issues were not confirmed by AMD.


4. Intel Crescent Island: The Air-Cooled, Large-Memory "Cost-Effective Oddball"​

Intel offers a completely different path: Crescent Island — a 350W, air-cooled, standard-PCIe-slot inference GPU designed for Agentic AI, with the key metric being tokens per watt.

MetricCrescent IslandNote
ArchitectureXe3P, 32 Xe cores, 32MB unified L2Disclosed at Hot Chips
MemoryIntel branded card 160GB / ODM up to 480GB LPDDR5XMore than Rubin's 288GB HBM4
Form factor350W air-cooled PCIePlugs into standard racks, no liquid-cooling retrofit
RASECC, dynamic page offline, hard-package repair, PCIe advanced error reportingAddresses "silent data corruption"

Intel's logic is clear: inference scenarios need far more memory capacity than bandwidth; using low-cost LPDDR5X for capacity and air cooling to skip liquid-cooling infrastructure drives down per-token cost. Combined with Diamond Rapids Xeon (256 performance cores, 1.28GB cache, 128 PCIe Gen6 lanes), Intel tries to surround from edge to datacenter with "CPU + inference GPU + open software stack."


5. Custom ASIC Parade: Google, OpenAI, Microsoft, Meta Together​

Tuesday afternoon's AI session was the most historic of the conference — a parade of "hyperscalers de-NVIDIA-izing":

ChipVendor / PartnerPositioningKey Specs / Progress
TPU 8t (Sunfish)Google × BroadcomTraining9,600 cards per pod, 121 FP4 ExaFLOPS, 2PB shared HBM
TPU 8i (Zebrafish)Google × MediaTekInference288GB HBM, 384MB on-chip SRAM (3× prev gen), ICI 19.2 Tb/s
JalapeñoOpenAI × BroadcomInference9-month end-to-end design, target ~50% token cost cut, commercial end of 2026
Maia 200Microsoft (TSMC 3nm)Inference140B+ transistors, 10+ PFLOPS FP4, 216GB HBM3E, serving GPT-5.2 at Des Moines datacenter
MTIA 300-500Meta (RISC-V) × BroadcomTraining + inferenceUp to 25× compute gain, one model every 6 months before 2027

Google split TPU into training (8t) and inference (8i) dedicated architectures for the first time — its biggest architectural shift in a decade. Norm Jouppi personally took the stage to present TPU v8.


6. Two Hidden Threads — Memory and Networking: HBM4 Year 1 + AI Factory OS​

Beyond GPUs/ASICs, two hidden threads mattered equally:

  • Memory: Samsung's HBM Base Die (logic-process base die) and SK hynix's advanced packaging appeared together; the HBM4-era "base-die foundry" industry shift begins; HBF (high-bandwidth flash), LPDDR5X-PIM, 3D DRAM, and CXL compute-storage showcased "compute-in-memory" moving from papers to products.
  • Networking: NVIDIA BlueField-4 (DPU) and Spectrum-X Multiplane architecture (presented by Gilad Shainer) — networking is becoming the decisive architecture for gigascale AI, scaling from hundreds of thousands to a million cards; Broadcom Thor Ultra Ethernet NIC keeps pressing; Mojo Vision showed chip-level optical I/O.

7. Three Routes, One Consensus​

At the same conference, three vendors offered three distinctly different AI compute delivery philosophies:

  1. NVIDIA: Full-stack closed integration — GPU, CPU, DPU, and switch chips all self-designed, pushing system performance to the extreme via ultimate software-hardware co-design, at the cost of deep customer lock-in.
  2. AMD: Open-standard catch-up — Uses larger HBM4 capacity + UALoE open interconnect for a "cost-effective + open" play, tearing open the inference gap with Meta and OpenAI's 12GW-class orders.
  3. Intel: Air-cooled cost-effectiveness — Abandons liquid cooling and HBM, uses LPDDR5X large memory + standard PCIe, betting that "most inference doesn't need a 200kW rack."

But all three agree: the unit of competition is no longer the chip, but the co-designed system (rack / system). For buyers, 2027 compute planning should compare not "single-card PFLOPS" but "tokens per megawatt, latency, availability, and full-lifecycle cost."

References​


This article is compiled from Hot Chips 2026 (Aug 23-25) official presentations and on-site reports from ServeTheHome, SemiAnalysis, TechPowerUp, etc. Performance data are vendor-disclosed figures; actual performance subject to mass-produced products.

HBM4 Mass-Production Year One: Samsung Yield Breaks 80%, Three Giants Pass NVIDIA Certification, the Last Bottleneck of AI Compute Supply

· 6 min read
Industry Research Team

If 2025 was the year of HBM3E capacity ramp-up, then 2026 is year one of HBM4 mass production. With NVIDIA Vera Rubin and AMD MI400 — two generations of flagship — both betting on HBM4, this "memory on the AI chip" has for the first time become a strategic commodity that dictates the delivery pace of entire racks. The yield and certification data disclosed densely in August is rewriting the global HBM supply map.


1. Golden Yield Breakthrough: Samsung Jumps from Under 60% to 80% in Six Months​

Per South Korea's Seoul Economic Daily on August 9, Samsung Electronics' HBM4 yield officially crossed the 80% "golden yield" threshold in early August — more than four months ahead of its original year-end target.

TimelineSamsung HBM4 YieldNotes
Feb 2026 (mass production start)Under 60%Line ramp-up period
Early Aug 2026~80%Crosses the mass-production / stable-profit watershed

The semiconductor industry has long held that "80% yield is the golden yield" — it is both a yardstick of foundry competitiveness and the financial break-even point for large-scale commercial supply. The key to this leap was Samsung's breakthrough in Thermal Compression Non-Conductive Film (TC-NCF) bonding, plus the stable base of its underlying 1c DRAM yield, already above 80%. In the same period, Samsung's HBM4E reliability test yield also broke 70%.

Industry assessments suggest SK Hynix's HBM4 yield has likewise entered the 80% range. The gap between the two giants in production quality is being rapidly erased.


2. Supply Map: SK Hynix Holds 60–70% of Rubin Allocation​

At a Seoul event on June 5, Jensen Huang publicly confirmed: Samsung, SK Hynix, and Micron have all passed HBM4 certification for Vera Rubin — the first time three memory makers have simultaneously received public certification for the same platform.

But certification is just the "entry ticket" — allocation share is where the real voice lies:

Vendor2026 Rubin HBM4 Allocation (est.)Notes
SK Hynix60%–70%Based on HBM3/3E-era customer relationships and MR-MUF packaging
Samsung25%–30%Rapid share gains after yield leap
MicronRemainderLimited HBM4 exposure, relatively stable share

Counterpoint Research forecasts the 2026 HBM4 market as SK Hynix 54% / Samsung 28% / Micron 18%. Samsung has set staged catch-up targets: Q3 HBM4 revenue up 3× QoQ, HBM4 exceeding 60% of total HBM revenue in H2, and year-end overall HBM market share approaching 38%.


3. The Real Bottleneck: From Wafers to "Back-End Stacking"​

As front-end yield stabilizes, the rhythm of the AI accelerator supply chain no longer depends on "how many wafers can be made," but on the speed of back-end stacking, bonding, testing, and shipment.

  • Industry analysts rank HBM stacking as the second-most severe bottleneck in the AI chip supply chain, second only to TSMC's CoWoS advanced packaging capacity.
  • HBM accounts for roughly 25% of 2026 DRAM wafer output; each HBM wafer consumes about 3–4× the resources of a standard DRAM wafer (extra TSV and stacking steps), so every wafer redirected pulls 3–4 units of commodity memory off the spot market.
  • Samsung is considering relocating part of its legacy memory back-end lines (Cheonan, Onyang) to Vietnam to free up HBM back-end capacity — a side confirmation that back-end throughput is now the tightest link in the chain.

4. HBM4 Spec Snapshot: Generational Leap in Bandwidth and Efficiency​

SpecHBM4 (12-Hi / 16-Hi)HBM4E
Per-stack capacity36 GB / 48 GB—
Pin rate11.7–13.0 Gbps16 Gbps
Per-stack bandwidthup to 3.3 TB/sup to 3.6 TB/s
Bus width2048-bit—
Energy efficiency+40% vs HBM3E—
Thermal resistance / cooling+10% improvement / +30%—

Samsung HBM4 entered mass production in Feb 2026; its 11.7 Gbps pin rate already exceeds the 8 Gbps industry baseline required for Vera Rubin compatibility; HBM4E samples were first shipped to major customers on May 29.


5. Pricing Power Extends Into 2027: Supply Remains Tight Balance​

TrendForce judges that HBM suppliers' pricing power will run through 2027, because supply remains constrained:

  • 2027 HBM bit shipments are expected to grow 50%–60% YoY, but will still lag demand growth, keeping the market tight;
  • The industry already anticipates significant price increases;
  • For NVIDIA and AMD, a stronger Samsung means more supply options and more comfortable lead times — in a market where memory is the tightest link in AI servers, the mere existence of second and third suppliers is itself a buffer.

For entire racks, HBM cost is already the biggest driver: the Rubin Ultra rack carries an estimated price tag as high as $21 million, with HBM making up a substantial portion.


6. Lessons for China: HBM Export Controls Accelerate Domestic Iteration​

HBM is one of the core fronts of current AI chip controls. As the overseas HBM4 arms race intensifies, domestic HBM technology iteration is being pushed forward in sync — Huawei's Ascend roadmap has explicitly written "drive domestic HBM technology iteration" into its product cadence (the 950 series advances domestic HBM pairing, with the 960/970 series planned for gradual rollout in 2027–2028).

In the short term, HBM4 scarcity will directly transmit to the delivery cadence of Rubin / MI400; in the long term, whoever can lock in stable HBM4 supply holds the valve on 2027 AI compute expansion.

References​


This article is compiled from August 2026 public reports by TrendForce, Seoul Economic Daily, TechTimes, etc. HBM allocation shares and market shares are third-party estimates, not official vendor-confirmed data.

NVIDIA Vera Rubin Officially Ships: First VR200 NVL72 Delivered, Samsung HBM4 Mass Production, Rubin Ultra Cabinet Sky-High Price

· 5 min read
Industry Research Team

July 2026, NVIDIA's next-gen AI compute platform Vera Rubin officially began its first shipments, succeeding the Blackwell architecture, with large-scale mass production planned for H2 2026. First customers include Microsoft, Google, Amazon, Meta, Oracle, and other large cloud providers.

1. World's First VR200 NVL72 Delivered (Milestone)​

CoreWeave jointly with Dell announced that the world's first NVIDIA Vera Rubin VR200 NVL72 cabinet has been officially delivered and passed the L11 full-cabinet hardware diagnostics on the first try. This marks Rubin's move from roadmap to physical product, with no major bottlenecks in core supply-chain links (HBM4, advanced packaging, liquid cooling, ultra-high-power power supply).

VR200 NVL72 Core Configuration​

MetricVera Rubin VR200 NVL72
Cabinet codenameOberon
GPU72 Rubin GPUs
CPU36 Vera CPUs
Per-GPU memory288 GB HBM4
Per-CPU memory1.5 TB LPDDR5X
Total cabinet HBM420.7 TB (20,736 GB)
Total cabinet LPDDR5X54 TB
InterconnectNVLink 6 full mesh
Inference performance~3.6 exaFLOPS class
CoolingLiquid cooling
Generational improvement~3.5× per-GPU compute, ~2.8× memory bandwidth (vs Blackwell)

Vera CPU integrates 88 custom Olympus ARM cores, with 1.8 TB/s interconnect to the GPU, usable as a GPU memory expansion pool. NVIDIA completed its first Vera CPU deliveries to Anthropic, OpenAI, xAI, and Oracle Cloud in May.

2. Samsung HBM4 Mass Production: Key Bottleneck Eases​

July 8, 2026, Samsung Electronics officially started HBM4 mass production for the Vera Rubin platform, with reported HBM4 mass-production yield reaching 70% (above the initial 60-65% expectation). Confirmation of this key supply-chain link clears obstacles for Rubin's large-scale deployment.

HBM Supply Landscape (2026 Q1)Share
SK hynix45%
Samsung40%
Micron15%

HBM4 uses 8-layer stacking (12-layer design planned for 2028), priced at about 2.8× HBM3e. TrendForce predicts HBM supply will grow 65% annually, with HBM4 reaching 35% of total output by 2027 Q4.

3. Rubin Ultra Sky-High Price: HBM Cost Dominates​

Per BofA Global Research estimates, the Rubin generation will push single-server cost to a new high:

Cost ItemRubin VR200 (Oberon)Comparison
Cabinet HBM4 usage20,736 GB—
HBM4 unit price~$18.40 / GBBlackwell (HBM3e) ~$11.26 / GB
HBM4 cost alone~$382KExcluding LPDDR5X
Rubin Ultra cabinet estimated price~$21MITHome / BofA estimate

4. Rubin Ultra Design Change: Original 4-die Cancelled (per SemiAnalysis)​

Semiconductor research firm SemiAnalysis (2026-06-30) disclosed that the original 4-die Rubin Ultra GPU unveiled at GTC 2026 has been cancelled; the version actually shipping in 2027 is roughly halved in scale and performance:

  • Reason for cancellation: The original integrated 4 compute dies + 16 HBM4E in a single CoWoS-L package; the substrate warped under the 4-die config, causing compute-die-to-substrate contact failure and yield collapse; the alternative CoPoS won't reach mass production until after late 2028, missing the 2027 node.
  • New approach: Changed to dual-die (same construction as standard Rubin) + HBM4E, ~384 GB HBM4E per GPU (higher than standard Rubin's 288 GB), but total compute and bandwidth only half the original; to approach the original's aggregate compute, NVIDIA plans to assemble "2+2" board-level configs within the Kyber rack to reach four-die equivalent scale.
  • Kyber rack delay: The companion Kyber NVL144 rack is delayed 12+ months to 2028 due to midplane PCB manufacturing difficulties; the 800V DC power scheme is likewise delayed to 2028.

⚠️ Note: NVIDIA has not commented officially on the above design change; some on X argue "the chip count hasn't changed, it's old news reheated." This section is compiled from SemiAnalysis public reports, subject to final NVIDIA disclosure. We have marked "specs pending official confirmation" on the Rubin Ultra preview card.

Industry Interpretation​

  1. "Never doubt" moment realized: Rubin's first delivery passed L11 on the first try, dispelling market doubts about "Rubin delay," locking in H2 2026 AI compute supply certainty ahead of time.
  2. Designed for Agentic AI: Rubin targets agentic workflows and ultra-long-context inference, further lowering the training/inference cost curve for trillion-parameter models.
  3. HBM is the full-chain winner: 20.7 TB HBM4 per cabinet is enormous usage; SK hynix, Samsung, Micron, advanced packaging (CoWoS-L), liquid cooling, and power retrofitting all benefit across the chain, while also becoming the biggest cost and capacity constraint.

References​


This article continuously tracks Vera Rubin mass-production ramp and HBM4 supply-chain dynamics.

AI Hardware Enters the "Era of Deployment": Five Major Shifts of 2026 and the Rules for Survival

· 9 min read
Industry Research Team

In 2026, the AI hardware market is undergoing a fundamental shift from the "training race" to "deployment as king." As large models move from technology demos to large-scale commercial deployment, hardware form factors, technology roadmaps, and the competitive landscape are undergoing systematic change.

Publisher: CSHIA Research (中智盟咨询) Author: Zhou Jun

Trend 1: Shift in compute demand structure — inference becomes the main engine of growth​

The biggest change in the 2026 AI hardware market is the shift in the center of gravity of compute demand from training to inference.

According to market data:

  • In 2026, global AI inference compute demand is expected to grow over 60% year-over-year
  • Inference compute will exceed training compute for the first time, becoming the dominant workload of AI infrastructure

This shift stems from AI applications moving from "model development" into the "large-scale deployment" stage — enterprises no longer train large models frequently, but instead transform AI capability into real business value through high-frequency inference calls.

Key manifestations​

  1. Inference chip market explosion: Shipments of dedicated inference chips (ASICs) are expected to grow 129%, with their share of AI servers rising from under 20% in 2025 to 27.8%.

  2. Cost structure optimization: NVIDIA's Rubin platform reduces inference token cost to 1/10 of the previous generation, pushing inference applications from "luxury" to "commodity."

  3. Workload characteristics change: Inference tasks show "high-frequency, long-pipeline, low-latency" characteristics, demanding higher real-time responsiveness from hardware.

Latest GTC 2026 developments (June 1, Taipei)​

NVIDIA CEO Jensen Huang announced several major inference compute advances at GTC 2026 Taipei:

  • Vera Rubin platform enters full production: The NVL72 rack system delivers agentic throughput 10× that of the previous-generation Grace Blackwell, designed for Agentic AI
  • Vera CPU officially launched: 88-core Armv9.2 custom Olympus architecture, highest single-thread IPC in the world, 1.5TB LPDDR5X memory, 1.2 TB/s bandwidth, native FP8 support
  • RTX Spark AI PC chip: Co-developed with MediaTek (codename N1X), Blackwell-architecture GPU with 1 PFLOP AI compute, 128GB unified memory, TSMC 3nm, reshaping the Windows PC ecosystem
  • AI Factory platform DSX: Four components — DSX Sim (digital-twin simulation), DSX OS (resource orchestration), DSX MaxLPS (power optimization), DSX Flex (grid coordination)

This trend means the competitive focus for hardware vendors is no longer "peak single-card compute" but "inference energy efficiency" and "system-level optimization capability."


Trend 2: Edge and on-device AI — the scaled deployment of compute moving downstream​

2026 is the pivotal year for edge AI hardware moving from proof-of-concept to scaled deployment.

As cloud inference cost pressure rises and privacy compliance requirements tighten, compute is accelerating its migration toward data sources, spawning explosive growth in hardware form factors such as edge servers, AI terminals, and smart devices.

Three deployment scenarios​

ScenarioHardware formCore characteristics2026 market size forecast
Edge serversCompact cabinets, edge compute nodesPower density 40-80kW/cabinet, liquid cooling supportedGlobal shipments grow 28%
AI terminalsAI phones, AI PCs, smart glassesOn-device NPU compute 60+ TOPS, offline inference1.5 billion units shipped
IoT devicesSmart cameras, sensors, robotsLow-power chips, real-time responseMarket size exceeds $1.5 trillion

Technology breakthroughs​

  1. On-device model compression: Through quantization, distillation and other techniques, models with tens of billions of parameters are compressed to run on-device.

  2. Heterogeneous compute architecture: CPU+NPU+GPU coordination maximizes performance under power constraints.

  3. Memory bandwidth optimization: Application of HBM technology in edge chips alleviates the "memory wall" problem.

The edge AI explosion means hardware design must balance "performance density" with "power efficiency," and traditional general-purpose chips face specialization challenges.


Trend 3: Dedicated chips and heterogeneous computing — breaking the monopoly of a single architecture​

In 2026 the AI chip market will show a "one superpower, many strong players, a hundred flowers blooming" competitive landscape.

Although NVIDIA maintains its advantage in training, in segmented markets such as inference, edge, and specific scenarios, dedicated chips (ASICs) and heterogeneous computing solutions are rising rapidly.

Major technology roadmap comparison​

Chip typeRepresentative vendorsCore advantageApplicable scenarios
General-purpose GPUNVIDIA, AMDMature ecosystem, flexible programmingCloud training, complex inference
Dedicated ASICGoogle TPU, CambriconHigh energy efficiency, cost advantageLarge-scale inference, specific algorithms
Compute-in-memoryMultiple startupsBreaks the "memory wall," low latencyEdge inference, real-time processing
FPGA/DPUXilinx, HuaweiReconfigurable, high flexibilityNetwork acceleration, data preprocessing

Market landscape changes​

  1. Domestic substitution accelerates: China's AI chip vendors raise their share in inference, edge and other scenarios to over 30%.

  2. Open-source ecosystem rises: Open-source frameworks such as ROCm and OpenML lower the barrier to dedicated-chip development.

  3. Chiplet technology popularizes: Integrating chips of different process nodes through advanced packaging achieves a balance of performance and cost.

  4. GTC 2026 new products accelerate deployment (June 1, Taipei):

    • Vera Rubin platform: NVL72 rack system, agentic throughput 10× Grace Blackwell
    • Vera CPU: 88-core Olympus custom architecture, designed for Agentic AI low latency
    • RTX Spark: In partnership with MediaTek and Microsoft, reshaping the Windows PC ecosystem, 1 PFLOP AI compute
    • Nemotron 3 Ultra: SSM+MoE hybrid architecture, 5× faster inference, 30% lower cost

The core logic of this trend is: no single chip can dominate all AI scenarios; scenario fragmentation spawns technology-roadmap diversification.


Trend 4: Energy efficiency and thermal management — from technical challenge to business bottleneck​

As AI chip power consumption breaks the kilowatt level (NVIDIA Rubin GPU reaches 2300W), energy efficiency and thermal management have been upgraded from "supporting technology" to "core bottleneck."

In 2026, single-cabinet power density will exceed 240kW, traditional air cooling completely fails, and liquid cooling changes from "optional" to "mandatory."

Key data​

  • Power cost share: The share of power cost in AI data center operating cost rises from 15% to 35%
  • Thermal value increases: A single GB300 server's liquid-cooling components are worth about $50,000, 15-20% of hardware cost
  • PUE optimization: Liquid-cooled data centers can bring PUE down to under 1.1, but upfront investment rises 30%

Technology evolution directions​

  1. Tiered liquid cooling: Cold-plate (mainstream), immersion (high density), two-phase cooling (frontier)

  2. Power architecture upgrade: From 12V to 48V/800V high-voltage DC, reducing conversion losses

  3. Intelligent thermal management: AI predictive cooling, dynamically adjusting cooling strategy based on load

This trend means a hardware vendor's competitiveness depends not only on chip performance but more on "system-level energy efficiency optimization capability"; the importance of supporting technologies such as thermal management, power delivery, and cabinet design rises substantially.


Trend 5: AI-native hardware ecosystem — from "compatibility" to "reconstruction"​

In 2026, AI hardware is undergoing a paradigm shift from "adapting to AI" to "built for AI."

Traditional general-purpose hardware architectures struggle to meet the unique demands of AI workloads, spurring the rise of AI-native hardware design philosophy.

Three reconstruction directions​

1. Compute architecture reconstruction​
  • Memory hierarchy optimization: HBM4 memory bandwidth breaks 3TB/s, compute-in-memory architecture reduces data movement
  • Interconnect upgrade: NVLink 6.0 reaches 1.8TB/s bandwidth, supporting direct GPU-to-GPU communication
  • Heterogeneous integration: Through advanced packaging, CPU, GPU and memory are stacked to boost bandwidth and reduce latency
2. Software-defined hardware​
  • Reconfigurable logic: FPGA and DPU support dynamic algorithm loading, adapting to different AI models
  • Compiler optimization: AI compilers (e.g., MLIR) automatically optimize hardware resource allocation
  • Hardware abstraction layer: Unified programming interfaces shield underlying hardware differences
3. Ecosystem co-evolution​
  • Model-hardware co-design: Large-model architectures account for hardware constraints (e.g., sparsification, quantization)
  • Open-source hardware design: Application of RISC-V in AI chips lowers the development barrier
  • Vertical integration: Cloud vendors' self-developed chips (e.g., AWS Graviton, Google TPU), software-hardware co-optimization

The essence of this trend is: the characteristics of AI workloads (matrix operations, high parallelism, memory sensitivity) are redefining hardware design principles, and the universality advantage of traditional x86 architecture is weakened in AI scenarios.


Key Conclusions and Outlook​

The inference demand explosion drives edge deployment, edge scenarios spawn dedicated chips, high power consumption forces an energy-efficiency revolution, and all changes ultimately point to the reconstruction of the AI-native hardware ecosystem.

The core driver of this round of change is AI moving from "technology demo" to "commercial deployment"; hardware must satisfy the industry requirements of "scale, low cost, high reliability."

2. Opportunity windows for industry participants​

For industry participants, the opportunities in 2026 lie in:

  • ✅ Capture the inference dividend: Deploy inference-specific chips and system optimization
  • ✅ Deepen vertical scenarios: Customize hardware solutions for specific industries/applications
  • ✅ Break the energy-efficiency bottleneck: Liquid cooling, high-voltage DC, AI thermal management and other technologies
  • ✅ Build an open ecosystem: Open-source frameworks, open standards, cross-industry collaboration

Vendors that can provide "end-to-end solutions" rather than "single-point chips" will gain an advantageous position in this reshuffle.

3. Dynamic adjustment and continuous evolution​

The above analysis is based on early-2026 market data and industry forecasts; actual development may adjust dynamically due to factors such as technology breakthroughs, policy adjustments, and market demand changes.


Industry Implications​

2026 is a watershed year for the AI hardware industry:

  • From "compute race" to "deployment as king"
  • From "single-point breakthroughs" to "system optimization"
  • From "general-purpose architecture" to "dedicated customization"
  • From "performance first" to "energy efficiency balance"

Vendors that can keenly capture trends, rapidly adjust strategy, and sustain technological innovation will seize the initiative in the AI hardware "era of deployment."


References:

  • CSHIA Research, "2026 AI Hardware: Five Transformations and the Rules for Survival"
  • "AI Hardware Enters the 'Era of Deployment'," Sohu Tech, February 10, 2026