封面
市場調查報告書
商品編碼
2100786

2026 年的發展趨勢:記憶體將滿足日益成長的人工智慧推理需求

2026 Trends: Memory for New AI Inference Demand

出版日期: | 出版商: TrendForce | 英文 11 Pages | 商品交期: 最快1-2個工作天內

價格
簡介目錄

2026年1月,NVIDIA發布了由BlueField-4 DPU管理的CMX上下文記憶體儲存平台。該平台擴展了本地SSD和共用儲存之間的記憶體層次結構,以滿足AI推理時代對海量鍵值快取儲存的需求。此外,NVIDIA和Arm相繼推出了CPU機架,以滿足基於代理的AI的CPU需求,從而開闢了CPU記憶體的新市場。

本報告詳細分析了以下三個面向:(1) 人工智慧推理的記憶體需求;(2) 由鍵值快取卸載驅動的固態硬碟 (SSD) 儲存單元 (POD) 需求;(3) 由基於代理的人工智慧驅動的 CPU 記憶體需求。其目的是解釋人工智慧推理時代記憶體容量需求不斷成長的原因,檢驗現有解決方案,並展望未來記憶體需求結構。

主要亮點

  • NVIDIA 推出了 BlueField-4 DPU 管理的 CMX 平台,該平台彌合了本地 SSD 和共用儲存之間的差距,可用於 AI 推理中的大規模鍵值快取。
  • NVIDIA 和 Arm 推出了 CPU 機架,以滿足基於代理的 AI 的 CPU 需求,這正在逐步增加對 CPU 記憶體的需求。
  • 本報告重點關注 AI 推理中的記憶體需求、與 KV 快取卸載相關的 SSD POD 需求,以及不同基於代理的 AI 架構引起的 CPU 記憶體結構變化。

目錄

  • 1. 人工智慧推理中的記憶體需求
  • 2. 對 SSD POD 的需求將由 KV 快取卸載驅動。
  • 3. 基於代理的人工智慧增加了對 CPU 的需求
  • 4. TRI的觀點
簡介目錄
Product Code: TRi-196

In January 2026, NVIDIA introduced the CMX Context Memory Storage Platform, managed by the BlueField‑4 DPU, to extend the memory hierarchy between local SSD and shared storage and address the massive KV cache storage demands of the AI inference era. In addition, NVIDIA and Arm have successively launched CPU racks to meet the CPU requirements of agentic AI, creating an incremental market for CPU RAM.

This report provides an in‑depth analysis of: (1) memory demand in AI inference; (2) SSD POD demand driven by KV cache offloading; and (3) CPU RAM demand driven by agentic AI. The goal is to explain why memory capacity needs are expanding in the AI inference era, review current solutions, and outline the future structure of emerging memory demand.

Key Highlights

  • NVIDIA introduced the CMX platform managed by BlueField‑4 DPU to extend between local SSD and shared storage for large KV cache in AI inference.
  • NVIDIA and Arm launched CPU racks to meet agentic AI CPU needs, creating incremental CPU memory demand.
  • The report focuses on AI inference memory needs, SSD POD demand from KV cache offload, and CPU memory structure changes driven by agentic AI.

Table of Contents

  • 1. Memory Demand in AI Inference
    • Figure 1: AI Models Average Output Tokens per Question (2023-2026)
    • Figure 2: Example of KV Cache Applications
    • Figure 3: Changes to CPU:GPU Ratio among Agentic AI Applications
  • 2. SSD POD Demand Driven by KV Cache Offloading
    • Figure 4: Sequence of KV Cache Offloading for NVIDIA’s Dynamo (G1-G4)
  • 3. CPU Demand Driven by Agentic AI
    • Figure 5: NVIDIA’s Vera CPU Architecture
    • Table 1: CPU Specifications of Various Suppliers (2023-2026)
    • Table 2: Analysis on Hypothetical Shipment Scenario of NVIDIA’s CPUs in 2026
    • Figure 6: Analysis Results on Demand Scenario of NVIDIA’s CPUs in 2026
    • Table 3: Summary of Memory Demand Drivers Introduced by AI Inference
  • 4. TRI’s View