---
格式版本: 2
标题: "Vistara for Hyperscale Efficiency.pptx"
原文链接: "https://computeexpresslink.org/wp-content/uploads/2026/07/Vistara-for-Hyperscale-Efficiency.pptx-1.pdf"
发布日期: "2026-07-21"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:scrape:strict_markdown_body"
发布时间证据: "July 21, 2026"
发布时间校准原因: "规则确认唯一严格发布时间，来源 scrape:strict_markdown_body"
发布时间校准置信度: "high"
发布时间候选数量: 2
发布时间严格候选数量: 1
发布时间原页读取状态: "原页面非文本类型 application/pdf"
发布时间未找到原因: ""
发布时间校准时间: "2026-07-26T02:14:19+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-07-26T02:08:10+08:00"
入库时间: "2026-07-25T18:14:19.103Z"
来源平台: "固定入口"
搜索渠道: "fixed_url"
搜索词: "https://www.computeexpresslink.org/"
匹配关键词:
  - "CXL"
  - "deployment"
  - "performance"
  - "latency"
  - "bandwidth"
  - "throughput"
相关厂家:
  - "Meta"
相关专家:
  []
内容类型: "PDF"
抓取工具: "PDF 下载"
清洗工具: "PyMuPDF 文本抽取"
原始附件:
  - "原文.pdf"
AI质检状态: "评分失败"
AI评分尝试次数: 1
AI评分错误类型: "service_error"
AI评分错误: "LLM 调用失败，已尝试模型链：deepseek-v4-flash -> deepseek-v4-pro｜模型 deepseek-v4-flash 未返回内容｜模型 deepseek-v4-pro 未返回内容"
AI评分开始时间: "2026-07-27T03:12:58.369Z"
AI评分结束时间: "2026-07-27T03:13:28.763Z"
采集批次: "2026年7月25日22点32分16秒"
采集批次ID: "20260725-223216-273"
去重键: "https://computeexpresslink.org/wp-content/uploads/2026/07/Vistara-for-Hyperscale-Efficiency.pptx-1.pdf"
---

# Scaling CXL® to Millions of Servers: Vistara for Hyperscale Efficiency

## 第1页

Scaling CXL® to Millions of Servers
Neha Gholkar and Hasan Al Maruf (Meta, Inc.)
July 21, 2026
CXL® Consortium 2026
Vistara for Hyperscale Efficiency

## 第2页

CXL® Consortium 2026
Memory is the Hyperscale Bottleneck
CXL promised to address all three at once
43%
of servers are memory-bound
not compute-bound
69%
of a server's embodied 
carbon footprint is DRAM
10–14yr
DRAM lifetime
servers retire at 5–7 yr
Fraction of Servers bottlenecked by resource (%)
Emissions by Component (%)

## 第3页

Six Years In– Where is CXL at Scale?
CXL® Consortium 2026
6 years
since CXL was introduced
yet limited reports of at-scale deployment.
The CXL Promise
One standard to expand memory capacity and bandwidth 
across the datacenter.
The Reality
Off-the-shelf solutions didn't ﬁt bill 
Software readiness to make CXL useful

## 第4页

It’s Architecture, Not the Protocol
Why Broad Adoption Has Been Slow?
CXL® Consortium 2026
Vistara, end-to-end CXL stack to turn CXL into production reality across millions of servers
CXL Memory Modules Bundled New DRAM
Required buying new DRAM chips with the controller– adds cost.
No DDR4 Support
Blocked reuse of decommissioned memory– killing the cost and carbon win.
Tiering is too Slow
A persistent myth: high tiering overheads and unstable tail latency rule out production.

## 第5页

What Fleet-Wide CXL Demands?
CXL® Consortium 2026
Cost-effective
Expansion
Lower
Carbon 
Easy
Adoption
Generic
& Broad
Operational 
Efficiency & 
Low 
Maintenance 
Overheads
Vistara: one co-designed stack that meets all the demands

## 第6页

Vistara: End-to-End CXL Stack
CXL® Consortium 2026
Software Linux Memory Tiering
The complexity of CXL is invisible to applications. Keeping tooling and maintenance simple 
ASIC Vistara
Power & cost efficient, performant expansion with DDR4 reuse
Platform MemServer
High memory-to-compute ratio for memory-bound services
One stack, co-designed end-to-end, running in production today

## 第7页

MemServer: 1TB Server
CXL® Consortium 2026
MemServer platform designed for memory constrained services with 3:1 local:CXL ratio
1 TB 
memory / server
3 : 1
local : CXL
250 ns 
CXL Idle Latency
+50 W
for both ASICs+DIMMs
Vistara ASIC (2x ASICs)
Host Connectivity x16 PCIe Gen5 (x8 used in prod)
Memory Interface 2x DDR4 2DPC channels
Conﬁguration
 4x 32GB DDR4 DIMMs 2400MT/s 
Production 
Operation 
Conditions
Local memory 
[60% BW util]
CXL Memory
[20% BW util]
234 ns
280 ns (+50ns)
48 GB/s
Peak Effective CXL BW

## 第8页

One Uniform Fleet– Opt Out As Needed
CXL® Consortium 2026
Transparent memory tiering. Minimal to no application changes. One fungible ﬂeet, Simple to run
Transparent Tiering
CXL appears as CPU-less NUMA node
Onlined as ZONE_MOVABLE by Linux Driver
Secondary tier via HMAT / CEDT
Node 
Full
CPU
CXL-Node
NUMA 1
Demote Cold Pages
Local Node
NUMA 0
Promote Hot Pages
TPP for Memory Tiering
OS allocates memory to local-ﬁrst
When local is full TPP demotes cold pages to CXL
Trapped hot pages in CXL gets promoted back
Hot working set mostly in local; cold pages in CXL
Special tiering-aware memory allocation policies 
beneﬁt speciﬁc workloads

## 第9页

One Uniform Fleet– Opt Out As Needed
CXL® Consortium 2026
Transparent memory tiering. Minimal to no application changes. One fungible ﬂeet, Simple to run
TMO for Proactive Demotion
User-space tool senpai monitors PSI and triggers reclamation
Stalls due to lack of memory monitored in the kernel space
Impact on a container adjusts the reclamation rate dynamically
Reclamation Pipeline
Local memory gets demoted to CXL, CXL gets offloaded to storage
user space
kernel space
Senpai
container 1
container 2
container N
…
pressure stall information
memory reclamation

## 第10页

One Uniform Fleet– Opt Out As Needed
CXL® Consortium 2026
Transparent memory tiering. Minimal to no application changes. One fungible ﬂeet, Simple to run
Resource Orchestrator
opt-out
Uniform Fleet
Uniform 1TB everywhere, 768GB local + 256GB CXL
Per-service opt-out is implemented at cgroup-level (cpuset.mems)
Service-level policies are speciﬁed at launch time
No BIOS changes or reboots needed
Metrics on resource usage, error rates, and device health, etc. are 
continuously reported to centralized monitoring systems
Service A
Service B

## 第11页

Local vs CXL Memory in Production
CXL® Consortium 2026
Practical trade-oƴs of deploying CXL memory shows its potential for cost-eƴective capacity expansion
Capacity
Local
CXL
768 GB
256 GB
BW
614 GB/s
58 GB/s
Power
132 W
30 W
Power/GB 1x
0.7x
Cost/GB
1x
0.13x

## 第12页

Memory Idle Time in Production
CXL® Consortium 2026
When workloads have large cold memory (memory with high idle time), CXL tier’s extra latency doesn’t impact much
22.5 seconds
28.3 minutes
1.3 hours
1.9 hours
P25
P50
P75
P99
Ads
4.3 minutes
19.4 minutes
43.8 minutes
1.4 hours
Cache
7.9 seconds
2.1 minutes
30.9 minutes
38.5 hours
Web1
4.2 seconds
1.7 minutes
27.1 minutes
72.9 hours
Web2

## 第13页

Already in Production
CXL® Consortium 2026
Server count reduction & DRAM reuse cut cost and carbon footprint
Development Infrastructure
More CI/CD jobs & VMs per server
−33%
ﬂeet size
ML Parameter Server
Fewer servers, less fan-out, fresher models
−25%
ﬂeet size
Cache
Bigger cache partitions cut backend traffic
−29%
avg. query 
processing time
+25%
throughput
−25%
ﬂeet size
Data Warehouse
+33% executor packing, fewer OOMs
−10%
execution time
−30%
ﬂeet size

## 第14页

Learnings from the Production
CXL® Consortium 2026
Stable CXL Latency 
270–370 ns for steady 
production workloads; tail 
latency variance is inline 
local memory 
Workload Impact
50–100 ns latency gap 
between local and CXL in 
production is imperceptible 
to the workloads
OS-Based Tiering
Simple OS-based tiering 
works; production design 
limits overheads under 0.5%
Hotness Detection
LRU-based hotness 
detection suffices to keep 
hot working set in local 
memory for a 3:1 setup

## 第15页

Vistara: Making CXL Real!
CXL® Consortium 2026
Co-design wins
Custom ASIC + transparent OS tiering, built production-grade solution.
Broad impact
Delivers performance gains across a wide range of real-world workloads.
CXL Memory expansion with DRAM Reuse pays
High memory-to-compute ratio achieved with DDR4 reuse dramatically cuts cost and carbon footprint landing huge TCO wins!
First at scale
The ﬁrst reported CXL memory expansion deployed at scale.

## 第16页

7/21/2026
16
Q&A
CXL® Consortium 2026

## 第17页

Thank You
CXL® Consortium 2026
