---
格式版本: 2
标题: "Agentic Flow with AMD PACE"
原文链接: "https://www.amd.com/en/developer/resources/technical-articles/2026/agentic-flow-with-amd-pace.html"
发布日期: "2026-09-03"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "llm:scrape:original_script_field"
发布时间证据: "datePublished: 2026-09-03T12:56:00-07:00"
发布时间校准原因: "候选c2为明确的datePublished字段，属于文章发布时间元数据，且无冲突。"
发布时间校准置信度: "1"
发布时间候选数量: 5
发布时间严格候选数量: 0
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-07T23:05:19+08:00"
发布时间仲裁状态: "confirmed"
发布时间仲裁尝试次数: 1
发布时间仲裁耗时毫秒: 3330
发现时间: "2026-09-07T22:07:21+08:00"
入库时间: "2026-09-07T15:05:31.728Z"
来源平台: "固定入口"
搜索渠道: "fixed_url"
搜索词: "https://www.amd.com/en/blogs.html"
匹配关键词:
  - "deployment"
  - "performance"
  - "latency"
  - "throughput"
  - "GPU"
  - "AI"
相关厂家:
  - "AMD"
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 41
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线为AMD PACE软件平台从推理引擎向Agentic AI编排器的演进，聚焦于LangGraph集成、CPU/GPU协同调度及WebVoyager/GAIA基准测试优化。虽为官方一手来源且具备一定技术细节，但核心内容属于软件栈、应用层编排和通用AI推理优化，未涉及超节点、AI Rack、机柜级硬件架构、互连、供电或液冷等机架级基础设施。命中强否决项：技术很深但只优化Serving、单一模型或单一工作负载，不能形成可复用机架级基础设施机制。"
AI质检模型: "qwen3.8-flash"
AI质检时间: "2026-09-07T23:06:37+08:00"
AI主题相关性: 2
AI来源权威性: 15
AI新颖性: 5
AI技术细节: 8
AI商业部署信号: 1
AI完整性: 10
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.100"
AI评分知识库SHA256: "7fe4568b9de9fc322c105ece08d18efab5b9b86f27dea6eb043b9a4d273bde60"
AI评分知识库检索词: "[\"AMD\",\"https://www.amd.com/en/blogs.html\",\"RAS\",\"NPU\",\"GPU\",\"PACE\",\"RAD\",\"EPYC\",\"PRO\",\"LLMs\",\"LLM\",\"KV\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0001\",\"title\":\"全球首颗2nm GPU来了！苏姿丰甩出“最强AI机架”，CPU性能干翻英伟达 - 智东西\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"AMD\",\"RAS\",\"NPU\",\"GPU\",\"EPYC\",\"PRO\",\"LLM\",\"KV\"],\"rank\":-18.342545089838076},{\"id\":\"july-correct-0065\",\"title\":\"AMD Pensando™ Vulcano 800 AI NIC: Built to Scale-Out and Across\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"AMD\",\"https://www.amd.com/en/blogs.html\",\"RAS\",\"GPU\",\"PACE\",\"RAD\",\"PRO\"],\"rank\":-14.45589120272508},{\"id\":\"july-correct-0068\",\"title\":\"Microsoft Azure Expanding AI Infra Choice with AMD Helios™\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"AMD\",\"https://www.amd.com/en/blogs.html\",\"RAS\",\"GPU\",\"RAD\",\"EPYC\",\"PRO\",\"LLM\"],\"rank\":-12.3231426425504},{\"id\":\"july-correct-0034\",\"title\":\"AMD Fires Back at Nvidia with Helios AI System, Epyc CPUs\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"AMD\",\"RAS\",\"GPU\",\"PACE\",\"EPYC\",\"PRO\"],\"rank\":-11.64721271639966},{\"id\":\"july-correct-0075\",\"title\":\"From EPYC to Helios, AMD and Meta are Scaling the Future of AI Together\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"AMD\",\"RAS\",\"GPU\",\"PACE\",\"RAD\",\"EPYC\",\"PRO\"],\"rank\":-11.099535594182303}]"
采集批次: "2026年9月7日22点05分18秒"
采集批次ID: "20260907-220517-361"
去重键: "https://www.amd.com/en/developer/resources/technical-articles/2026/agentic-flow-with-amd-pace.html"
---

## Introduction

AMD PACE (AMD Platform Aware Compute Engine) is an Open source Research and Development (RAD) project and a high-performance inference server that squeezes maximum throughput out of the AMD EPYC™ processor for Transformer workloads.

With this release, PACE takes its next major step: it graduates from being purely an inference engine to becoming a full agentic AI orchestrator. In this blog we describe why agentic workloads demand a new class of runtime, how PACE now uses LangGraph as its orchestration layer, and the four capabilities that make PACE a platform for building and serving agents:

1. Deterministic Agent Execution and Benchmarking
2. Flexible Deployment Across Local and Remote AI Infrastructure
3. Native Support for Custom LangGraph Agents
4. Optimized tool ops execution for end-to-end measurable performance gains

You can learn more about AMD PACE at [https://github.com/amd/AMD-PACE](https://github.com/amd/AMD-PACE).

## Overview: From Modeling to Reasoning to Agents to Multi-Agents

Applied AI has climbed a steady ladder of autonomy. It started with  **modeling** - training LLMs to predict the next token, where the goal was one good answer. Next came  **reasoning**, as techniques like chain-of-thought let models break down problems, think through steps, and check their own work. That led to  **agents**: models placed in a loop that can plan, call tools, and take real actions instead of just producing text. Today's frontier is  **multi-agent systems**, where specialized agents - planners, researchers, coders, and critics - work together to solve problems no single agent could handle alone. Every step up this ladder adds capability by running hundreds of connected LLM calls, tool calls, and decisions. PACE now enables orchestration for agentic flow.

#### LangGraph

LangGraph has become one of the most popular frameworks for building stateful, controllable agents. Instead of treating an agent as an opaque while-loop, it models the agent as an explicit graph: nodes are units of work (an LLM call, a tool call, a decision); edges define the flow between them. This structure gives developers first-class support for loops, branching, human-in-the-loop interrupts and persistent memory, so a run can be paused, inspected, resumed, or fanned out to sub-agents. That mix of clear structure, and durable state is what makes LangGraph a natural orchestration layer, and it's the foundation PACE builds on.

#### PACE as an Orchestrator

Originally, PACE optimized every layer of the LLM pipeline - scheduling, KV cache, attention and MLP kernels, to deliver leadership inference performance on AMD EPYC CPUs. That high-performance core stays.

What's new is the scope: **PACE now supports LangGraph as its native orchestration layer**, turning it from a single-model inference server into an engine that runs entire agentic graphs end to end. PACE executes the LangGraph state machine, dispatching nodes, routing edges, threading shared state, and managing tool calls, while transparently mapping the LLM calls onto its own optimized backends or external accelerators. The result is one stack that unifies  **how agents are defined**  (LangGraph graphs) with  **how they run efficiently**  (PACE's platform-aware runtime). Developers write agents in familiar LangGraph, and PACE handles orchestration and scheduling - extending the platform-awareness that made PACE fast for inference into multi-step, multi-agent workflows. Around this orchestration core, PACE focuses on four features.

1\. **Deterministic Agent Execution and Benchmarking**  
Agentic systems are hard to debug and reproduce. Because they chain up many LLM calls, tool responses, and branches - each subject to sampling randomness; non-determinism can produce very different runs. This  **non-determinism leads to inconsistent results**, making failures hard to reproduce, and evaluations hard to trust. PACE addresses this with a first-class **replay mode**: during a run, it records the full trace - graph state, prompts, sampling parameters, tool inputs/outputs, and event ordering, then re-executes the graph against that trace to reproduce the run step-for-step. This makes it easy to reproduce latency, defects and build reliable regression tests by turning agent benchmarking reliable and a repeatable, engineering-grade workflow. Additionally, PACE supports a **profiling mode** that provides detailed insight into the time spent across different operations in the agentic flow.

2\. **Flexible Deployment Across Local and Remote AI Infrastructure**  
PACE can **orchestrate agents on CPU while offloading heavy inference to GPUs**: PACE can connect to  **AMD Radeon™ graphics or AMD Instinct™ GPUs running vLLM or other inference serving engines with Open AI compatibility**, dispatching model calls to them over the same interface. One can achieve the best of both worlds - PACE's efficient CPU orchestration and deterministic control plane, paired with high-throughput GPU inference, all behind one consistent API.

3\. **Native Support for Custom LangGraph**  
PACE enables bringing your own LangGraph workflows and using our example graphs and reference implementations to build and accelerate agentic applications.

4\. **Optimized Tool Execution for Agentic Benchmarks**  
To make agent quality and performance measurable, the current version of agentic workflow supports two leading agentic benchmarks: **WebVoyager** and  **GAIA**.

**WebVoyager**  tests  **web-browsing agents** on realistic, end-to-end tasks on live websites - searching, navigating, filling forms, and extracting information from dynamic pages, often using a multimodal (screenshot + DOM) view of the browser. It stresses the skills that matter for real-world autonomy: long-horizon planning, robust tool use, and recovery from unexpected page states. Figure 1 shows the normalized time spent on non-LLM operations across page weights (cost of observing a page) and the overhead increases with page weights.

![Figure 1. LLM vs non-LLM time](https://www.amd.com/content/dam/amd/en/images/blogs/designs/technical-blogs/agentic-flow-with-amd-pace/figure1.png)

Figure 1. LLM vs non-LLM time

**GAIA (General AI Assistants)**  tests general assistant reasoning through questions that are simple for humans but require agents to combine multi-step reasoning, tool use, web search, file handling, and multimodal understanding.

Finally, PACE treats an agent as a single end-to-end flow to optimize. Because it owns the entire LangGraph execution - orchestration, scheduling, inference, and tool dispatch; PACE can optimize across the full graph.

PACE currently enables orchestration across WebVoyager and GAIA tasks and enables configuration to choose

- Models and devices or Open AI API
- Cores per CPU
- Parallel workers

## Evaluation

| **Feature** | **Specification** |
| --- | --- |
| **CPU** | AMD EPYC™ 9755 Series server CPUs, (“Turin”) |
| **Architecture** | Zen 5 |
| **Cores** | 128 cores per socket |
| **RAM** | 1.5 TB |
| **Precision** | BF16 |
| **Sockets** | 2 |

| **Feature** | **Specification** |
| --- | --- |
| **GPU** | AMD Radeon™ AI PRO R9700S |
| **Architecture** | AMD RDNA™ 4 |

| **Page weight** | **Normalized Page Weight** | **Operator Level Speedup** | **End to End Speedup** |
| --- | --- | --- | --- |
| **Very light** | 1x | 7.2x | 1.02x |
| **Light** | 2.6x | 8.8x | 1.04x |
| **Medium** | 5.9x | 10.0× | 1.08× |
| **Heavy** | 15.8x | 9.1x | 1.21x |
| **Very heavy** | 68.2x | 7.9x | 2.20x |

Agentic loops repeatedly invoke tools, observe their output, and call the LLM. As Table 2 shows, PACE significantly reduces the time spent observing page state, which translates into an end-to-end performance gain of 1.02×–2.20× on WebVoyager tasks. Here, *page weight* denotes the cost of observing a page — i.e., serializing its DOM/accessibility tree; normalized to the lightest task in the set (=1×); heavier pages have larger DOM trees and therefore higher observation cost. The gain grows with page weight because observation dominates a larger share of total task time on heavy pages. See the [AMD PACE repository](https://github.com/amd/AMD-PACE) for sample tasks and configurations; further optimizations are ongoing.

**Applicability of PACE:** PACE can further be applied to characterize agentic workflows, such as determining the optimal choice between an LLM and an SLM for a given benchmark.

**Work in Progress**: Subsequent releases are expected to add features such as semantic-routing-based model switching and support for recent Qwen models.

## Conclusion

In summary, PACE enables **agentic AI orchestrators**. By implementing LangGraph on top of its platform-aware runtime, PACE unifies how agents are defined with how they are executed efficiently on AMD hardware. Its four pillars - deterministic execution, flexible deployment, native custom LangGraph support and optimized tool execution are demonstrated with real performance gains; together provide PACE the required foundation for the next generation of autonomous and multi-agent AI systems. As the field continues its climb from modeling to reasoning agents to multi-agent collaboration, PACE aims to be the efficient, reliable engine underneath it all.

## Resources

- [AMD PACE GitHub repository](https://github.com/amd/AMD-PACE)
- [vLLM project](https://docs.vllm.ai/)
- [Original AMD PACE blog](https://www.amd.com/en/developer/resources/technical-articles/2026/amd-pace---high-performance-platform-aware-compute-engine.html)
- [5th Gen AMD EPYC processor architecture white paper](https://www.amd.com/content/dam/amd/en/documents/epyc-business-docs/white-papers/5th-gen-amd-epyc-processor-architecture-white-paper.pdf)
- [AMD EPYC 9005 Series processor data sheet](https://www.amd.com/content/dam/amd/en/documents/epyc-business-docs/datasheets/amd-epyc-9005-series-processor-datasheet.pdf)

Footnotes

## Footnote

A system configured with an AMD EPYC™ 9755 series processor and AMD Radeon™ AI PRO R9700S was used to evaluate AMD PACE end to end orchestration. Testing done by AMD on 21 <sup>st</sup> August. Results may vary based on configuration, usage, software version, and optimizations. SYSTEM CONFIGURATION: Supermicro; AMD EPYC 9755 128-Core Processor (2 sockets, 128 cores per socket, 2 threads per core); 1 NUMA node per socket; 1536 GB memory (24 DIMMs, 6400 MT/s, 64 GiB/DIMM); Ubuntu 24.04.2 LTS, kernel 6.8.0-86-generic.

---

![](https://www.amd.com/content/dam/amd/en/images/backgrounds/dividers/divider-white-pearl-gradient-medium.jpg "white pearl gradient medium color divider")
