---
格式版本: 2
标题: "Simplify AMD Instinct™ GPU Workloads on Oracle Kubernetes Engine with the AMD GPU Operator | cloud-infrastructure"
原文链接: "https://blogs.oracle.com/cloud-infrastructure/simplify-amd-gpu-workloads-on-oke"
发布日期: "2026-08-25"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:configured_publication_date_rule"
发布时间证据: "doc-fixed-49fb23095330-publication-date html:original: August 25, 2026"
发布时间校准原因: "信源发布日期识别规则直接确认发布时间"
发布时间校准置信度: "high"
发布时间候选数量: 1
发布时间严格候选数量: 1
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-26T16:00:09+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-26T15:59:29+08:00"
入库时间: "2026-08-26T08:00:10.591Z"
来源平台: "固定入口"
搜索渠道: "fixed_url"
搜索词: "https://blogs.oracle.com/?page=news"
匹配关键词:
  - "GPU"
  - "performance"
  - "bandwidth"
  - "AI"
相关厂家:
  - "Oracle"
  - "AMD"
相关专家:
  []
内容类型: "网页"
抓取工具: "CDP Render"
清洗工具: "CDP Text + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 50
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是Oracle官方介绍AMD GPU Operator现已作为OKE增强集群可选插件，新增事实为其相较原AMD Device Plugin扩展了驱动、发现、调度、监控、验证及GPU分区的统一生命周期管理。固定知识库未见该OKE集成事件，但Top 5并非完整历史，且文章未提供发布日期、客户部署、规模数据、独立实测或机架级架构变化。当前页面虽为Oracle一手来源且正文完整，技术内容主要属于Kubernetes运维与单厂商生态集成，命中“教程与运维选型”硬否决，不属于可复用的机架级AI基础设施新增。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-08-26T16:00:21+08:00"
AI主题相关性: 4
AI来源权威性: 15
AI新颖性: 9
AI技术细节: 9
AI商业部署信号: 3
AI完整性: 10
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819"
AI评分知识库SHA256: "e8daaadd1f28923bf4557a3e84042f83b8070615c03a40d448c106e6bfe5fcdc"
AI评分知识库检索词: "[\"Oracle\",\"AMD\",\"GPU\",\"https://blogs.oracle.com/?page=news\",\"PCIe\",\"RAS\",\"Intel\",\"OKE\",\"GPU-accelerated\",\"ML\",\"HPC\",\"GPUs\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0081\",\"title\":\"AAI 2026: 6th Gen AMD EPYC Server CPUs Power the Agentic Data Center\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"AMD\",\"GPU\",\"PCIe\",\"RAS\",\"Intel\",\"ML\",\"HPC\",\"GPUs\"],\"rank\":-16.229910662201736},{\"id\":\"july-correct-0078\",\"title\":\"AAI 2026: AMD Launches AMD Instinct MI400 Series GPUs for Frontier AI, HPC\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"AMD\",\"GPU\",\"RAS\",\"ML\",\"HPC\",\"GPUs\"],\"rank\":-14.172551390529172},{\"id\":\"july-correct-0034\",\"title\":\"AMD Fires Back at Nvidia with Helios AI System, Epyc CPUs\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"AMD\",\"GPU\",\"RAS\",\"Intel\",\"OKE\",\"HPC\",\"GPUs\"],\"rank\":-12.433674181399121},{\"id\":\"july-correct-0026\",\"title\":\"The Rackscale AI System Roadmaps That AMD Is Using To Chase Money\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"AMD\",\"GPU\",\"RAS\",\"Intel\",\"OKE\",\"HPC\",\"GPUs\"],\"rank\":-12.332751209932022},{\"id\":\"july-correct-0016\",\"title\":\"AMD launches Instinct MI400 Series GPUs for AI workloads\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"AMD\",\"GPU\",\"RAS\",\"HPC\",\"GPUs\"],\"rank\":-12.31592437253742}]"
AI摘要: "Oracle Cloud Infrastructure 的 Kubernetes Engine（OKE）现已将 AMD GPU Operator 作为增强集群的可选插件提供。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T02:08:50.082Z"
采集批次: "2026年8月26日15点57分35秒"
采集批次ID: "20260826-155735-186"
去重键: "https://blogs.oracle.com/cloud-infrastructure/simplify-amd-gpu-workloads-on-oke"
---

Oracle Cloud Infrastructure Kubernetes Engine (OKE) now offers the AMD GPU Operator as an optional cluster add-on for enhanced clusters. This add-on helps platform teams deploy and manage the AMD GPU software components required to run GPU-accelerated Kubernetes workloads through OKE.

For artificial intelligence (AI), machine learning (ML), inference, and high-performance computing (HPC) workloads, provisioning GPU worker nodes is only the beginning. Teams also need the software that makes GPUs available to containers, visible to Kubernetes, observable by operators, and ready for specialized configurations such as GPU partitioning.

The AMD GPU Operator brings these capabilities together in an OKE-managed add-on experience.

## Why GPU software management matters

A production-ready Kubernetes environment for workloads backed by AMD Instinct GPU powered infrastructure requires multiple components before GPU accelerated workloads can be deployed. Depending on the workload and environment, this can include AMD GPU drivers, node discovery and labeling, the [Kubernetes device plugins](https://kubernetes.io/docs/concepts/extend-kubernetes/compute-storage-net/device-plugins/) or [Dynamic Resource Allocation](https://kubernetes.io/docs/concepts/scheduling-eviction/dynamic-resource-allocation/), GPU metrics, health monitoring, validation tools, and device configuration.

These components must align with the Kubernetes version, worker node host OS image, GPU hardware, driver version, in use by your workloads. Installing and maintaining this large number of variables independently can increase operational complexity for platform teams and make updates more difficult to plan and validate.

The AMD GPU Operator helps reduce that complexity by managing key AMD GPU software components through a unified operator model.

## Managed cluster add-ons in OKE

OKE cluster add-ons extend Kubernetes clusters with optional capabilities that administrators can enable and configure. On enhanced clusters, administrators can:

- Enable or disable supported add-ons.
- Select a supported add-on version.
- Choose automatic updates or manage the deployed version.
- Apply supported configuration arguments.

With the AMD GPU Operator deployed as an OKE add-on, teams can manage its lifecycle through OKE rather than separately installing and operating the operator in each cluster.

## The AMD GPU Operator

OKE previously supported AMD Instinct GPUs through the [AMD GPU Plugin add-on](https://docs.oracle.com/en-us/iaas/releasenotes/conteng/conteng-AMD-GPU-device-plugin-addon.htm), which managed the AMD Device Plugin for Kubernetes. The [AMD GPU Operator add-on](https://docs.oracle.com/en-us/iaas/releasenotes/conteng/conteng-AMD-GPU-Operator-addon.htm) expands on that foundation with broader lifecycle management capabilities.

The operator coordinates components that support the complete lifecycle of AMD Instinct GPU workloads, including driver management, GPU discovery, scheduling, monitoring, validation, and configuration.

### Discover and prepare AMD Instinct GPU nodes

The operator uses Node Feature Discovery (NFD) to detect AMD Instinct GPU hardware and advertise node capabilities through Kubernetes labels. These labels help the operator identify the appropriate nodes and enable platform teams to target GPU workloads based on hardware characteristics. [Node Feature Discovery is another add-on](https://docs.oracle.com/en-us/iaas/Content/ContEng/Tasks/configuration-arguments-node-feature-discovery-plugin.htm) supported by OKE.

Kernel Module Management (KMM) manages the lifecycle of GPU driver kernel modules. Together with the Controller Manager, it supports driver installation, upgrades, and removal according to the desired configuration.

This coordinated approach helps ensure that worker nodes are prepared before GPU workloads are scheduled.

### Make GPUs schedulable for Kubernetes workloads

The AMD GPU Device Plugin integrates AMD Instinct GPUs with the Kubernetes device-plugin framework. It registers AMD Instinct GPUs as allocatable resources—such as amd.com/gpu—so application teams can request them in pod specifications.

Kubernetes can then schedule workloads to nodes with the required available GPU capacity. Platform teams can use node selectors, taints, and tolerations to help reserve GPU nodes for workloads that explicitly request accelerated resources.

The operator’s node labeler can also apply detailed GPU-specific labels, allowing more targeted workload placement when applications require specific GPU capabilities.

## Monitor GPU health and utilization

GPU operations require visibility beyond standard CPU and memory metrics. The Device Metrics Exporter provides GPU metrics in Prometheus format, including data that can help teams monitor GPU utilization, temperature, and health.

These metrics can support monitoring and alerting workflows while helping teams identify unhealthy devices. The operator can also use health information alongside the device plugin so unhealthy GPUs are not presented as schedulable capacity.

For larger clusters, teams should size the resources allocated to operator-managed components appropriately. The number of GPU nodes, GPUs per node, monitoring frequency, and workload intensity can all affect operational resource requirements.

## Validate GPU readiness

The AMD GPU Operator includes a test runner for hardware validation, diagnostics, and benchmarking. Teams can use it to run configurable tests on GPU worker nodes, schedule or manually trigger validation workflows, and report test outcomes as Kubernetes events.

The test runner can also run pre-start tests as init containers for GPU workload pods. This can be useful for long-running jobs where validating GPU health and stability before execution is important.

Depending on the selected test tooling, validation scenarios can include GPU stress tests, PCIe bandwidth benchmarks, memory tests, and burn-in tests.

## Configure GPU partitioning

For supported AMD Instinct GPU environments, the Device Config Manager (DCM) provides a Kubernetes-native way to manage GPU partitioning. DCM runs on GPU worker nodes and uses configuration profiles stored in Kubernetes ConfigMaps.

Platform teams can define partitioning profiles and apply them to nodes using labels. DCM monitors the selected profiles and node labels, then applies the appropriate configuration to the targeted GPUs.

This approach can help teams support different workload profiles across a GPU fleet while maintaining a consistent configuration process. For example, environments can use separate compute and memory partitioning profiles for workloads with different resource requirements.

## Getting started

The AMD GPU Operator is available as an optional OKE cluster add-on for enhanced clusters. Before enabling it, review the supported OKE Kubernetes versions, AMD GPU worker node images, GPU hardware, drivers, and add-on versions.

OCI Console showing the OKE Add-ons configuration page.

Then select the add-on version and supported configuration arguments that fit your environment. Consider node labels, taints and tolerations, monitoring integration, driver-management requirements, and whether your workloads require validation or GPU partitioning capabilities.

The edit page for the AMD GPU Operator

With the AMD GPU Operator on OKE, platform teams can simplify the work required to prepare and maintain Kubernetes environments for AMD GPU workloads. This lets AI, ML, and HPC teams focus more on building and running accelerated applications.
