---
格式版本: 2
标题: "Amazon SageMaker HyperPod enhances support for Ray"
原文链接: "https://aws.amazon.com/cn/about-aws/whats-new/2026/08/amazon-sagemaker-hyperpod-ray/"
发布日期: "2026-08-24"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "llm:local:original_script_field"
发布时间证据: "postDateTime: 2026-08-24T15:00:00Z"
发布时间校准原因: "该候选来自文章原始脚本字段，明确标注为发布时间，符合发布时间判定优先级。"
发布时间校准置信度: "1"
发布时间候选数量: 1
发布时间严格候选数量: 0
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-25T17:28:04+08:00"
发布时间仲裁状态: "confirmed"
发布时间仲裁尝试次数: 1
发布时间仲裁耗时毫秒: 2939
发现时间: "2026-08-25T17:27:42+08:00"
入库时间: "2026-08-25T09:28:07.999Z"
来源平台: "固定入口"
搜索渠道: "fixed_url"
搜索词: "https://aws.amazon.com/new"
匹配关键词:
  - "GPU"
  - "AI"
  - "throughput"
相关厂家:
  - "AWS"
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 36
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "内容为AWS SageMaker HyperPod对Ray的软件功能增强，聚焦分布式训练、推理与可观测性，未涉及超节点、AI Rack、机柜级硬件架构、供电散热或高速互连等核心主题，与项目关注范围无关。"
AI质检模型: "ali-deepseek-v4-flash"
AI质检时间: "2026-08-25T17:30:14+08:00"
AI主题相关性: 3
AI来源权威性: 12
AI新颖性: 5
AI技术细节: 5
AI商业部署信号: 3
AI完整性: 8
AI摘要: "Amazon SageMaker HyperPod 增强了对 Ray 的支持，提供内置可观测性、弹性训练、加速推理和托管开发环境。数据科学家可通过 SageMaker Studio 网页界面管理 Ray 集群，并交互式开发；"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T02:09:49.729Z"
采集批次: "2026年8月25日17点27分36秒"
采集批次ID: "20260825-172736-501"
去重键: "https://aws.amazon.com/cn/about-aws/whats-new/2026/08/amazon-sagemaker-hyperpod-ray"
---

Amazon SageMaker HyperPod now enhances support for Ray with built-in observability, resilient training, accelerated inference and managed development environments. Ray is a popular open-source framework for scaling AI workloads on a unified compute layer, from data processing and distributed training to reinforcement learning and model serving. Running Ray on Kubernetes at production scale can be an operational burden: job hangs, low GPU utilization from static team allocations, and multi-step observability setup. Also, lack of interactive development environment means every code change needs another job submission and familiarity with kubectl.

HyperPod now brings easier development, resilient training, and accelerated inference to Ray. Data scientists create, edit, monitor, and delete Ray clusters from a web-based interface in Amazon SageMaker Studio, then attach JupyterLab, Code Editor, or a local IDE to a running Ray cluster and iterate interactively against cluster-scale compute. A multi-node Ray cluster behaves like a local development environment, so you test each change immediately, without waiting for a new job to queue and start. For Observability, HyperPod provisions Grafana dashboards with metrics in Amazon Managed Service for Prometheus and allows one-click access to the Ray Dashboard through a secure browser link, giving you visibility into your workloads from the first run. For training at scale, HyperPod node auto recovery and hung job detection handle GPU faults, job hangs, loss spikes, and degraded throughput. Tiered checkpointing restores state from cluster memory to maximize goodput, and task governance improves compute utilization through quotas, priorities, and preemption. Together, these keep your long training runs progressing through failures and maximize the useful work done per GPU-hour. For inference with Ray Serve, a tiered KV cache reuses cached prefixes to reduce time to first token, and you can deploy Amazon SageMaker JumpStart models directly.

Open-source Ray code runs unchanged and you can either adopt the purpose-built experience in SageMaker Studio or take individual capabilities to integrate into your own ML platform.

Ray support is available for HyperPod clusters orchestrated by Amazon EKS, in AWS Regions where SageMaker HyperPod is supported. To learn more, see the [SageMaker HyperPod documentation](https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-ray.html), and explore the [interactive demo](https://d1dpyy0tl92esj.cloudfront.net/overview).
