---
格式版本: 2
标题: "Generative AI Inference Recommendation for Amazon SageMaker now available in the SageMaker AI Studio"
原文链接: "https://aws.amazon.com/cn/about-aws/whats-new/2026/08/generative-ai-inference-recommendation-for-amazon-sagemaker-now-available-in-the-sagemaker-ai-studio/"
发布日期: "2026-08-20"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "llm:local:original_script_field"
发布时间证据: "postDateTime: 2026-08-20T17:25:00Z"
发布时间校准原因: "候选日期来自页面脚本字段postDateTime，通常表示文章正式发布时间，且无其他更优先来源。"
发布时间校准置信度: "1"
发布时间候选数量: 1
发布时间严格候选数量: 0
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-21T09:55:38+08:00"
发布时间仲裁状态: "confirmed"
发布时间仲裁尝试次数: 1
发布时间仲裁耗时毫秒: 6953
发现时间: "2026-08-21T09:50:45+08:00"
入库时间: "2026-08-21T01:55:47.474Z"
来源平台: "固定入口"
搜索渠道: "fixed_url"
搜索词: "https://aws.amazon.com/new"
匹配关键词:
  - "AI"
  - "performance"
  - "latency"
  - "throughput"
  - "GPU"
相关厂家:
  - "AWS"
  - "NVIDIA"
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 37
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "该内容为AWS SageMaker推理配置推荐功能更新，属软件层服务优化，未涉及超节点、AI Rack、机柜级AI基础设施、供电散热互连或量产落地等主题，与项目关注范围无关。"
AI质检模型: "ali-deepseek-v4-flash"
AI质检时间: "2026-08-21T09:56:30+08:00"
AI主题相关性: 2
AI来源权威性: 15
AI新颖性: 5
AI技术细节: 5
AI商业部署信号: 0
AI完整性: 10
AI摘要: "AWS在SageMaker AI Studio推出生成式AI推理推荐功能，帮助用户以低代码/无代码方式自动在真实GPU上基准测试多种配置，按延迟、吞吐或成本目标返回排序后的生产就绪推荐，将验证周期从数周缩至数小时。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T02:13:40.327Z"
采集批次: "2026年8月21日9点50分40秒"
采集批次ID: "20260821-095040-035"
去重键: "https://aws.amazon.com/cn/about-aws/whats-new/2026/08/generative-ai-inference-recommendation-for-amazon-sagemaker-now-available-in-the-sagemaker-ai-studio"
---

Amazon SageMaker AI now offers Generative AI Inference Recommendations in SageMaker AI Studio, giving customers a guided, low-code, no-code path to find the best inference configuration for their workload. This builds on the API-based launch in April 2026, extending the same benchmarking infrastructure to teams that prefer a visual workflow over programmatic access.

Deploying generative AI models in production requires finding the right combination of instance type, serving container, and optimization strategy. Getting this right typically involves weeks of manual benchmarking, configuration tuning, and trial-and-error, with no easy way to know if the final setup is actually optimal. With the new experience, customers describe their workload and what matters most, whether that's latency, throughput, or cost, and SageMaker AI does the rest. It benchmarks multiple configurations on real GPU infrastructure using NVIDIA AIPerf, applies goal-aligned techniques like speculative decoding for throughput or kernel tuning for latency, and returns ranked, production-ready recommendations with measured performance data. Teams get to a validated configuration in hours instead of weeks, without needing to decide which techniques to apply or how to configure them.

With the new experience, customers describe their workload and what matters most, whether that's latency, throughput, or cost, and SageMaker AI does the rest. It benchmarks multiple configurations on real GPU infrastructure using NVIDIA AIPerf, applies goal-aligned techniques like speculative decoding for throughput or kernel tuning for latency, and returns ranked, production-ready recommendations with measured performance data. Teams get to a validated configuration in hours instead of weeks, without needing to decide which techniques to apply or how to configure them.

In SageMaker AI Studio under Jobs, Inference optimization, customers select a use-case profile (Interact, Generate, Summarize, or Custom), choose an optimization goal (minimize latency, maximize throughput, or minimize cost), and pick their model from JumpStart, S3, Model Registry, or an existing SageMaker model. Recommendations are ranked by TTFT, inter-token latency, throughput, and cost, and can be compared visually before deploying to a SageMaker real-time endpoint directly from Studio.

There is no additional cost for generating recommendations. Standard compute costs apply for optimization jobs and endpoints provisioned during benchmarking. This capability is available in US East (N. Virginia), US West (Oregon), US East (Ohio), Europe (Ireland), Europe (Frankfurt), Asia Pacific (Singapore), Asia Pacific (Tokyo). To learn more, visit the [blog post](https://aws.amazon.com/blogs/machine-learning/launching-ui-for-generative-ai-inference-recommendations-in-amazon-sagemaker-ai/) or the [documentation](https://docs.aws.amazon.com/sagemaker/latest/dg/generative-ai-inference-recommendations.html).
