---
格式版本: 2
标题: "NVIDIA and AWS Collaborate to Bring AI to Production at Scale | NVIDIA Blog"
原文链接: "https://blogs.nvidia.com/blog/nvidia-aws-ai-production-scale/"
发布日期: "2026-06-30"
发布时间校准状态: "found"
发布时间来源: "llm:strict_original_body"
发布时间证据: "time class=related-news-date nvidia-article-date datetime=2026-06-30T08:00:57-07:00: Jun 30, 2026"
发布时间校准原因: "该日期来自标题下方的文章发布日期字段，且为最早日期，符合文章发布时间特征。"
发布时间校准置信度: "100"
发布时间候选数量: 36
发布时间严格候选数量: 12
发布时间原页读取状态: "原页面来自已抓取 HTML"
发布时间未找到原因: "候选日期无效或 LLM 未确认"
发布时间校准时间: "2026-07-20T12:26:27+08:00"
发现时间: "2026-07-20T09:26:26+08:00"
入库时间: "2026-07-20T04:38:20.862Z"
来源平台: "NVIDIA Blog 搜索"
搜索渠道: "source_template"
搜索词: "https://blogs.nvidia.com/?s=AWS"
匹配关键词:
  - "GPU"
相关厂家:
  - "AWS"
  - "NVIDIA"
相关专家:
  []
内容类型: "网页"
抓取工具: "AgentKey Scrape"
清洗工具: "AgentKey Markdown + LLM 正文裁剪"
原始附件:
  []
AI优质: "否"
AI打分: 68
AI分档: "召回候选"
AI质检状态: "不通过"
AI打分理由: "NVIDIA官方发布，涉及GB300 Exemplar Cloud认证及EC2 G7实例规格，具商业部署信号。但正文聚焦云实例与软件库，未深入超节点/机柜级架构、供电散热互连等核心硬件细节，技术深度不足。"
AI质检模型: "qwen3.6-plus"
AI质检时间: "2026-07-20T12:38:20+08:00"
AI主题相关性: 12
AI来源权威性: 14
AI新颖性: 16
AI技术细节: 8
AI商业部署信号: 10
AI完整性: 8
图片摘要:
  - "✗ ./assets/img-83d2f103.jpg | other | 图片Alt文本指向能效主题，且为通用服务器机架渲染图，与正文主要讲述的EC2 G7实例和OpenSearch合作内容无直接关联，疑似推荐阅读配图。"
  - "✗ ./assets/img-1f6575f1.png | diagram | 图片展示Vera CPU架构，正文未提及该产品，属于无关技术配图。"
  - "✗ ./assets/img-3511966c.jpg | photo | NVIDIA公司Logo及建筑实拍，属于品牌宣传/装饰性配图，无具体技术信息。"
  - "✗ ./assets/img-e47e99dc.png | other | 抽象渲染图，Alt文本指向推理软件栈，但图片本身无具体架构或数据，且正文主要侧重硬件实例与向量搜索库，关联度弱。"
采集批次: "2026年7月20日9点23分34秒"
采集批次ID: "20260720-092334-062"
去重键: "https://blogs.nvidia.com/blog/nvidia-aws-ai-production-scale"
---

Building AI systems at scale is demanding, requiring low-latency inference, fast vector search, strong GPU price-performance and infrastructure that can grow without multiplying operational complexity.

NVIDIA’s latest work with Amazon Web Services (AWS) addresses each of those constraints. Across Amazon OpenSearch and Amazon EC2, NVIDIA AI infrastructure is giving enterprises more practical paths to deploy AI at production scale.

EC2 G7 instances powered by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs expand the compute layer for AI, graphics, video and data analytics workloads, while the NVIDIA cuVS library accelerates the retrieval layer by making GPU-powered vector indexing the default in OpenSearch Serverless. And with AWS achieving NVIDIA Exemplar Cloud status for NVIDIA GB300, customers can trust they’re receiving peak optimized performance for their training workloads.

## NVIDIA RTX PRO 4500 Blackwell Server Edition Multi-Workload GPUs Power New Amazon EC2 G7 Instances

Amazon EC2 G7 instances bring NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs to AWS for AI inference, graphics, spatial computing and GPU-accelerated data analytics — delivering a new instance type engineered for production workloads that need performance without the operational overhead of a customer-managed GPU platform.

Compared with G6 instances, G7 delivers up to 4.6x AI inference performance, up to 2.1x graphics performance and significantly faster GPU-accelerated data analytics on Amazon EMR using the NVIDIA cuDF library for Apache Spark workloads.

With support for up to eight GPUs, 256GB of total GPU memory, 700 Gbps of EFA-enabled networking and up to 7.6TB of local NVMe SSD storage — across one-, two-, four- and eight- GPU configurations plus bare metal, coming soon — G7 instances let customers right-size infrastructure for their workloads instead of over-provisioning for them.

The platform’s versatility means AI teams get lower-latency inference. Media and entertainment teams get high-resolution video workflows and rendering. Simulation, computer-aided design, virtual desktop infrastructure, gaming and spatial computing teams get the same instance type for graphics-intensive applications. And data teams can apply the GPU memory, local storage and networking improvements to analytics pipelines and vector database workloads.

G7 instances are accessible through AWS Deep Learning Amazon Machine Images (AMIs), Amazon Deep Learning Containers, Amazon EMR, Amazon EKS, Amazon ECS and graphics AMIs — and coming soon to Amazon SageMaker AI.

## NVIDIA cuVS Makes GPU-Accelerated Vector Search the Default in Amazon OpenSearch

The next generation of Amazon OpenSearch Serverless powers agentic AI and dynamic workloads with no infrastructure management required. It uses GPU-accelerated vector indexing, powered by NVIDIA cuVS, as the default compute choice for all vector collections.

For teams building [retrieval-augmented generation](https://blogs.nvidia.com/blog/what-is-retrieval-augmented-generation/), semantic search, recommendation systems and agentic AI applications, that shift matters. It turns GPU-powered vector search from a specialized optimization project into a standard AWS capability.

The customer impact is direct: vector indexing up to 10x faster at a quarter of the cost, compared with CPU-only builds — making billion-scale vector databases practical to build in under an hour.

By making NVIDIA cuVS the default in OpenSearch Serverless, AWS customers get a much faster path from raw data to production-ready AI retrieval infrastructure — with serverless scaling that reduces operational overhead when workloads are idle.

## AWS Achieves NVIDIA Exemplar Cloud Status for GB300 Training Performance

AWS has achieved NVIDIA Exemplar Cloud status on NVIDIA GB300 for training workloads. This means AWS meets the rigorous performance thresholds that NVIDIA uses to benchmark AI workloads against its reference architecture.

This achievement is the result of deep co-engineering efforts between AWS and NVIDIA teams. Through the NVIDIA Exemplar Clouds initiative, developers and AI leaders can be confident they’re using consistent, high-performance cloud infrastructure for large-scale training, helping teams evaluate cloud providers with greater confidence, improve total cost of ownership and move AI projects from planning to production more efficiently.

Together, these advancements reinforce every layer of the AI infrastructure stack on AWS. The throughline is the same: production-grade AI infrastructure that performs at scale, without adding operational burden to the teams running it.

*Learn more in* [*this AWS blog*](https://aws.amazon.com/about-aws/whats-new/2026/06/amazon-ec2-g7-generally-available/)*.*
