---
格式版本: 2
标题: "NVIDIA NIM Microservices for AI Inference"
原文链接: "https://www.nvidia.com/en-us/ai-data-science/products/nim-microservices/"
发布日期: "2026-07-10"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "llm:scrape:original_meta_loose"
发布时间证据: "nv-pub-date: 2026-07-10T06:14:07.000Z"
发布时间校准原因: "nv-pub-date 表示页面发布日期，符合优先级规则，且无其他更优先候选。"
发布时间校准置信度: "1"
发布时间候选数量: 6
发布时间严格候选数量: 0
发布时间原页读取状态: "原页面已读取"
发布时间未找到原因: ""
发布时间校准时间: "2026-07-27T00:11:48+08:00"
发布时间仲裁状态: "confirmed"
发布时间仲裁尝试次数: 2
发布时间仲裁耗时毫秒: 16511
发现时间: "2026-07-26T23:55:44+08:00"
入库时间: "2026-07-26T16:12:05.824Z"
来源平台: "NVIDIA 站内搜索"
搜索渠道: "source_template"
搜索词: "https://www.nvidia.com/en-us/search/?q=Foxconn&page=1"
匹配关键词:
  - "deployment"
  - "performance"
  - "latency"
  - "throughput"
相关厂家:
  - "Foxconn"
  - "NVIDIA"
  - "OpenAI"
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 33
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "主题完全不相关：正文为NVIDIA NIM推理微服务软件介绍，未涉及超节点、AI Rack、机柜级硬件架构、供电、液冷、高速互连等核心主题，仅软件层面内容。"
AI质检模型: "deepseek-v4-flash"
AI质检时间: "2026-07-27T11:07:39+08:00"
AI主题相关性: 2
AI来源权威性: 15
AI新颖性: 3
AI技术细节: 3
AI商业部署信号: 2
AI完整性: 8
采集批次: "2026年7月26日2点27分21秒"
采集批次ID: "20260726-022721-515"
去重键: "https://www.nvidia.com/en-us/ai-data-science/products/nim-microservices"
---

![NVIDIA AI](data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mNkYAAAAAYAAjCB0C8AAAAASUVORK5CYII= "NVIDIA AI")

## NVIDIA NIM Microservices

Designed for rapid, reliable deployment of accelerated generative AI inference anywhere.

## Overview

## What Is NVIDIA NIM?

NVIDIA NIM™ provides prebuilt, optimized inference microservices for rapidly deploying the latest AI models on any NVIDIA-accelerated infrastructure—cloud, data center, workstation, and edge.

### Sovereign AI Agents Think Local, Act Global With NVIDIA AI Factories

Validated design for AI factories pairs accelerated infrastructure with software, including new NVIDIA NIM™ capabilities and an expanded suite of NVIDIA blueprints.

### Free Development Access to NIM

Get access to unlimited prototyping with hosted APIs for NIM accelerated by DGX Cloud, or download and self-host NIM microservices for research and development as part of the NVIDIA Developer program.

## Accelerate AI Deployment With NVIDIA NIM

NVIDIA NIM combines the ease of use and operational simplicity of managed APIs with the flexibility and security of self-hosting models on your preferred infrastructure. NIM microservices come with everything AI teams need—the latest AI foundation models, optimized inference engines, industry-standard APIs, and runtime dependencies—prepackaged in enterprise-grade software containers ready to deploy and scale anywhere.

![NVIDIA NIM Stack Diagram](data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mNkYAAAAAYAAjCB0C8AAAAASUVORK5CYII= "NVIDIA NIM Stack Diagram")

### Benefits

## Enterprise Generative AI That Does More for Less

Easy, production-ready microservices are built for high-performance AI and designed to work seamlessly and scale affordably. Get started building AI agents and other enterprise generative AI applications faster with the latest AI models for reasoning, simulation, speech, and more.

### Demo

## Build AI Agents With NIM

[![Video thumbnail showing someone using an AI agent](https://www.nvidia.com/content/nvidiaGDC/us/en_US/ai-data-science/products/nim-microservices/_jcr_content/root/responsivegrid/nv_container_1676348_2037606428/nv_image.coreimg.jpeg/1773699326401/build-simple-ai-agent-ari.jpeg "Video thumbnail showing someone using an AI agent")](#demo-video)

Learn how to set up two AI agents—one for content generation and another for digital graphic design—and see how easy it is to get up and running with NIM microservices.

---

### Technology

## Building Blocks for Agentic AI

![Reasoning NIM icon](https://www.nvidia.com/content/nvidiaGDC/us/en_US/ai-data-science/products/nim-microservices/_jcr_content/root/responsivegrid/nv_container_1676348/nv_container/nv_teaser.coreimg.svg/1773699326702/m48-thinking-reasoning-ffffff.svg "Reasoning NIM icon")

### Get the Latest AI Models

Access the latest AI models for reasoning, language, retrieval, speech, vision and more—ready to deploy in five minutes on any NVIDIA-accelerated infrastructure.

### Benchmarks

## Boost Throughput With NIM

NVIDIA NIM provides optimized throughput and latency out of the box to maximize token generation, support concurrent users at peak times, and improve responsiveness. NIM microservices are continuously updated with the latest optimized inference engines, boosting performance on the same infrastructure over time.

Configuration: Llama 3.1 8B instruct, 1x H100 SXM; concurrent requests: 200. NIM ON: FP8, throughput 1201 tokens/s, ITL 32ms. NIM OFF: FP8, throughput 613 tokens/sec, ITL 37ms.

### Models

## Unlock Enterprise-Ready Inference for Thousands of Open Models

Deploy large language models (LLMs) supported by NVIDIA® TensorRT™-LLM, vLLM, or SGLang for low-latency, high-throughput inferencing on NVIDIA-accelerated infrastructure.

<iframe title="" width="100%" height="800px" frameborder="0" allow="microphone; fullscreen; clipboard-read; clipboard-write;" src="https://build.nvidia.com/carousel?nv_source=www-ai"></iframe>

---

### Features

## The Easy Button for AI Development and Deployment

Designed to run anywhere, NIM microservices expose industry-standard APIs for easy integration with enterprise systems and applications and scale seamlessly on Kubernetes to deliver high-throughput, low-latency inference at cloud scale.

### Deploy NIM

Deploy NIM for your model with a single command. You can also easily run NIM with LLMs supported by NVIDIA TensorRT-LLM, vLLM, or SGLang, including fine-tuned models.

### Run Inference

Get NIM up and running with the optimal runtime engine based on your NVIDIA-accelerated infrastructure.

### Build

Integrate self-hosted NIM endpoints with just a few lines of code.

Deploy

Run

Build

docker run nvcr.io/nim/publisher\_name/model\_name

curl -X 'POST' \\ 'http://0.0.0.0:8000/v1/completions' \\ -H 'accept: application/json' \\ -H 'Content-Type: application/json' \\ -d '{
```
"model" : "model_name",
```
```
"prompt" : "Once upon a time",
```
```
"max_tokens" : 64
```
}'

import openai client = openai.OpenAI(
```
base_url = "YOUR_LOCAL_ENDPOINT_URL",
```
```
api_key="YOUR_LOCAL_API_KEY"
```
) chat\_completion = client.chat.completions.create(
```
model="model_name",
```
messages=\[{"role": "user", "content": "Write me a love song" }\], temperature=0.7 )

### Use Cases

## How NIM Is Being Used

See how NVIDIA NIM supports industry use cases, and jump-start your AI development with curated examples.

### Starting Options

## Ways to Get Started With NVIDIA NIM

### Start Prototyping for Free

Get started with easy-to-use API endpoints for NIM, powered by DGX Cloud.

- Access fully accelerated AI infrastructure.
- Ensure your data isn't used for model training.
- Access for development and testing as part of the [NVIDIA Developer Program](https://developer.nvidia.com/developer-program).

### Download and Deploy

Run NVIDIA NIM to scale optimized AI models in the cloud or data center of your choice.

- Ensure data never leaves your secure enclave.
- Seamlessly transition from cloud endpoints to self-hosted APIs without code changes.
- Start with free access for development and testing, and move to an NVIDIA AI Enterprise license for production.

### Get in Touch

Talk to an NVIDIA AI specialist about moving generative AI pilots to production with the security, API stability, and support that comes with NVIDIA AI Enterprise.

- Explore your generative AI use cases.
- Discuss your technical requirements.
- Align NVIDIA AI solutions to your goals and requirements.
