---
格式版本: 2
标题: "NVIDIA Triton Accelerates Inference on Oracle Cloud | NVIDIA Blogs"
原文链接: "https://blogs.nvidia.com/blog/ai-inference-oci-triton/"
发布日期: "2026-06-30"
发布时间校准状态: "found"
发布时间来源: "llm:strict_original_body"
发布时间证据: "time class=related-news-date nvidia-article-date datetime=2026-06-30T08:00:57-07:00: Jun 30, 2026"
发布时间校准原因: "该日期来自标题下方的文章发布日期标签，符合博客文章发布时间的特征。"
发布时间校准置信度: "1"
发布时间候选数量: 36
发布时间严格候选数量: 12
发布时间原页读取状态: "原页面来自已抓取 HTML"
发布时间未找到原因: "候选日期无效或 LLM 未确认"
发布时间校准时间: "2026-07-20T13:24:23+08:00"
发现时间: "2026-07-20T09:26:32+08:00"
入库时间: "2026-07-20T05:27:14.421Z"
来源平台: "NVIDIA Blog 搜索"
搜索渠道: "source_template"
搜索词: "https://blogs.nvidia.com/?s=Oracle"
匹配关键词:
  []
相关厂家:
  - "Oracle"
  - "NVIDIA"
相关专家:
  []
内容类型: "网页"
抓取工具: "AgentKey Scrape"
清洗工具: "AgentKey Markdown + LLM 正文裁剪"
原始附件:
  []
AI优质: "否"
AI打分: 42
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "文章主要讨论NVIDIA Triton推理服务器在Oracle云上的软件部署与性能优化，未涉及超节点、AI Rack、机柜级系统、供电、散热、高速互连等硬件架构或基础设施细节，主题相关性低。"
AI质检模型: "qwen3.6-plus"
AI质检时间: "2026-07-20T13:27:14+08:00"
AI主题相关性: 5
AI来源权威性: 12
AI新颖性: 8
AI技术细节: 5
AI商业部署信号: 7
AI完整性: 5
图片摘要:
  - "✗ ./assets/img-83d2f103.jpg | diagram | Alt文本指向能效文章，与正文Triton/OCI主题无关"
  - "✗ ./assets/img-1f6575f1.png | diagram | Alt文本指向Vera CPU文章，正文未提及"
  - "✗ ./assets/img-3511966c.jpg | photo | NVIDIA品牌Logo，属于品牌宣传"
  - "✗ ./assets/img-e47e99dc.png | diagram | Alt文本指向Token Cost文章，正文主要讲CV推理"
采集批次: "2026年7月20日9点23分34秒"
采集批次ID: "20260720-092334-062"
去重键: "https://blogs.nvidia.com/blog/ai-inference-oci-triton"
---

An avid cyclist, Thomas Park knows the value of having lots of gears to maintain a smooth, fast ride.

So, when the software architect designed an AI inference platform to serve predictions for Oracle Cloud Infrastructure’s (OCI) Vision AI service, he picked [NVIDIA Triton Inference Server](https://www.nvidia.com/en-us/ai-data-science/products/triton-inference-server/). That’s because it can shift up, down or sideways to handle virtually any AI model, framework and hardware and operating mode — quickly and efficiently.

“The NVIDIA AI inference platform gives our worldwide cloud services customers tremendous flexibility in how they build and run their AI applications,” said Park, a Zurich-based computer engineer and competitive cycler who’s worked for four of the world’s largest cloud services providers.

Specifically, Triton reduced OCI’s total cost of ownership by 10%, increased prediction throughput up to 76% and reduced inference latency up to 51% for OCI Vision and Document Understanding Service models that were migrated to Triton. The services run globally across more than 45 regional data centers, according to [an Oracle blog](https://blogs.oracle.com/ai-and-datascience/post/oci-ai-vision-nvidia-triton-inference-server) Park and a colleague posted earlier this year.

## Computer Vision Accelerates Insights

Customers rely on OCI Vision AI for a wide variety of object detection and image classification jobs. For instance, a U.S.-based transit agency uses it to automatically detect the number of vehicle axles passing by to calculate and bill bridge tolls, sparing busy truckers wait time at toll booths.

OCI AI is also available in Oracle NetSuite, a set of business applications used by more than 37,000 organizations worldwide. It’s used, for example, to automate invoice recognition.

Thanks to Park’s work, Triton is now being adopted across other OCI services, too.

## A Triton-Aware Data Service

“Our AI platform is Triton-aware for the benefit of our customers,” said Tzvi Keisar, a director of product management for OCI’s Data Science service, which handles machine learning for Oracle’s internal and external users.

“If customers want to use Triton, they don’t have to worry about the configuration because it will be done automatically by the service, launching a Triton-powered inference endpoint for them,” said Keisar.

Triton is included in [NVIDIA AI Enterprise](https://www.nvidia.com/en-us/data-center/products/ai-enterprise/), a platform that provides full security and support businesses need — and it’s available on [OCI Marketplace](https://cloudmarketplace.oracle.com/marketplace/en_US/listing/155314141).

## A Massive SaaS Platform

OCI’s Data Science service is the machine learning platform for both Oracle NetSuite and Oracle Fusion Applications.

“These business application suites are massive, with tens of thousands of customers who are also building their frameworks on top of our service,” he said.

It’s a wide swath of mainly enterprise users in manufacturing, retail, transportation and other industries. They’re building and using AI models of nearly every shape and size.

Inference was one of the group’s first services, and Triton came on the team’s radar not long after its launch.

## A Best-in-Class Inference Framework

“We saw Triton pick up in popularity as a best-in-class serving framework, so we started experimenting with it,” Keisar said. “We saw really good performance, and it closed a gap in our existing offerings, especially on multi-model inference — it’s the most versatile and advanced inferencing framework out there.”

Launched on OCI [in March](https://blogs.oracle.com/ai-and-datascience/post/oci-nvidia-triton-inference-server), Triton has already attracted the attention of many internal teams at Oracle hoping to use it for inference jobs that require serving predictions from multiple AI models running concurrently.

“Triton has a very good track record and performance on multiple models deployed on a single endpoint,” he said.

## Accelerating the Future

Looking ahead, Keisar’s team is evaluating [NVIDIA TensorRT-LLM](https://developer.nvidia.com/tensorrt#inference) software to supercharge inference on the complex large language models ([LLMs](https://www.nvidia.com/en-us/glossary/data-science/large-language-models/)) that have captured the imagination of many users.

An active blogger, Keisar’s [latest article](https://blogs.oracle.com/ai-and-datascience/post/quantize-deploy-llama2-70b-costeffective-a10s-oci) detailed quantization techniques for running a Llama 2 LLM with a whopping 70 billion parameters on NVIDIA A10 Tensor Core GPUs.

“Even down to four-bit parameters, the quality of model outputs is still quite good,” he said. “Deploying on NVIDIA GPUs gives us the flexibility to find a good balance in latency, throughput and cost.”

After announcements this fall that Oracle is deploying the latest [NVIDIA H100 Tensor Core GPUs](https://www.nvidia.com/en-us/data-center/h100/), H200 GPUs, L40S GPUs and [Grace Hopper Superchips](https://www.nvidia.com/en-us/data-center/grace-hopper-superchip/), it’s just the start of many accelerated efforts to come.
