---
格式版本: 2
标题: "Not All LLM Workloads Are Equal: Benchmarking TPU Performance on Classification vs. Generation"
原文链接: "https://cloud.google.com/blog/topics/developers-practitioners/not-all-llm-workloads-are-equal-benchmarking-tpu-performance-on-classification-vs-generation"
发布日期: "2026-09-04"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:configured_publication_date_rule"
发布时间证据: "google-cloud-blog-date-div html:original: September 4, 2026"
发布时间校准原因: "信源发布日期识别规则直接确认发布时间"
发布时间校准置信度: "high"
发布时间候选数量: 1
发布时间严格候选数量: 1
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-07T22:09:39+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-07T22:06:41+08:00"
入库时间: "2026-09-07T14:25:53.057Z"
来源平台: "固定入口"
搜索渠道: "fixed_url"
搜索词: "https://cloud.google.com/blog/"
匹配关键词:
  - "performance"
  - "deployment"
  - "latency"
  - "bandwidth"
  - "throughput"
  - "AI"
相关厂家:
  - "Google"
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 57
AI分档: "召回候选"
AI质检状态: "不通过"
AI打分理由: "正文主线是在Google Cloud TPU v6e上用vLLM对Gemma 3 12B/27B做分类与生成两类LLM推理负载的基准测试，属于推理性能调优，而非超节点/机架级AI基础设施新架构、标准、量产或部署。来源为Google Cloud官方开发者博客，权威性较好，但文章是运维/选型类实测指南，没有发布新的机架级硬件或协议标准，没有具名客户、量产或生产部署里程碑。历史知识库已有多篇Google TPU v8/AI Hypercomputer机架级信息，本文未新增超节点级可核验事实；虽有TPU v6e 2x2拓扑、vLLM参数、吞吐和延迟数据，技术细节尚可，但无法进入任何高价值准入通道，且命中硬否决第1条教程与运维选型，按最高54分本应更严；鉴于基准数据具体、可作为TPU v6e推理选型参考，且未完全脱离AI基础设施，总分定在55—74的召回候选区间。"
AI质检模型: "zj-deepseek-v4-flash"
AI质检时间: "2026-09-07T22:26:15+08:00"
AI主题相关性: 9
AI来源权威性: 13
AI新颖性: 6
AI技术细节: 15
AI商业部署信号: 5
AI完整性: 9
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.97"
AI评分知识库SHA256: "093c22b302b0d377ed12eab9c169d422553a099393d75a2ea63f76a9eb1d8782"
AI评分知识库检索词: "[\"Google\",\"performance\",\"https://cloud.google.com/blog/\",\"RAS\",\"NPU\",\"LLM\",\"TPU\",\"LLMs\",\"v6e\",\"TPUs\",\"CPU/Memory\",\"E2E\"]"
AI评分知识库命中: "[{\"id\":\"runtime-e8b946c3131877fe7add476b\",\"title\":\"COMPUTE Peeling Apart That Supposed $120 Billion Chip Deal Google Inked With Marvell August 27, 2026\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-27\",\"matchedTerms\":[\"Google\",\"RAS\",\"TPU\",\"TPUs\"],\"rank\":-16.715094752368376},{\"id\":\"historical-jan-apr-02\",\"title\":\"二、Google Cloud Next '26：AI Hypercomputer 与第八代 TPU 发布\",\"sourceType\":\"curated_item\",\"time\":\"2026-01_to_2026-04\",\"matchedTerms\":[\"Google\",\"performance\",\"TPU\"],\"rank\":-15.126166862584078},{\"id\":\"historical-jun-025\",\"title\":\"英伟达、谷歌与国产超节点的三种网络选择\",\"sourceType\":\"curated_item\",\"time\":\"2026-06\",\"matchedTerms\":[\"NPU\",\"TPU\"],\"rank\":-12.192404405576479},{\"id\":\"runtime-ecb59d92bc701c622f102404\",\"title\":\"Broadcom slams AI silicon-fueled Q3, promises double-double on the way\",\"sourceType\":\"ai_excellent_article\",\"time\":\"\",\"matchedTerms\":[\"Google\",\"RAS\",\"TPU\",\"TPUs\"],\"rank\":-11.72233163777794},{\"id\":\"july-correct-0024\",\"title\":\"Salience Labs Wants To Scale Up AI With Silicon Photonics Optical Switch\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"Google\",\"performance\",\"TPU\",\"TPUs\"],\"rank\":-11.677184292121638}]"
采集批次: "2026年9月7日22点05分18秒"
采集批次ID: "20260907-220517-361"
去重键: "https://cloud.google.com/blog/topics/developers-practitioners/not-all-llm-workloads-are-equal-benchmarking-tpu-performance-on-classification-vs-generation"
---

Developers & Practitioners

##### Rupjit Chakraborty

AI Engineer

Moving Large Language Models (LLMs) from experimental prototypes into enterprise production exposes a critical truth: your infrastructure dictates both your performance ceilings and your unit economics. Standard hardware benchmarks often ignore a fundamental reality—not all LLM requests stress the silicon in the same way.

In this post, we dive into a comprehensive benchmarking exercise comparing Gemma 3 12B and Gemma 3 27B on Google Cloud TPU v6e to answer a crucial architectural question: How does TPU infrastructure actually perform when tasked with structurally distinct workloads at scale?

## Key Findings and Suggestions

Before diving into the methodology, here are the critical takeaways for architects deploying Gemma 3 on TPU v6e:

### The Generation Performance Wall

For decode-heavy generation tasks, the Gemma 3 27B model hits a strict performance wall past 64 concurrent users, plateauing at a 4.12x normalized throughput multiplier at 128 users. In contrast, the 12B model scales up to an 8.19x multiplier.

**Suggestion**: If your workload requires high-concurrency generation, downsize to the 12B model, or set strict pod-autoscaling limits capping concurrent requests at 64 per replica for the 27B model.

### The Classification Parity

For prefill-heavy classification tasks, model parameter size matters significantly less. Both the 12B and 27B models achieve similar peak scaling (around 6.0x to 6.4x normalized throughput at 128 users) without saturating the TPUs.

**Suggestion**: You can safely deploy larger, more capable models for summarization or classification workflows without paying a throughput penalty. The average --max-num-seqs or --max-model-len should be kept judiciously based on the average user load and average tokens per request, without which there might be request drops.

### Designing Around the Wall

Hardware saturation manifests as severe latency spikes and silent request dropouts. To mitigate this, do not rely on standard CPU/Memory scaling triggers. Instead, scale based on End-to-End (E2E) latency metrics, and implement aggressive vLLM bucket padding optimizations (VLLM\_TPU\_BUCKET\_PADDING\_GAP) to conserve memory.

## The Architecture Setup

The inference stack can be divided into three core pillars:

**1\. Infrastructure: GKE & TPU**

The foundation of our deployment is a Google Kubernetes Engine (GKE) Autopilot cluster. Connected to this is a single-host TPU v6e node pool configured with a 2x2 chip topology.

**2\. Software & Tools: vllm**

For the serving framework, we leveraged vllm via [vllm-project/tpu-inference](https://github.com/vllm-project/tpu-inference).

**3\. Models: Gemma 3 12B and 27B**

We evaluated two highly capable open-weights models: Gemma 3 12B and Gemma 3 27B. These models were accessed via HuggingFace.

## The Workloads: Classification vs. Generation

Not all LLM requests stress the system equally. We benchmarked two distinct scenarios: Classification and Generation, across 16, 32, 64, and 128 concurrent users:

- **Classification (High Input, Low Output):** This use case mimics an e-commerce compliance task. The prompt includes large blocks of product rules, item descriptions, and OCR-extracted text. The output is exceptionally small—typically just classifying an item as "Allow" or "Prohibit". Input Sequence Length (ISL) is ~4,000 tokens and Output Sequence Length (OSL) is ~10 tokens.
- **Generation (Low/Medium Input, High Output):** This use case mimics long-form text generation. The prompt requests a detailed, analytical policy brief on the future of AI in the labor market. The model spends the majority of its time decoding and streaming out hundreds of tokens. Input Sequence Length (ISL) is 500 tokens and Output Sequence Length (OSL) is ~1,000 tokens.

## Results and Observations

We measured metrics like Throughput (requests/sec), End-to-End Latency and the results provided some fascinating insights into how parameter size and hardware bandwidth interact. To ensure architectural consistency, every benchmark was executed using the [vllm-project/tpu-inference](https://github.com/vllm-project/tpu-inference) hardware plugin, leveraging a standardized global serving configuration of **max-model-len=128000**, **max-num-batched-tokens=8192**,and **max-num-seqs=512**.

### Generation Scaling Divergence

In Generation tasks, both models perform similarly up to 64 concurrent users. However, at 128 concurrent users, the Gemma 3 12B model shows significantly better scaling, achieving an 8.19x normalized throughput multiplier compared to a 4.12x plateau for the Gemma 3 27B model (normalized against the Gemma 3 12B baseline at 16 users). This suggests that the larger 27B model hits memory or compute limits much earlier under high generation loads.

| **Concurrent Users** | **Gemma 3 12B Throughput (req/s)** | **Gemma 3 27B Throughput (req/s)** |
| --- | --- | --- |
| 16 users | 1.00 x | 1.05 x |
| 32 users | 1.98 x | 1.97 x |
| 64 users | 2.96 x | 4.00 x |
| 128 users | 8.19 x | 4.12 x |

![https://storage.googleapis.com/gweb-cloudblog-publish/images/generation_scaling.max-1800x1800.jpg](https://storage.googleapis.com/gweb-cloudblog-publish/images/generation_scaling.max-1800x1800.jpg)

### Classification Performance Parity

In Classification tasks, there is negligible difference in scaling behavior between the Gemma 3 12B and Gemma 3 27B models. Both models operate efficiently within the hardware's capacity and scale well, reaching peak normalized throughputs of approximately 6.04x to 6.37x at 128 concurrent users (normalized against the Gemma 3 12B baseline at 16 users).

| **Concurrent Users** | **Gemma 3 12B Throughput (req/s)** | **Gemma 3 27B Throughput (req/s)** |
| --- | --- | --- |
| 16 users | 1.00 x | 0.76x |
| 32 users | 1.18x | 1.53x |
| 64 users | 2.04x | 3.15x |
| 128 users | 6.37x | 6.04x |

![https://storage.googleapis.com/gweb-cloudblog-publish/images/classification_perf_table_image2.max-1800x1800.png](https://storage.googleapis.com/gweb-cloudblog-publish/images/classification_perf_table_image2.max-1800x1800.png)

### Latency Threshold Analysis

End-to-End (E2E) latency exhibits different scaling behaviors depending on the model size and task. When using identical serving hyperparameters (--max-num-seqs=512), the Gemma 3 12B model's Classification latency roughly doubles when moving from 32 users to 64 users, indicating resource contention. However, for the larger Gemma 3 27B model, Classification latency remains relatively flat between 32 and 64 users before doubling at the 128-user mark.

| **Model** | **Task** | **16 Users** | **32 Users** | **64 Users** | **128 Users** |
| --- | --- | --- | --- | --- | --- |
| Gemma 3 12B | Generation | 1.00x | 1.13x | 1.40x | 1.70x |
| Gemma 3 12B | Classification | 1.00x | 0.99x | 1.79x | 2.90x |
| Gemma 3 27B | Generation | 1.20x | 1.68x | 2.93x | 3.33x |
| Gemma 3 27B | Classification | 1.20x | 1.95x | 1.95x | 3.88x |

![https://storage.googleapis.com/gweb-cloudblog-publish/images/latency_threshold_image3.max-1800x1800.png](https://storage.googleapis.com/gweb-cloudblog-publish/images/latency_threshold_image3.max-1800x1800.png)

## Conclusion

Benchmarking Gemma 3 12B and 27B models on Google Cloud TPU v6e architecture reveals that raw parameter count is not the sole predictor of inference performance; rather, the interaction between the serving framework, hardware topology, and workload token ratios dictates efficiency. For generation tasks (low input, high output), the 12B model proves superior at high concurrency, sustaining an 8.19x relative throughput multiplier where the 27B model saturates at 4.12x. Conversely, for prefill-heavy classification tasks, both models perform similarly, allowing organizations to deploy larger models without a severe scaling penalty. Our evaluation also mapped exact hardware saturation thresholds—such as End-to-End latency doubling at 64 users for classification and hitting a cliff at 128 users for generation—enabling precise, data-driven auto-scaling triggers rather than costly over-provisioning. Ultimately, achieving these peak metrics requires aggressive tuning of vllm parameters, such as adjusting batched tokens and configuring TPU-specific bucket padding to prevent compute waste, proving that cost-effective AI infrastructure must strictly align model selection and serving configurations to the unique input/output profiles of production workloads.

## Ready to scale your LLM workloads?

Don't let unoptimized infrastructure bottleneck your enterprise AI rollouts. Now that you know how different workload shapes impact hardware saturation, it's time to put these insights into practice:

- **Use these benchmarks to right-size your production architecture**. Safely leverage the larger Gemma 3 27B for prefill-heavy classification tasks without a throughput penalty, but consider switching to the 12B model to maintain linear scaling for decode-heavy generation at high concurrency.
- **Deploy using** **[Google Kubernetes Engine (GKE)](https://docs.cloud.google.com/kubernetes-engine/docs/tutorials/serve-vllm-tpu) with TPU v6e node pools to build a highly scalable, managed AI foundation and dedicated [vllm-project/tpu-inference](https://github.com/vllm-project/tpu-inference)** **hardware plugin**. Alternatively, you can also deploy via [Model Garden on Gemini Enterprise Agent Platform](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/open-models/vllm/use-vllm-tpu) or you can spin up [TPU VMs](https://github.com/AI-Hypercomputer/tpu-recipes/tree/main/inference/trillium/vLLM) for serving Gemma 3 models.

Have you encountered similar performance walls in your own production deployments? Share your scaling strategies, ask questions, and join the discussion in the [Google Cloud Community forums](https://www.googlecloudcommunity.com/).

Posted in
