---
格式版本: 2
标题: "Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs"
原文链接: "https://arxiv.org/abs/2608.11624"
发布日期: "2026-08-12"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "**\\[v1\\]** Wed, 12 Aug 2026 04:11:22 UTC (561 KB)"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 18
发布时间严格候选数量: 6
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-13T18:34:55+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-13T18:24:32+08:00"
入库时间: "2026-08-13T10:34:56.185Z"
来源平台: "arXiv 学术论文搜索"
搜索渠道: "source_template"
搜索词: "https://arxiv.org/search/?query=AI&searchtype=all"
匹配关键词:
  - "AI"
相关厂家:
  []
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI质检状态: "评分失败"
AI评分尝试次数: 1
AI评分错误类型: "service_error"
AI评分错误: "LLM call failed; tried model chain: ali-deepseek-v4-flash -> tx-deepseek-v4-flash | Model ali-deepseek-v4-flash failed 503: {\"error\":{\"code\":\"model_not_found\",\"message\":\"No available channel for model ali-deepseek-v4-flash under group vip (distributor) (request id: 202608131034596758049938268d9d66bzKtdnM)\",\"type\":\"new_api_error\"}} | Model tx-deepseek-v4-flash failed 503: {\"error\":{\"code\":\"model_not_found\",\"message\":\"No available channel for model tx-deepseek-v4-flash under group vip (distributor"
AI评分开始时间: "2026-08-13T10:34:56.196Z"
AI评分结束时间: "2026-08-13T10:35:01.037Z"
AI摘要: "研究提出对抗性强化学习框架，训练劝说者用单个说服论据诱导LLM改变答案，即使论据虚假也能使模型准确率接近零；该方法的劝说成功率可达93%，并通过课程学习将GPT-4o-mini的攻击成功率提升至38%。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T03:40:23.668Z"
采集批次: "2026年8月13日18点24分27秒"
采集批次ID: "20260813-182427-1564e572"
去重键: "https://arxiv.org/abs/2608.11624"
---

## Computer Science > Computation and Language

## Title:Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs

[View PDF](https://arxiv.org/pdf/2608.11624) [HTML (experimental)](https://arxiv.org/html/2608.11624v1)

> Abstract:Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. As LLMs increasingly debate, advise, and think collaboratively with humans and each other, resistance to harmful persuasion becomes a core requirement for reliable behavior. Yet we show that this requirement is far from met: a single targeted persuasive argument is enough to collapse model accuracy to near zero, even when the argument is factually false. We formalize this threat as adversarial persuasion and introduce an adversarial reinforcement learning framework that trains persuader agents to change a target model's answer in a single interaction. First, we show that optimizing persuasion strategies through trial and error exposes vulnerabilities that static prompting misses: RL-trained persuaders raise persuasion success from approximately 24% to over 93% against the training-time persuadee. Second, we find that these learned strategies transfer to unseen models, achieving 83% attack success on Qwen-14B, 79% on Llama-3.1-8B, and 25% on GPT-4o-mini. Third, we demonstrate that a curriculum that bootstraps on more persuadable open-weight models before targeting harder models further increases GPT-4o-mini attack success from 25% to 38%. Moreover, our results reveal that optimized persuaders increasingly rely on credibility-based tactics, including fabricated citations and false authoritative evidence. Together, these findings expose a critical weakness in current LLM agents: even when they initially reason correctly, they can be steered toward false conclusions by optimized natural language influence. This positions persuasion robustness as a necessary safety criterion for multi-agent and human-AI decision-making systems.

| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| --- | --- |
| Cite as: | [arXiv:2608.11624](https://arxiv.org/abs/2608.11624) \[cs.CL\] |
|  | (or [arXiv:2608.11624v1](https://arxiv.org/abs/2608.11624v1) \[cs.CL\] for this version) |
|  | [https://doi.org/10.48550/arXiv.2608.11624](https://doi.org/10.48550/arXiv.2608.11624) |

## Submission history

From: Nimet Beyza Bozdag \[[view email](https://arxiv.org/show-email/bf49a0d9/2608.11624)\]  
**\[v1\]** Wed, 12 Aug 2026 04:11:22 UTC (561 KB)

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2608.11624) | Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
