--- 格式版本: 2 标题: "Safety overview: GPT-6 Astra" 原文链接: "https://openai.com/index/safety-overview-gpt-6-astra" 发布日期: "2026-09-06" 发布时间校准状态: "found" 发布时间需复核: "否" 发布时间来源: "llm:scrape:provider_published_at" 发布时间证据: "provider publishedAt: 2026-09-06" 发布时间校准原因: "该日期来自页面发布商元数据,标识为 provider publishedAt,且与 YAML 发布日期一致,可作为文章发布时间。" 发布时间校准置信度: "1" 发布时间候选数量: 1 发布时间严格候选数量: 0 发布时间原页读取状态: "source template page reused from URL open" 发布时间未找到原因: "" 发布时间校准时间: "2026-09-07T22:33:28+08:00" 发布时间仲裁状态: "confirmed" 发布时间仲裁尝试次数: 1 发布时间仲裁耗时毫秒: 4769 发现时间: "2026-09-07T22:06:43+08:00" 入库时间: "2026-09-07T14:33:38.258Z" 来源平台: "固定入口" 搜索渠道: "fixed_url" 搜索词: "https://openai.com/news/security/" 匹配关键词: - "deployment" 相关厂家: - "OpenAI" 相关专家: [] 内容类型: "网页" 抓取工具: "Free Fetch + Defuddle" 清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取" 原始附件: [] AI优质: "否" AI打分: 39 AI分档: "非优质" AI质检状态: "不通过" AI打分理由: "正文主线是GPT-6 Astra模型安全发布与对齐、防越狱、CoT监控等AI模型安全内容,未讨论超节点、AI Rack、机柜级系统或关键部件;主题相关性极低。来源为OpenAI官方,权威性高,发布时间2026-09-03属新发布且有模型安全能力升级信息,但技术细节集中在模型行为、监控和红队测试,无rack-scale硬件、互连、供电或液冷架构细节;无超节点商业部署信号。完整性较好,官方安全概述含多项评估链接。历史知识库中无同页重复,但本文不属于机架级AI基础设施业务范围,未命中任何高价值准入通道。硬否决:属单一模型/应用安全发布而非机架级基础设施,按应用与模型效率类否决处理。" AI质检模型: "zj-deepseek-v4-flash" AI质检时间: "2026-09-07T22:33:52+08:00" AI主题相关性: 1 AI来源权威性: 15 AI新颖性: 14 AI技术细节: 0 AI商业部署信号: 0 AI完整性: 9 AI评分提示词版本: "v17-精简生产版" AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d" AI评分知识库版本: "knowledge_base_v1-20260819+runtime.97" AI评分知识库SHA256: "093c22b302b0d377ed12eab9c169d422553a099393d75a2ea63f76a9eb1d8782" AI评分知识库检索词: "[\"OpenAI\",\"https://openai.com/news/security/\",\"GPT-6\",\"GPT\",\"deploymentsafety.openai.com/gpt-6-astra\"]" AI评分知识库命中: "[{\"id\":\"runtime-b9406ae170bd91336ab3bb25\",\"title\":\"Jalapeño’s first results show industry-leading speed and efficiency in AI inference\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-25\",\"matchedTerms\":[\"OpenAI\",\"GPT\"],\"rank\":-6.982962704341991},{\"id\":\"runtime-51cfe3db04ed0bd3efb5e0e3\",\"title\":\"The full stack behind abundant intelligence\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-25\",\"matchedTerms\":[\"OpenAI\",\"GPT\"],\"rank\":-6.8682007745958185},{\"id\":\"july-correct-0077\",\"title\":\"AMD and OpenAI: Collaborating Across Every Layer of the Stack\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"OpenAI\",\"GPT\"],\"rank\":-6.245589686617948},{\"id\":\"runtime-4abbccc42d96af674efc7768\",\"title\":\"OpenAI Jalapeño: Better Than Nvidia Blackwell\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-25\",\"matchedTerms\":[\"OpenAI\",\"GPT\"],\"rank\":-5.617411662858215},{\"id\":\"july-correct-0015\",\"title\":\"AMD to join the optical interconnect party with 2027 Instinct GPUs\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"OpenAI\",\"GPT\"],\"rank\":-4.864669687823775}]" 采集批次: "2026年9月7日22点05分18秒" 采集批次ID: "20260907-220517-361" 去重键: "https://openai.com/index/safety-overview-gpt-6-astra" --- Today, we are releasing GPT‑6 Astra, the most capable model we have ever broadly deployed. Astra is our first model to reach the Critical level of cybersecurity capability under our Preparedness Framework. The most important things to know about the safety of this launch are as follows: 1. **GPT‑6 Astra is a significant step up in cyber capabilities and meets our Critical threshold.** This means that, with the right tools and access, GPT‑6 Astra can find previously unknown security flaws and develop new ways to exploit them across many well-protected systems without a person guiding each step. Accordingly, we significantly strengthened our protections against the model taking harmful cyber actions, whether that’s due to misuse or misalignment. We also took steps to secure our internal development and deployment of Astra and similar models, including stricter isolation, checkpoint encryption, universal monitoring of full trajectories including chains of thought (CoT), and a blocking alignment evaluation process before internal use. 2. **GPT‑6 Astra is significantly more robust than its predecessors**. Incorporating new robustness safety training techniques, GPT‑6 Astra is significantly more robust to jailbreaks than GPT‑5.6 Sol, including across longer trajectories. We know this from offline tests and our program of rigorous internal and external jailbreak testing and remediation. For users flagged as potentially high risk, we have additionally trained in the ability to adjust the model’s refusal boundary to be more conservative and cover a broader range of dual use risks. We use regression testing to make sure Astra is robust against jailbreaks found in previous testing periods, and performed new rounds of automated red-teaming with our latest internal red-teaming attackers to validate our improvements. 3. **GPT‑6 Astra is better aligned than GPT‑5.6 Sol.** Astra is a significant step forward in model alignment, and includes alignment improvements spanning the composition of our pre-training data to our grading during reinforcement learning. We report a [new suite of alignment evaluations⁠](https://deploymentsafety.openai.com/gpt-6-astra/alignment), and find that GPT‑6 Astra is stronger at respecting safety and security boundaries and staying within its authorized scope. In a simulation using more than [54,000 internal Codex tasks⁠](https://deploymentsafety.openai.com/gpt-6-astra/forecasting-misaligned-behavior-with-deployment-simulation-of-internal-codex-traffic), Astra also received roughly half as many flags for higher-severity misaligned behavior as Sol. 4. **We are deploying misalignment monitoring broadly**. We view model alignment as the primary lever to prevent potential misaligned behavior from our models. However, monitoring provides broad visibility into frontier model behavior, illuminating opportunities to further improve alignment and safety. In addition, monitoring serves as an additional layer of protection against misaligned behavior that is detected. For these reasons, we have additionally added misalignment monitoring to all tool-using inference involved in our external deployment of Astra, with significant compute cost. This system parallels our internal setup. 5. **GPT‑6 Astra’s monitorability has decreased relative to GPT‑5.6 Sol.** We have performed significant investigations on the [monitorability⁠](https://deploymentsafety.openai.com/gpt-6-astra/monitorability) and [controllability⁠](https://deploymentsafety.openai.com/gpt-6-astra/cot-controllability) of GPT‑6 Astra. We have found that GPT‑6 Astra is more capable of controlling its own CoT than GPT‑5.6 Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform [certain sabotage tasks⁠](https://deploymentsafety.openai.com/gpt-6-astra/monitor-evasion). We have not seen evidence of steganographic CoT reasoning (e.g., concealing hidden reasoning within ordinary and seemingly unrelated text), indicating that the evasion risk may largely be bounded to lower reasoning tasks. - These findings indicate that the Astra class models could evade our CoT monitors under adversarial conditions. Right now, these findings are largely based on adversarial evaluations (i.e., when we instruct the model to evade monitoring), and overall our alignment evaluations show that Astra is less likely than GPT‑5.6 Sol to violate security and safety restrictions overall. - However, we take the trend seriously, and we are continuing to investigate these findings and their implications for monitorability as our models become more capable. OpenAI has worked to preserve and utilize chain of thought monitoring, and preserving CoT monitorability is a core goal of the research program. However, these results also underscore the importance of developing alignment auditing techniques beyond examining the model’s chain of thought. 6. **GPT‑6 Astra more responsibly navigates browsing and workplace settings:** GPT‑6 Astra is significantly more robust to prompt injections than GPT‑5.6 Sol. We have additionally tested the model’s behavior in realistic browsing and professional computer environments, and find that the model is significantly less likely to perform misaligned and potentially destructive actions (for instance unauthorized transactions, data loss, excessive access, or circumvention of controls) compared to GPT‑5.6 Sol. It also acts more safely when handling harmful requests in agentic settings, such as requests to assist with violent attack planning or commit fraud. 7. **GPT‑6 Astra is significantly safer in higher-risk scenarios.** GPT‑6 Astra responds more safely than GPT‑5.6 Sol to challenging requests drawn from production and adversarial human red-teaming. Astra achieves a Pareto improvement in safely completing unsafe requests and avoiding unnecessary refusals to harmless requests. These improvements extend to high-severity scenarios where the risk of harm emerges from the broader context rather than an explicit request. Astra also applies age-appropriate safety boundaries more consistently for users under 18. For more information, see the [full system card⁠](http://deploymentsafety.openai.com/gpt-6-astra).