---
格式版本: 2
标题: "Issue #399 - The ML Engineer 🤖"
原文链接: "https://machinelearning.substack.com/p/issue-399-the-ml-engineer"
发布日期: "2026-08-09"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:scrape:strict_html_body"
发布时间证据: "div class=pencraft pc-display-flex pc-gap-12 pc-alignItems-center pc-reset byline-wrapper: Aug 09, 2026"
发布时间校准原因: "规则确认唯一严格发布时间，来源 scrape:strict_html_body"
发布时间校准置信度: "high"
发布时间候选数量: 11
发布时间严格候选数量: 2
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-31T11:33:15+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-31T11:29:39+08:00"
入库时间: "2026-08-31T03:33:15.517Z"
来源平台: "Substack 数据中心相关博客搜索"
搜索渠道: "source_template"
搜索词: "site:substack.com Open AI Infra Summit"
匹配关键词:
  - "AI"
  - "GPU"
  - "deployment"
  - "performance"
相关厂家:
  - "NVIDIA"
  - "OpenAI"
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 26
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文是ML工程资讯汇编，主线为AI安全、Agent记忆、模型发布及应用开发，不讨论超节点、AI整机柜或机架级基础设施。来源为Substack博客，主要转述外部内容，且发布日期缺失。虽新增列举Qwen3.8-Max、Shieldstral等模型规格和发布信息，但没有机架拓扑、互连、供电、液冷、RAS或规模部署事实；固定知识库也未提供可支持其为超节点新增事件的历史对照。命中应用与模型效率强否决项，当前页面不具备超节点业务信息源价值。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-08-31T11:33:35+08:00"
AI主题相关性: 1
AI来源权威性: 5
AI新颖性: 7
AI技术细节: 4
AI商业部署信号: 1
AI完整性: 8
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.59"
AI评分知识库SHA256: "630ed34f914ee61e1f2c29bf5210d738043637a2eaf7d9e8a3ea68614a92fc92"
AI评分知识库检索词: "[\"Open AI Infra Summit\",\"site:substack.com Open AI Infra Summit\",\"NPU\",\"GPU\",\"NVIDIA\",\"Meta\",\"Intel\",\"ML\",\"AIAlignment\",\"UI\",\"NA\",\"PR\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0088\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes - 智源社区论文\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NPU\",\"NVIDIA\",\"ML\",\"NA\",\"PR\"],\"rank\":-12.427463611417796},{\"id\":\"historical-may-025\",\"title\":\"附下载｜2026 Open AI Infra Summit 成果发布：标准化引领 AI 基础设施规模化落地\",\"sourceType\":\"curated_item\",\"time\":\"2026-05\",\"matchedTerms\":[\"Open AI Infra Summit\"],\"rank\":-8.720874692758809},{\"id\":\"july-correct-0081\",\"title\":\"AAI 2026: 6th Gen AMD EPYC Server CPUs Power the Agentic Data Center\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"GPU\",\"NVIDIA\",\"Intel\",\"ML\",\"UI\",\"NA\",\"PR\"],\"rank\":-7.8809782404555815},{\"id\":\"july-correct-0034\",\"title\":\"AMD Fires Back at Nvidia with Helios AI System, Epyc CPUs\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"GPU\",\"NVIDIA\",\"Meta\",\"Intel\",\"UI\",\"NA\",\"PR\"],\"rank\":-7.78315064967736},{\"id\":\"historical-jun-038\",\"title\":\"数据中心液冷“出问题”，怎么第一时间发现，这套监测体系给了答案\",\"sourceType\":\"curated_item\",\"time\":\"2026-06\",\"matchedTerms\":[\"Open AI Infra Summit\",\"GPU\"],\"rank\":-7.7822704058468215}]"
AI摘要: "Institute for Ethical AI Alignment & Safety重建官网并发布本周ML工程通讯，核心进展包括Qwen3.8-Max将首次开放2.4万亿参数权重。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-07T02:08:29.914Z"
采集批次: "2026年8月31日11点14分25秒"
采集批次ID: "20260831-111425-728"
去重键: "https://machinelearning.substack.com/p/issue-399-the-ml-engineer"
---

## The New Institute for Ethical AIAlignment & Safety

We are excited to announce the new face of [the Institute for Ethical AI Alignment & Safety](https://ethical.institute/)!

We have rebuilt **the Institute website** from the ground up, together with some of our key initiatives.

**Since our founding in 2017**, we have built a track record of contributions across public and private institutions.

Our work runs from **individual practice** to **national regulation;** by principle, by process, by standards, by regulation.

Several of our recommendations have been **adopted across the EU and UK policy.** Our mission has not changed since 2017.

We look forward to continue contributing to a **future** where frontier **AI is safe, aligned and accountable** to people and society.

**Check it out:** [https://ethical.institute](https://ethical.institute/)

## This week in ML Engineering:

- Multi-Tenant Memory [for AI Agents](https://hackernoon.com/whose-memory-is-it-building-multi-tenant-multi-tier-memory-for-ai-agents-part-1)
- LLMs [Reward Expertise](https://www.seangoedecke.com/llms-reward-expertise/)
- Qwen3.8-Max [and Open Weights](https://qwen.ai/blog?id=qwen3.8)
- Harness Design [for Long-Running Apps](https://www.anthropic.com/engineering/harness-design-long-running-apps)
- Mistral’s [Shieldstral Safety Classifier](https://mistral.ai/news/shieldstral/)
- Open Source [ML Frameworks](http://localhost:4321/open-source/production-ml-list/)
- Awesome AI Guidelines [to check out this week](http://localhost:4321/open-source/ai-guidelines/)
- \+ more 🚀

## Multi-Tenant Multi-Tier Memory for AI Agents

Our 4-part series on agent memory has been published & featured in the front-page of HackerNoon homepage as a top story! LLMs are stateless by design, so without a memory layer every session starts from zero, and the number of dedicated memory tools has been growing almost daily. This came out of my recent work extending the Kubernetes Agent Orchestration System (KAOS) to support multi-tiered memory persistence (aka short-, medium- and long-term memory). Along the way I hit most of the same issues that anyone would when building or integrating multi-tiered memory into an agentic system, so I thought it would be useful to compile all the learnings, design choices and examples. Check it out, together with the rest of the series!

## LLMs Reward Expertise

Do language models still reward expertise, or has that ship sailed? Sean Goedecke argues for the second; the distinction is between getting something usable out of a model and extracting the maximum value from it. A non-expert can get sort-of-okay Python, but only someone who knows what a good answer looks like can evaluate the output critically. His main evidence is Terence Tao’s published ChatGPT conversation on the Jacobian Conjecture, where Tao’s messages are short and to the point, and the model answers in a talking-to-mathematicians register rather than an explaining-to-amateurs one. Tao pushes back and he almost never takes the model’s advice. He is careful with the caveat that non-experts still get real value, and that OpenAI had a team of expert mathematicians filtering the suggestions, a step you cannot currently skip. For production ML practitioners the takeaway is that knowledge is still more improtant than ever, and now is even becoming the bottleneck; does this mean we need to accelerate our learning?

## Qwen3.8-Max and Open Weights

Qwen3.8-Max bas been released! And for the first time they say a Qwen-Max-class model will get open weights: the model scales to 2.4 trillion parameters with 95B active, is built on the Qwen 3.5 architecture, and is available through QwenCloud now with the weights promised next week. The more interesting part is actually the long-horizon runs, as they report a ten day autonomous coding run building the oh-my-cli project, which after roughly 16 days of fully autonomous operation had accumulated 265 commits, 127 PRs and 151 issues! Whether that is good or bad work is yet to be seen… Open weights at this scale would be quite something; let’s see what actually lands next week.

## Harness Design for Long-Running Apps

Anthropic has written up the harness they use to have Claude build entire applications over multi-hour runs, and more interestingly what they deleted from it as the model improved: The architecture is three agents framed as a separation of generator and judge, with a planner that expands a one to four sentence prompt into a full product spec, a generator that implements against it, and an evaluator that drives the running app through Playwright MCP like a real user, checking UI, API endpoints and database state. It is also quite a well timed piece as it comes next to their piece on [how they contain Claude across products](https://www.anthropic.com/engineering/how-we-contain-claude), where filesystem and egress boundaries are what holds once the model-layer defences fail.

## Mistral’s Shieldstral Safety Classifier

Europe strikes again! Mistral has released Shieldstral, an open-weights multimodal safety classifier under Apache 2.0: Shieldstral is a 3B parameter model covering text and images, it runs on a single 16GB NVIDIA GPU, and the weights are on Hugging Face. It is interesting how they are proposing to reframe moderation as binary question answering rather than fixed label classification. Basically a plain-language yes/no question and the document being judged as the three inputs; the output is a calibrated probability read off a single token, so a policy change means rewriting the question instead of retraining. They claim it outperforms models up to 7x its size and claim a new state of the art on multimodal safety, although the post does not publish per-benchmark numbers to check that against. One classifier instead of a guardrail model per deployment is a good direction, and great to see it shipped under a permissive licence!

## Upcoming MLOps Events

The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.

### Events we are speaking at this year:

- [Signals Conference](https://signalsconf.io/#tickets) - September @ Berlin
- [World Summit AI Europe](https://worldsummit.ai/) - September @ Amsterdam

### Other relevant events:

- [AI Infra Summit 2026](https://www.ai-infra-summit.com/) - Sept @ California
- [Code.Talks 2026](https://codetalks.com/) - Nov @ Hamburg
- [MLOps World 2026](https://mlopsworld.com/) - Nov @ Austin

### In case you missed our talks, check our recordings below:

- The State of AI in 2025 - [WeAreDevelopers 2025](https://www.youtube.com/watch?v=v2LENQOG-Xg)
- Prod Generative AI in 2024 - [KubeCon AI Day 2025](https://www.youtube.com/watch?v=0uJGmMZGUJE&list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&index=3)
- The State of AI in 2024 - [WeAreDevelopers 2024](https://www.youtube.com/live/AtA2XXo_b5s)
- Responsible AI Workshop Keynote - [NeurIPS 2021](http://www.youtube.com/watch?v=57YpXjcj0Ho)
- Practical Guide to ML Explainability - [PyCon London](http://www.youtube.com/watch?v=vq8mDiDODhc)
- ML Monitoring: Outliers, Drift, XAI - [PyCon Keynote](http://www.youtube.com/watch?v=QcevzK9ZuDg)
- Metadata for E2E MLOps - [Kubecon NA 2022](https://www.youtube.com/watch?v=OSbH4dfswCY)
- ML Performance Evaluation at Scale - [KubeCon Eur 2021](http://www.youtube.com/watch?v=8ORl8lu1Eeo)
- Industry Strength LLMs - [PyData Global 2022](https://www.youtube.com/watch?v=RVUi_rAFfzU)
- ML Security Workshop Keynote - [NeurIPS 2022](http://www.youtube.com/watch?v=7XSy5aw8oU8)

## Open Source MLOps Tools

Check out the fast-growing ecosystem of production ML tools & frameworks at [the github repository](http://localhost:4321/open-source/production-ml-list/) which has reached over 20,000 ⭐ github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here’s a few featured open source libraries that we maintain:

- [SARC](https://github.com/besanson/sarc-governance/tree/main) - Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.
- [KAOS](https://github.com/axsaucedo/kaos) - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.
- [Kompute](https://github.com/KomputeProject/kompute) - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.
- [Production ML Tools](http://localhost:4321/open-source/production-ml-list/) - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.
- [AI Policy List](https://github.com/EthicalML/awesome-artificial-intelligence-regulation) - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.
- [Agentic Systems Tools](https://github.com/EthicalML/awesome-production-genai/) - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain

Please do support some of our open source projects by sharing, contributing or adding a star ⭐

## About us

The Institute for Ethical AI & Machine Learning is a European research centre that carries out world-class research into responsible machine learning.

[Check out our website](https://ethical.institute/)

✉️ [Email](mailto:EMAIL?subject=Check%20out%20the%20Machine%20Learning%20Engineering%20Newsletter!&body=Check%20out%20this%20weekly%20newsletter%20on%20Machine%20Learning!%20Join%20for%20free%20here%3A%20https%3A%2F%2Fethical.institute%2Fmle.html), 🐦 [Twitter](http://twitter.com/EthicalML), 💼 [Linkedin](https://www.linkedin.com/company/the-institute-for-ethical-machine-learning/)
