---
格式版本: 2
标题: "Testing Agent Skills Systematically with Evals | OpenAI Developers"
原文链接: "https://developers.openai.com/blog/eval-skills"
发布日期: "2026-09-06"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:local:strict_original_body"
发布时间证据: "Published Time: Sun, 06 Sep 2026 10:26:38 GMT"
发布时间校准原因: "规则确认唯一严格发布时间，来源 local:strict_original_body"
发布时间校准置信度: "high"
发布时间候选数量: 4
发布时间严格候选数量: 1
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-06T18:26:39+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-06T17:44:27+08:00"
入库时间: "2026-09-06T10:28:59.656Z"
来源平台: "固定入口"
搜索渠道: "fixed_url"
搜索词: "https://developers.openai.com/blog"
匹配关键词:
  - "deployment"
  - "performance"
  - "latency"
  - "throughput"
  - "AI"
相关厂家:
  - "OpenAI"
  - "Microsoft"
  - "AWS"
  - "Google"
  - "Oracle"
相关专家:
  []
内容类型: "网页"
抓取工具: "Jina Reader"
清洗工具: "Jina Reader Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 27
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文是OpenAI官方开发者教程，主线为使用Codex、JSONL事件和结构化评分器测试Agent Skills，实际正文日期为2026年1月22日。内容完整且来源权威，但不涉及超节点、AI机架、互连、供电、液冷或机架级部署；固定知识库也未提供与该教程相关的历史新增依据。新增内容仅是应用层评测流程与示例，命中“教程与运维选型”及应用/模型效率类硬否决，不具备超节点业务准入价值。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-09-08T16:34:44+08:00"
AI主题相关性: 0
AI来源权威性: 15
AI新颖性: 3
AI技术细节: 0
AI商业部署信号: 0
AI完整性: 9
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.88"
AI评分知识库SHA256: "187ddf9763173d03e5c13ef1e1afe58e86093a26ab37f0e4d189c8eb19dacf5b"
AI评分知识库检索词: "[\"OpenAI\",\"Microsoft\",\"AWS\",\"Google\",\"Oracle\",\"deployment\",\"performance\",\"latency\",\"throughput\",\"https://developers.openai.com/blog\",\"RAS\",\"NPU\"]"
AI评分知识库命中: "[{\"id\":\"july-correct-0018\",\"title\":\"AMD, Cerebras partner on joint Helios rack-scale AI inference platform\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"OpenAI\",\"Microsoft\",\"Oracle\",\"performance\",\"latency\",\"throughput\",\"RAS\"],\"rank\":-20.408108891439618},{\"id\":\"july-correct-0033\",\"title\":\"Microsoft, Alphabet, Meta Pivot from Buy to Build in AI\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"OpenAI\",\"Microsoft\",\"AWS\",\"Google\",\"deployment\",\"performance\",\"RAS\"],\"rank\":-18.845160932854792},{\"id\":\"july-correct-0080\",\"title\":\"AMD and Cerebras Announce Industry-Leading Ultra-Low-Latency and High Throughput AI Inference Solution\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"OpenAI\",\"Microsoft\",\"AWS\",\"deployment\",\"performance\",\"latency\",\"throughput\",\"RAS\"],\"rank\":-18.001565535425346},{\"id\":\"runtime-15c79e6868a9b43f99c061a9\",\"title\":\"Nvidia’s AI Boom Tests Data Center Infrastructure Limits\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-27\",\"matchedTerms\":[\"OpenAI\",\"AWS\",\"Oracle\",\"deployment\",\"performance\",\"RAS\"],\"rank\":-15.23727524593058},{\"id\":\"runtime-4abbccc42d96af674efc7768\",\"title\":\"OpenAI’ Jalapeño: Better Than Nvidia Blackwell\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-25\",\"matchedTerms\":[\"OpenAI\",\"Microsoft\",\"AWS\",\"Google\",\"deployment\",\"performance\",\"latency\",\"throughput\",\"RAS\",\"NPU\"],\"rank\":-14.196371038268706}]"
采集批次: "2026年9月6日17点41分46秒"
采集批次ID: "20260906-174146-0dbc83f7"
去重键: "https://developers.openai.com/blog/eval-skills"
---

Title: Testing Agent Skills Systematically with Evals | OpenAI Developers

URL Source: https://developers.openai.com/blog/eval-skills

Published Time: Sun, 06 Sep 2026 10:26:38 GMT

Markdown Content:
For the complete documentation index, see [llms.txt](https://developers.openai.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL. 

[![Image 1: OpenAI Developers](https://developers.openai.com/OpenAI_Developers.svg)ChatGPT](https://developers.openai.com/)

[Home](https://developers.openai.com/)

[API](https://developers.openai.com/api/docs)

[Codex](https://learn.chatgpt.com/docs)

[Docs Guides, concepts, and product docs for Codex](https://learn.chatgpt.com/docs)[Use cases Example workflows and tasks teams can take on with ChatGPT or Codex](https://learn.chatgpt.com/use-cases)

[Docs](https://developers.openai.com/codex)

[Use cases](https://developers.openai.com/codex/use-cases)

[Training](https://developers.openai.com/training)

[Resources](https://developers.openai.com/codex/resources)

[ChatGPT](https://developers.openai.com/chatgpt)

[Plugins Extend ChatGPT and Codex](https://developers.openai.com/plugins)[Workspace Agents Trigger published ChatGPT workspace agents](https://developers.openai.com/workspace-agents)[Commerce Build commerce flows in ChatGPT](https://developers.openai.com/commerce)[Ads Publish and measure ads in ChatGPT](https://developers.openai.com/ads)

[Resources](https://developers.openai.com/learn)

[Showcase Demo apps to get inspired](https://developers.openai.com/showcase)[Blog Learnings and experiences from developers](https://developers.openai.com/blog)[Cookbook Notebook examples for building with OpenAI models](https://developers.openai.com/cookbook)[Learn Docs, videos, and demo apps for building with OpenAI](https://developers.openai.com/learn)[Community Programs, meetups, and support for builders](https://developers.openai.com/community)

Start searching

[API Dashboard](https://platform.openai.com/login)

[Try ChatGPT](https://chatgpt.com/)

## Search developer resources

Search docs 

### Suggested

responses create reasoning_effort realtime prompt caching

 Primary navigation 

 API  Codex  ChatGPT  Docs  Use cases  Training  Resources  Resources 

Search docs 

### Suggested

responses create reasoning_effort realtime prompt caching

 Overview  Models  Agents  Tools  Voice & Audio  Production  API reference 

Docs Overview

*   [Home](https://developers.openai.com/api/docs)

### Get started

*   [Quickstart](https://developers.openai.com/api/docs/quickstart)
*   [Using GPT-6 Astra](https://developers.openai.com/api/docs/guides/latest-model)
*   [Key concepts](https://developers.openai.com/api/docs/concepts)

### Core concepts

*   [Responses API](https://developers.openai.com/api/docs/guides/migrate-to-responses)
*   [Conversation state](https://developers.openai.com/api/docs/guides/conversation-state)
*   [Background mode](https://developers.openai.com/api/docs/guides/background)
*   [Streaming](https://developers.openai.com/api/docs/guides/streaming-responses)
*   [WebSocket mode](https://developers.openai.com/api/docs/guides/websocket-mode)
*   [Mid-turn steering](https://developers.openai.com/api/docs/guides/steering)
*   [Multi-agent](https://developers.openai.com/api/docs/guides/responses-multi-agent)
*   [Webhooks](https://developers.openai.com/api/docs/guides/webhooks)
*   [File inputs](https://developers.openai.com/api/docs/guides/file-inputs)
*   [Compaction](https://developers.openai.com/api/docs/guides/compaction)
*   [Counting tokens](https://developers.openai.com/api/docs/guides/token-counting)

### SDKs and CLI

*   [OpenAI SDK](https://developers.openai.com/api/docs/libraries)
*   [OpenAI CLI](https://developers.openai.com/api/docs/libraries/openai-cli)

### Resources

*   [Changelog](https://developers.openai.com/api/docs/changelog)
*   [Deprecations](https://developers.openai.com/api/docs/deprecations)
*   [Supported countries](https://developers.openai.com/api/docs/supported-countries)
*   [OpenAI Crawlers](https://developers.openai.com/api/docs/bots)
*   [Terms and policies](https://openai.com/policies)

### Legacy APIs

*   
Agent Builder
    *   [Overview](https://developers.openai.com/api/docs/guides/agent-builder)
    *   [Migration guide](https://developers.openai.com/api/docs/guides/agent-builder/migrate-from-agent-builder)
    *   [Node reference](https://developers.openai.com/api/docs/guides/node-reference)
    *   [Safety in building agents](https://developers.openai.com/api/docs/guides/agent-builder-safety)

*   
Evals
    *   [Getting started](https://developers.openai.com/api/docs/guides/evaluation-getting-started)
    *   [Working with evals](https://developers.openai.com/api/docs/guides/evals)
    *   [Prompt optimizer](https://developers.openai.com/api/docs/guides/prompt-optimizer)
    *   [External models](https://developers.openai.com/api/docs/guides/external-models)
    *   [Best practices](https://developers.openai.com/api/docs/guides/evaluation-best-practices)
    *   [Graders](https://developers.openai.com/api/docs/guides/graders)

*   
Fine-tuning
    *   [Optimization cycle](https://developers.openai.com/api/docs/guides/model-optimization)
    *   [Supervised fine-tuning](https://developers.openai.com/api/docs/guides/supervised-fine-tuning)
    *   [Vision fine-tuning](https://developers.openai.com/api/docs/guides/vision-fine-tuning)
    *   [Direct preference optimization](https://developers.openai.com/api/docs/guides/direct-preference-optimization)
    *   [Reinforcement fine-tuning](https://developers.openai.com/api/docs/guides/reinforcement-fine-tuning)
    *   [RFT use cases](https://developers.openai.com/api/docs/guides/rft-use-cases)
    *   [Best practices](https://developers.openai.com/api/docs/guides/fine-tuning-best-practices)

*   
Assistants API
    *   [Migration guide](https://developers.openai.com/api/docs/assistants/migration)

*   [Model catalog](https://developers.openai.com/api/docs/models)

### Choose a model

*   [Pricing](https://developers.openai.com/api/docs/pricing)
*   [Model selection](https://developers.openai.com/api/docs/guides/model-selection)

### Text and code

*   [Text generation](https://developers.openai.com/api/docs/guides/text)
*   [Code generation](https://developers.openai.com/api/docs/guides/code-generation)
*   [Structured output](https://developers.openai.com/api/docs/guides/structured-outputs)

### Prompting

*   [Overview](https://developers.openai.com/api/docs/guides/prompting)
*   [Prompt engineering](https://developers.openai.com/api/docs/guides/prompt-engineering)
*   [Citation formatting](https://developers.openai.com/api/docs/guides/citation-formatting)
*   [Migration guide](https://developers.openai.com/api/docs/guides/prompting/migrate-from-prompt-object)
*   [Prompt generation](https://developers.openai.com/api/docs/guides/prompt-generation)
*   [Frontend prompting](https://developers.openai.com/api/docs/guides/frontend-prompt)

### Reasoning

*   [Reasoning models](https://developers.openai.com/api/docs/guides/reasoning)
*   [Reasoning best practices](https://developers.openai.com/api/docs/guides/reasoning-best-practices)

### Images and video

*   
[Images and vision](https://developers.openai.com/api/docs/guides/images-vision)
    *   [Image input cost calculator](https://developers.openai.com/api/docs/guides/image-cost-calculator)

*   [Image generation](https://developers.openai.com/api/docs/guides/image-generation)
*   [Video generation](https://developers.openai.com/api/docs/guides/video-generation)

### Realtime and audio

*   [Audio and speech](https://developers.openai.com/api/docs/guides/audio)
*   [Overview](https://developers.openai.com/api/docs/guides/realtime)
*   [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents)

### Specialized models

*   [Deep research](https://developers.openai.com/api/docs/guides/deep-research)
*   [Embeddings](https://developers.openai.com/api/docs/guides/embeddings)
*   [Moderation](https://developers.openai.com/api/docs/guides/moderation)

*   [Overview](https://developers.openai.com/api/docs/guides/agents)

### Agents SDK

*   [Quickstart](https://developers.openai.com/api/docs/guides/agents/quickstart)
*   [Agent definitions](https://developers.openai.com/api/docs/guides/agents/define-agents)
*   [Models and providers](https://developers.openai.com/api/docs/guides/agents/models)
*   [Running agents](https://developers.openai.com/api/docs/guides/agents/running-agents)
*   [Sandbox agents](https://developers.openai.com/api/docs/guides/agents/sandboxes)
*   [Orchestration](https://developers.openai.com/api/docs/guides/agents/orchestration)
*   [Guardrails](https://developers.openai.com/api/docs/guides/agents/guardrails-approvals)
*   [Results and state](https://developers.openai.com/api/docs/guides/agents/results)
*   [Integrations and observability](https://developers.openai.com/api/docs/guides/agents/integrations-observability)
*   [Evaluate agent workflows](https://developers.openai.com/api/docs/guides/agent-evals)

### ChatKit

*   [Overview](https://developers.openai.com/api/docs/guides/chatkit)
*   [Customize](https://developers.openai.com/api/docs/guides/chatkit-themes)
*   [Widgets](https://developers.openai.com/api/docs/guides/chatkit-widgets)
*   [Actions](https://developers.openai.com/api/docs/guides/chatkit-actions)
*   [Advanced integrations](https://developers.openai.com/api/docs/guides/custom-chatkit)

*   [Overview](https://developers.openai.com/api/docs/guides/tools)
*   [Function calling](https://developers.openai.com/api/docs/guides/function-calling)

### Search and retrieval

*   [Web search](https://developers.openai.com/api/docs/guides/tools-web-search)
*   [File search](https://developers.openai.com/api/docs/guides/tools-file-search)
*   [Retrieval](https://developers.openai.com/api/docs/guides/retrieval)

### Connect tools and data

*   [MCP and Connectors](https://developers.openai.com/api/docs/guides/tools-connectors-mcp)
*   [Secure MCP Tunnel](https://developers.openai.com/api/docs/guides/secure-mcp-tunnels)

### Build tool workflows

*   [Skills](https://developers.openai.com/api/docs/guides/tools-skills)
*   [Tool search](https://developers.openai.com/api/docs/guides/tools-tool-search)
*   [Programmatic tool calling](https://developers.openai.com/api/docs/guides/tools-programmatic-tool-calling)
*   [Async tool calling](https://developers.openai.com/api/docs/guides/async-tool-calling)

### Computer and code

*   [Shell](https://developers.openai.com/api/docs/guides/tools-shell)
*   [Computer use](https://developers.openai.com/api/docs/guides/tools-computer-use)
*   [Apply Patch](https://developers.openai.com/api/docs/guides/tools-apply-patch)
*   [Local shell](https://developers.openai.com/api/docs/guides/tools-local-shell)
*   [Code interpreter](https://developers.openai.com/api/docs/guides/tools-code-interpreter)

### Media

*   [Image generation](https://developers.openai.com/api/docs/guides/tools-image-generation)

*   [Overview](https://developers.openai.com/api/docs/guides/realtime)

### Get started

*   [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents)
*   [Live translation](https://developers.openai.com/api/docs/guides/realtime-translation)
*   [Realtime prompting guide](https://developers.openai.com/api/docs/guides/realtime-models-prompting)

### Audio

*   [Audio and speech](https://developers.openai.com/api/docs/guides/audio)
*   [Transcription](https://developers.openai.com/api/docs/guides/transcription)
*   [File transcription](https://developers.openai.com/api/docs/guides/speech-to-text)
*   [Realtime transcription](https://developers.openai.com/api/docs/guides/realtime-transcription)
*   [Speech generation](https://developers.openai.com/api/docs/guides/text-to-speech)

### Connection methods

*   [WebRTC](https://developers.openai.com/api/docs/guides/realtime-webrtc)
*   [WebSocket](https://developers.openai.com/api/docs/guides/realtime-websocket)
*   [SIP](https://developers.openai.com/api/docs/guides/realtime-sip)

### Sessions and operations

*   [Managing conversations](https://developers.openai.com/api/docs/guides/realtime-conversations)
*   [Voice activity detection](https://developers.openai.com/api/docs/guides/realtime-vad)
*   [Realtime with tools](https://developers.openai.com/api/docs/guides/realtime-mcp)
*   [Webhooks and server-side controls](https://developers.openai.com/api/docs/guides/realtime-server-controls)
*   [Managing costs](https://developers.openai.com/api/docs/guides/realtime-costs)

### Go live

*   [Production best practices](https://developers.openai.com/api/docs/guides/production-best-practices)
*   [Deployment checklist](https://developers.openai.com/api/docs/guides/deployment-checklist)

### Performance and quality

*   [Latency optimization](https://developers.openai.com/api/docs/guides/latency-optimization)
*   [Predicted Outputs](https://developers.openai.com/api/docs/guides/predicted-outputs)
*   [Fast mode](https://developers.openai.com/api/docs/guides/fast-mode)
*   [Accuracy optimization](https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy)

### Cost and throughput

*   [Cost optimization](https://developers.openai.com/api/docs/guides/cost-optimization)
*   [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching)
*   [Batch](https://developers.openai.com/api/docs/guides/batch)
*   [Flex processing](https://developers.openai.com/api/docs/guides/flex-processing)

### Safety and governance

*   [Safety best practices](https://developers.openai.com/api/docs/guides/safety-best-practices)
*   [Red teaming](https://developers.openai.com/api/docs/guides/red-teaming)
*   
Safety checks
    *   [Safety classifiers](https://developers.openai.com/api/docs/guides/safety-checks)
    *   [Cybersecurity checks](https://developers.openai.com/api/docs/guides/safety-checks/cybersecurity)
    *   [Misalignment monitoring](https://developers.openai.com/api/docs/guides/safety-checks/misalignment-monitoring)

*   [Under-18 guidance](https://developers.openai.com/api/docs/guides/safety-checks/under-18-api-guidance)
*   [CSAM guidance](https://developers.openai.com/api/docs/guides/csam-guidance)
*   [Content provenance](https://developers.openai.com/api/docs/guides/content-provenance)
*   [Your data](https://developers.openai.com/api/docs/guides/your-data)
*   [Permissions](https://developers.openai.com/api/docs/guides/rbac)

### Infrastructure and access

*   
[Terraform provider](https://developers.openai.com/api/docs/guides/terraform)
    *   [Overview](https://developers.openai.com/api/docs/guides/terraform)
    *   [Projects and access](https://developers.openai.com/api/docs/guides/terraform/projects-and-access)
    *   [Service accounts](https://developers.openai.com/api/docs/guides/terraform/service-accounts)
    *   [Rate limits and spend](https://developers.openai.com/api/docs/guides/terraform/rate-limits-and-spend)
    *   [Model, tool, and data controls](https://developers.openai.com/api/docs/guides/terraform/project-controls)
    *   [Import and reconciliation](https://developers.openai.com/api/docs/guides/terraform/import-and-reconcile)

*   [Private Link](https://developers.openai.com/api/docs/guides/private-link)
*   [IP allowlist](https://developers.openai.com/api/docs/guides/ip-allowlist)
*   [Mutual TLS](https://developers.openai.com/api/docs/guides/mutual-tls)
*   
[Workload identity federation](https://developers.openai.com/api/docs/guides/workload-identity-federation)
    *   [Codex setup](https://developers.openai.com/codex/enterprise/workload-identity)
    *   [Federation rules](https://developers.openai.com/api/docs/guides/workload-identity-federation/federation-rules)
    *   [Admin API](https://developers.openai.com/api/docs/guides/workload-identity-federation/admin-api)
    *   [X.509 certificates](https://developers.openai.com/api/docs/guides/workload-identity-federation/x509)
    *   [Kubernetes](https://developers.openai.com/api/docs/guides/workload-identity-federation/kubernetes)
    *   [AWS](https://developers.openai.com/api/docs/guides/workload-identity-federation/aws)
    *   [Microsoft Azure](https://developers.openai.com/api/docs/guides/workload-identity-federation/microsoft-azure)
    *   [Google Cloud](https://developers.openai.com/api/docs/guides/workload-identity-federation/google-cloud)
    *   [Oracle Cloud Infrastructure](https://developers.openai.com/api/docs/guides/workload-identity-federation/oracle-cloud)
    *   [GitHub Actions](https://developers.openai.com/api/docs/guides/workload-identity-federation/github-actions)
    *   [SPIFFE](https://developers.openai.com/api/docs/guides/workload-identity-federation/spiffe)

*   [IP egress ranges](https://developers.openai.com/api/docs/guides/ip-addresses)
*   [Amazon Bedrock](https://developers.openai.com/api/docs/guides/amazon-bedrock)

### Operations

*   [Rate limits](https://developers.openai.com/api/docs/guides/rate-limits)
*   [Spend limits](https://developers.openai.com/api/docs/guides/spend-limits)
*   [Admin APIs](https://developers.openai.com/api/docs/guides/admin-apis)
*   [Error codes](https://developers.openai.com/api/docs/guides/error-codes)

[Docs](https://learn.chatgpt.com/docs)[Use cases](https://learn.chatgpt.com/use-cases)

Docs Docs

 Plugins  Workspace Agents  Commerce  Ads 

Docs Select...

*   [Home](https://developers.openai.com/plugins)
*   [Quickstart](https://developers.openai.com/plugins/quickstart)

### Core concepts

*   [Plugin architecture](https://developers.openai.com/plugins/concepts/plugins)
*   [Skills](https://developers.openai.com/plugins/concepts/skills)
*   [MCP server](https://developers.openai.com/plugins/concepts/mcp-server)

### Plan

*   [Brainstorm use cases](https://developers.openai.com/plugins/plan/use-case)
*   [Define tools](https://developers.openai.com/plugins/plan/tools)

### Build

*   [Build an MCP server](https://developers.openai.com/plugins/build/mcp-server)
*   [Add UI to your MCP server (optional)](https://developers.openai.com/plugins/build/chatgpt-ui)
*   [Authenticate users](https://developers.openai.com/plugins/build/auth)
*   [Build skills](https://developers.openai.com/plugins/build/skills)
*   [Package your plugin](https://developers.openai.com/plugins/build/plugins)
*   [Examples](https://developers.openai.com/plugins/build/examples)

### Test and publish

*   [Connect and test your plugin](https://developers.openai.com/plugins/deploy/connect-chatgpt)
*   [Submit and publish](https://developers.openai.com/plugins/deploy/submission)
*   [Submission error reference](https://developers.openai.com/plugins/deploy/submission-errors)

### Conversion specs

*   [Restaurant reservation spec](https://developers.openai.com/plugins/guides/restaurant-reservation-conversion-spec)
*   [Get Quote spec](https://developers.openai.com/plugins/guides/local-services-request-quote-conversion-spec)
*   [Product checkout spec](https://developers.openai.com/plugins/guides/product-checkout-conversion-spec)

### Guides

*   [UI guidelines](https://developers.openai.com/plugins/concepts/ui-guidelines)
*   [Optimize Metadata](https://developers.openai.com/plugins/guides/optimize-metadata)
*   [Submit a Claude Code plugin](https://developers.openai.com/plugins/guides/submit-claude-plugin)
*   [Security & Privacy](https://developers.openai.com/plugins/guides/security-privacy)
*   [Troubleshooting](https://developers.openai.com/plugins/deploy/troubleshooting)

### Resources

*   [Changelog](https://developers.openai.com/plugins/changelog)
*   [Plugin guidelines](https://developers.openai.com/plugins/app-guidelines)
*   [MCP server review requirements](https://developers.openai.com/plugins/deploy/app-review)
*   [Plugin UI reference](https://developers.openai.com/plugins/reference)
*   [Checkout API reference](https://developers.openai.com/plugins/build/monetization)

*   [Home](https://developers.openai.com/workspace-agents)

### Get started

*   [Trigger workspace agent runs](https://developers.openai.com/workspace-agents/trigger-runs)
*   [Authenticate with Workspace Agent access tokens](https://developers.openai.com/workspace-agents/authentication)

*   [Home](https://developers.openai.com/commerce)

### Guides

*   [Get started](https://developers.openai.com/commerce/guides/get-started)
*   [Best practices](https://developers.openai.com/commerce/guides/best-practices)

### File Upload

*   [Overview](https://developers.openai.com/commerce/specs/file-upload/overview)
*   [Products](https://developers.openai.com/commerce/specs/file-upload/products)

### API

*   [Overview](https://developers.openai.com/commerce/specs/api/overview)
*   [Feeds](https://developers.openai.com/commerce/specs/api/feeds)
*   [Products](https://developers.openai.com/commerce/specs/api/products)
*   [Promotions](https://developers.openai.com/commerce/specs/api/promotions)

*   [Ads Overview](https://developers.openai.com/ads)

### Measurement

*   [Measurement Pixel](https://developers.openai.com/ads/measurement-pixel)
*   [Multiple Pixels (Advanced)](https://developers.openai.com/ads/multiple-pixels)
*   [Image Tag](https://developers.openai.com/ads/image-tag)
*   [Conversions API](https://developers.openai.com/ads/conversions-api)
*   [Supported Events](https://developers.openai.com/ads/supported-events)

### Advertiser API

*   [Overview](https://developers.openai.com/ads/api-overview)
*   [API Partner Setup](https://developers.openai.com/ads/api-partner-setup)
*   [Quickstart](https://developers.openai.com/ads/api-quickstart)
*   [Bulk API](https://developers.openai.com/ads/bulk-api)
*   [Product Feeds](https://developers.openai.com/ads/product-feeds)
*   [Delta Feeds API](https://developers.openai.com/ads/delta-feeds)
*   [Campaign Targeting](https://developers.openai.com/ads/campaign-targeting)
*   [Conversion-Optimized Campaigns](https://developers.openai.com/ads/conversion-optimized-campaigns)
*   [Custom Audiences](https://developers.openai.com/ads/custom-audiences)

### API Reference

*   [Authentication](https://developers.openai.com/ads/api-reference/authentication)
*   [Ad Account](https://developers.openai.com/ads/api-reference/ad-account)
*   [Campaigns](https://developers.openai.com/ads/api-reference/campaigns)
*   [Ad Groups](https://developers.openai.com/ads/api-reference/ad-groups)
*   [Ads](https://developers.openai.com/ads/api-reference/ads)
*   [Insights](https://developers.openai.com/ads/api-reference/insights)
*   [Files](https://developers.openai.com/ads/api-reference/files)
*   [Conversion Setup](https://developers.openai.com/ads/api-reference/conversion-setup)

 Overview  Features  Configuration  Developers  Security  Administration  Use Cases  Resources 

Docs Overview

*   [Home](https://developers.openai.com/codex)

### Get started

*   [Quickstart](https://developers.openai.com/codex/quickstart)
*   [Use ChatGPT](https://developers.openai.com/codex/use-chatgpt)
*   [Get started with Work](https://developers.openai.com/codex/get-started-with-work)
*   [Import from another agent](https://developers.openai.com/codex/import)

### Foundations

*   [Prompting](https://developers.openai.com/codex/prompting)
*   [Personalize ChatGPT](https://developers.openai.com/codex/personalize)
*   [Skills & Plugins](https://developers.openai.com/codex/skills-and-plugins)
*   [Permissions](https://developers.openai.com/codex/permission-modes)

### Explore

*   [What's new](https://developers.openai.com/codex/whats-new)
*   [Models](https://developers.openai.com/codex/models)
*   [Pricing](https://developers.openai.com/codex/pricing)
*   [Glossary](https://developers.openai.com/codex/glossary)

### Available on

*   [ChatGPT desktop app](https://developers.openai.com/codex/app)
*   [Remote](https://developers.openai.com/codex/remote)
*   [ChatGPT on the web](https://developers.openai.com/codex/web)
*   [Codex CLI](https://developers.openai.com/codex/cli)
*   [Codex IDE extension](https://developers.openai.com/codex/ide)
*   [Codex cloud](https://developers.openai.com/codex/cloud)

### Releases

*   [Changelog](https://developers.openai.com/codex/changelog)
*   [Feature Maturity](https://developers.openai.com/codex/feature-maturity)
*   [Open Source](https://developers.openai.com/codex/open-source)

*   [Overview](https://developers.openai.com/codex/features)

### Workflows

*   [Projects and chats](https://developers.openai.com/codex/projects)
*   [Sites](https://developers.openai.com/codex/sites)
*   [Visualizations](https://developers.openai.com/codex/visualizations)
*   [Scheduled tasks](https://developers.openai.com/codex/automations)
*   [Long-running work](https://developers.openai.com/codex/long-running-work)
*   [Notifications](https://developers.openai.com/codex/notifications)
*   [Pets](https://developers.openai.com/codex/pets)
*   [Codex Micro](https://developers.openai.com/codex/features/codex-micro)

### Capabilities

*   [Browser](https://developers.openai.com/codex/browser)
*   [Computer use](https://developers.openai.com/codex/computer-use)
*   [Voice](https://developers.openai.com/codex/features/voice)
*   [Plugins](https://developers.openai.com/codex/plugins)
*   [Web search](https://developers.openai.com/codex/web-search)
*   [Image generation](https://developers.openai.com/codex/image-generation)
*   [Image inputs](https://developers.openai.com/codex/image-inputs)
*   [Appshots](https://developers.openai.com/codex/appshots)
*   [Browser extension](https://developers.openai.com/codex/chrome-extension)
*   [Work with files](https://developers.openai.com/codex/artifacts-viewer)

### Reference

*   [Commands](https://developers.openai.com/codex/reference/commands)
*   [Slash commands](https://developers.openai.com/codex/reference/slash-commands)
*   [Settings](https://developers.openai.com/codex/reference/settings)
*   [Troubleshooting](https://developers.openai.com/codex/reference/troubleshooting)

*   [Overview](https://developers.openai.com/codex/configuration)

### Customization

*   [Overview](https://developers.openai.com/codex/customization/overview)
*   [Memories](https://developers.openai.com/codex/customization/memories)
*   [Computer History](https://developers.openai.com/codex/customization/computer-history)

### Config file

*   [Config Basics](https://developers.openai.com/codex/config-file/config-basic)
*   [Advanced Config](https://developers.openai.com/codex/config-file/config-advanced)
*   [Config Reference](https://developers.openai.com/codex/config-file/config-reference)
*   [Environment Variables](https://developers.openai.com/codex/config-file/environment-variables)
*   [Sample Config](https://developers.openai.com/codex/config-file/config-sample)

### Agent configuration

*   [AGENTS.md](https://developers.openai.com/codex/agent-configuration/agents-md)
*   [Subagents](https://developers.openai.com/codex/agent-configuration/subagents)
*   [Speed](https://developers.openai.com/codex/agent-configuration/speed)
*   [Rules](https://developers.openai.com/codex/agent-configuration/rules)

### Extend ChatGPT and Codex

*   [Record & Replay](https://developers.openai.com/codex/extend/record-and-replay)
*   [MCP](https://developers.openai.com/codex/extend/mcp)

### Linux

*   [Desktop app](https://developers.openai.com/codex/linux/linux-app)

### Windows

*   [Desktop app](https://developers.openai.com/codex/windows/windows-app)
*   [Windows sandbox](https://developers.openai.com/codex/windows/windows-sandbox)
*   [WSL](https://developers.openai.com/codex/windows/wsl)

*   [Overview](https://developers.openai.com/codex/developers)

### Development workflows

*   [Code review](https://developers.openai.com/codex/code-review)
*   [Integrated terminal](https://developers.openai.com/codex/integrated-terminal)

### Extend and automate

*   [Build skills](https://developers.openai.com/codex/build-skills)
*   [Build plugins](https://developers.openai.com/codex/build-plugins)
*   [Site tools (WebMCP)](https://developers.openai.com/codex/webmcp)
*   [Hooks](https://developers.openai.com/codex/hooks)

### Environments

*   [Modes](https://developers.openai.com/codex/environments/modes)
*   [Local environments](https://developers.openai.com/codex/environments/local-environment)
*   [Cloud environment](https://developers.openai.com/codex/environments/cloud-environment)
*   [Git worktrees](https://developers.openai.com/codex/environments/git-worktrees)

### Build with Codex

*   [Codex SDK](https://developers.openai.com/codex/codex-sdk)
*   [App Server](https://developers.openai.com/codex/app-server)
*   [MCP Server](https://developers.openai.com/codex/mcp-server)
*   [GitHub Action](https://developers.openai.com/codex/github-action)
*   [Non-interactive mode](https://developers.openai.com/codex/non-interactive-mode)

### Third-party integrations

*   [GitHub](https://developers.openai.com/codex/third-party/github)
*   [GitLab (Beta)](https://developers.openai.com/codex/third-party/gitlab)
*   [Slack](https://developers.openai.com/codex/third-party/slack)
*   [Linear](https://developers.openai.com/codex/third-party/linear)

### Reference

*   [CLI customization](https://developers.openai.com/codex/cli-customization)
*   [Developer commands](https://developers.openai.com/codex/developer-commands)
*   [Developer settings](https://developers.openai.com/codex/developer-settings)

*   [Overview](https://developers.openai.com/codex/security-administration)

### Permissions

*   [Profiles](https://developers.openai.com/codex/permissions)
*   [Sandboxing](https://developers.openai.com/codex/sandboxing)
*   [Auto-review](https://developers.openai.com/codex/sandboxing/auto-review)
*   [Agent approvals & security](https://developers.openai.com/codex/agent-approvals-security)
*   [Internet access](https://developers.openai.com/codex/cloud/internet-access)

### Codex Security

*   [Overview](https://developers.openai.com/codex/security)
*   
Codex Security plugin
    *   [Quickstart](https://developers.openai.com/codex/security/plugin)
    *   [Run a security scan](https://developers.openai.com/codex/security/plugin/scans)
    *   [Run a deep scan](https://developers.openai.com/codex/security/plugin/deep-scans)
    *   [Review code changes](https://developers.openai.com/codex/security/plugin/code-changes)
    *   [Use the Security workbench](https://developers.openai.com/codex/security/plugin/workbench)
    *   [Triage a backlog](https://developers.openai.com/codex/security/plugin/triage-backlog)
    *   [Fix findings](https://developers.openai.com/codex/security/plugin/fix-findings)
    *   [Propose security hardening](https://developers.openai.com/codex/security/plugin/security-hardening)
    *   [Write vulnerability reports](https://developers.openai.com/codex/security/plugin/vulnerability-reports)
    *   [Export and track findings](https://developers.openai.com/codex/security/plugin/export-findings)
    *   [Changelog](https://developers.openai.com/codex/security/plugin/changelog)

*   
Codex Security CLI
    *   [Quickstart](https://developers.openai.com/codex/security/cli)
    *   [Run bulk scans](https://developers.openai.com/codex/security/cli/bulk-scans)
    *   [Run scans in CI](https://developers.openai.com/codex/security/cli/ci)
    *   [GitLab CI/CD](https://developers.openai.com/codex/security/cli/ci/gitlab)
    *   [Reference](https://developers.openai.com/codex/security/cli/reference)
    *   [FAQ](https://developers.openai.com/codex/security/cli/faq)

*   [TypeScript SDK](https://developers.openai.com/codex/security/sdk)
*   
Codex Security cloud
    *   [Setup](https://developers.openai.com/codex/security/setup)
    *   [Security Review](https://developers.openai.com/codex/security/security-review)
    *   [Improving the threat model](https://developers.openai.com/codex/security/threat-model)
    *   [FAQ](https://developers.openai.com/codex/security/faq)

### Cyber safety

*   [Models & Trusted Access](https://developers.openai.com/codex/cyber-safety)
*   [Recommended configuration](https://developers.openai.com/codex/cyber-safety/recommended-configuration)

*   [Overview](https://developers.openai.com/codex/administration)

### Getting started

*   [Admin rollout guide](https://developers.openai.com/codex/enterprise/admin-setup)

### ChatGPT Work

*   [ChatGPT Work Overview](https://developers.openai.com/codex/enterprise/chatgpt-work-overview)
*   [ChatGPT Work cloud security](https://developers.openai.com/codex/enterprise/chatgpt-work-cloud-security)
*   [ChatGPT Work local security](https://developers.openai.com/codex/enterprise/chatgpt-work-local-security)
*   [ChatGPT Work admin FAQ](https://developers.openai.com/codex/enterprise/work-admin-faq)
*   [ChatGPT Work: usage and cost](https://developers.openai.com/codex/enterprise/chatgpt-work-usage-and-cost)

### Identity and authentication

*   [Authentication overview](https://developers.openai.com/codex/auth)
*   [Workload identity](https://developers.openai.com/codex/enterprise/workload-identity)
*   [Personal Access Tokens](https://developers.openai.com/codex/enterprise/access-tokens)
*   [Service accounts](https://developers.openai.com/codex/enterprise/service-accounts)

### Workspace access, policy, and models

*   [Groups and provisioning](https://developers.openai.com/codex/enterprise/groups-and-provisioning)
*   [User lifecycle management](https://developers.openai.com/codex/enterprise/user-lifecycle)
*   [Roles and workspace permissions](https://developers.openai.com/codex/enterprise/roles-and-workspace-permissions)
*   [GPTs and Sharing](https://developers.openai.com/codex/enterprise/gpts-and-sharing)
*   [Managed configuration](https://developers.openai.com/codex/enterprise/managed-configuration)
*   [Prisma AIRS](https://developers.openai.com/codex/enterprise/prisma-airs)
*   [HIPAA configuration](https://developers.openai.com/codex/hipaa-configuration)
*   [Workspace model availability](https://developers.openai.com/codex/enterprise/workspace-model-availability)

### Plugin and connector controls

*   [Plugin controls](https://developers.openai.com/codex/enterprise/apps-and-connectors)
*   [Plugin management](https://developers.openai.com/codex/enterprise/plugin-management)
*   [Skill controls](https://developers.openai.com/codex/enterprise/skills)

### Usage, governance, and compliance

*   [Governance](https://developers.openai.com/codex/enterprise/governance)
*   [Admin plugin](https://developers.openai.com/codex/enterprise/admin-plugin)
*   [Workspace analytics](https://developers.openai.com/codex/enterprise/workspace-analytics)
*   [Analytics API](https://developers.openai.com/codex/enterprise/analytics-api)
*   [Compliance API and audit events](https://developers.openai.com/codex/enterprise/compliance-api)

### Deployment and model providers

*   [Manage app updates](https://developers.openai.com/codex/enterprise/manage-app-updates)
*   [Windows app deployment](https://developers.openai.com/codex/enterprise/windows-deployment)
*   [Remote connections](https://developers.openai.com/codex/remote-connections)
*   [Amazon Bedrock](https://developers.openai.com/codex/amazon-bedrock)

*   [Explore use cases](https://developers.openai.com/codex/use-cases)
*   [Collections](https://developers.openai.com/codex/use-cases/collections)

*   [Home](https://developers.openai.com/codex/resources)
*   [Videos](https://developers.openai.com/codex/videos)
*   [Showcase](https://developers.openai.com/showcase)
*   [OpenAI Academy](https://openai.com/academy/)
*   [Online trainings](https://academy.openai.com/home/events)

### Community

*   [Codex Ambassadors](https://developers.openai.com/community/codex-ambassadors)
*   [Codex for Students](https://developers.openai.com/community/students)
*   [Codex for Open Source](https://developers.openai.com/community/codex-for-oss)
*   [Meetups](https://developers.openai.com/community/meetups)

### Blog

*   [Company blog](https://openai.com/news/)
*   [Developer blog](https://developers.openai.com/blog)

*   [Explore use cases](https://developers.openai.com/codex/use-cases)
*   [Collections](https://developers.openai.com/codex/use-cases/collections)

*   [Home](https://developers.openai.com/codex/resources)
*   [Videos](https://developers.openai.com/codex/videos)
*   [Showcase](https://developers.openai.com/showcase)
*   [OpenAI Academy](https://openai.com/academy/)
*   [Online trainings](https://academy.openai.com/home/events)

### Community

*   [Codex Ambassadors](https://developers.openai.com/community/codex-ambassadors)
*   [Codex for Students](https://developers.openai.com/community/students)
*   [Codex for Open Source](https://developers.openai.com/community/codex-for-oss)
*   [Meetups](https://developers.openai.com/community/meetups)

### Blog

*   [Company blog](https://openai.com/news/)
*   [Developer blog](https://developers.openai.com/blog)

[Showcase](https://developers.openai.com/showcase) Blog  Cookbook  Learn  Community 

Docs Blog

*   [All posts](https://developers.openai.com/blog)

### Recent

*   [Architectural visualization with Astra](https://developers.openai.com/blog/architectural-visualization-with-astra)
*   [Building games with Astra](https://developers.openai.com/blog/how-to-build-games-with-astra)
*   [Meet Rosalind Workbench: Empowering every scientist to be their own research team](https://developers.openai.com/blog/rosalind-workbench)
*   [Automating repetitive work at OpenAI with Codex](https://developers.openai.com/blog/automating-repetitive-work-at-openai-with-codex)
*   [Meet the winners of OpenAI Build Week](https://developers.openai.com/blog/build-week-winners)

### Topics

*   [General](https://developers.openai.com/blog/topic/general)
*   [API](https://developers.openai.com/blog/topic/api)
*   [Apps SDK](https://developers.openai.com/blog/topic/apps-sdk)
*   [Audio](https://developers.openai.com/blog/topic/audio)
*   [Codex](https://developers.openai.com/blog/topic/codex)
*   [Life sciences](https://developers.openai.com/blog/topic/life-sciences)

*   [Home](https://developers.openai.com/cookbook)

### Topics

*   [Agents](https://developers.openai.com/cookbook/topic/agents)
*   [Evals](https://developers.openai.com/cookbook/topic/evals)
*   [Multimodal](https://developers.openai.com/cookbook/topic/multimodal)
*   [Text](https://developers.openai.com/cookbook/topic/text)
*   [Guardrails](https://developers.openai.com/cookbook/topic/guardrails)
*   [Optimization](https://developers.openai.com/cookbook/topic/optimization)
*   [ChatGPT](https://developers.openai.com/cookbook/topic/chatgpt)
*   [Codex](https://developers.openai.com/cookbook/topic/codex)
*   [gpt-oss](https://developers.openai.com/cookbook/topic/gpt-oss)

### Contribute

*   [Cookbook on GitHub](https://github.com/openai/openai-cookbook)

*   [Home](https://developers.openai.com/learn)
*   [OpenAI Developers plugin](https://developers.openai.com/learn/developers-codex-plugin)
*   [Docs MCP](https://developers.openai.com/learn/docs-mcp)

### Categories

*   [Demo apps](https://developers.openai.com/learn/code)
*   [Videos](https://developers.openai.com/learn/videos)

### Topics

*   [Agents](https://developers.openai.com/learn/agents)
*   [Audio & Voice](https://developers.openai.com/learn/audio)
*   [Computer Use](https://developers.openai.com/learn/cua)
*   [Codex](https://developers.openai.com/learn/codex)
*   [Evals](https://developers.openai.com/learn/evals)
*   [gpt-oss](https://developers.openai.com/learn/gpt-oss)
*   [Fine-tuning](https://developers.openai.com/learn/fine-tuning)
*   [Image generation](https://developers.openai.com/learn/imagegen)
*   [Scaling](https://developers.openai.com/learn/scaling)
*   [Tools](https://developers.openai.com/learn/tools)
*   [Video generation](https://developers.openai.com/learn/videogen)

*   [Community](https://developers.openai.com/community)

### Programs

*   [Codex Ambassadors](https://developers.openai.com/community/codex-ambassadors)
*   [Codex for Students](https://developers.openai.com/community/students)
*   [Codex for Open Source](https://developers.openai.com/community/codex-for-oss)
*   [OpenAI for Startups](https://openai.com/business/why-openai/startups/)

### Events

*   [Meetups](https://developers.openai.com/community/meetups)

### Spaces

*   [Developer Forum](https://community.openai.com/)
*   [Discord](https://discord.com/invite/openai)
*   [Reddit](https://www.reddit.com/r/OpenAI/)
*   [X](https://x.com/OpenAIDevs)

[API Dashboard](https://platform.openai.com/login)

[Try ChatGPT](https://chatgpt.com/)

*   [All posts](https://developers.openai.com/blog)

### Recent

*   [Architectural visualization with Astra](https://developers.openai.com/blog/architectural-visualization-with-astra)
*   [Building games with Astra](https://developers.openai.com/blog/how-to-build-games-with-astra)
*   [Meet Rosalind Workbench: Empowering every scientist to be their own research team](https://developers.openai.com/blog/rosalind-workbench)
*   [Automating repetitive work at OpenAI with Codex](https://developers.openai.com/blog/automating-repetitive-work-at-openai-with-codex)
*   [Meet the winners of OpenAI Build Week](https://developers.openai.com/blog/build-week-winners)

### Topics

*   [General](https://developers.openai.com/blog/topic/general)
*   [API](https://developers.openai.com/blog/topic/api)
*   [Apps SDK](https://developers.openai.com/blog/topic/apps-sdk)
*   [Audio](https://developers.openai.com/blog/topic/audio)
*   [Codex](https://developers.openai.com/blog/topic/codex)
*   [Life sciences](https://developers.openai.com/blog/topic/life-sciences)

Copy Page

Copy Page

Jan 22, 2026 Codex

# Testing Agent Skills Systematically with Evals

A practical guide to turning agent skills into something you can test, score, and improve over time.

Authors: Dominik Kundel, Gabriel Chua

![Image 2: Testing Agent Skills Systematically with Evals](https://developers.openai.com/images/blog/eval-skills.png)

When you’re iterating on a skill for an agent like Codex, it’s hard to tell whether you’re actually improving it or just changing its behavior. One version feels faster, another seems more reliable, and then a regression slips in: the skill doesn’t trigger, it skips a required step, or it leaves extra files behind.

At its core, a skill is an [organized collection of prompts and instructions](https://developers.openai.com/codex/build-skills) for an LLM. The most reliable way to improve a skill over time is to evaluate it the same way you would [any other prompt for LLM applications](https://platform.openai.com/docs/guides/evaluation-best-practices).

_Evals_ (short for _evaluations_) check whether a model’s output, and the steps it took to produce it, match what you intended. Instead of asking “does this feel better?” (or relying on vibes), evals let you ask concrete questions like:

*   Did the agent invoke the skill?
*   Did it run the expected commands?
*   Did it produce outputs that follow the conventions you care about?

Concretely, an eval is: a prompt → a captured run (trace + artifacts) → a small set of checks → a score you can compare over time.

In practice, evals for agent skills look a lot like lightweight end-to-end tests: you run the agent, record what happened, and score the result against a small set of rules.

This post walks through a clear pattern for doing that with Codex, starting from defining success, then adding deterministic checks and rubric-based grading so improvements (and regressions) are clear.

## **1. Define success before you write the skill**

Before writing the skill itself, write down what “success” means in terms you can actually measure. A useful way to think about this is to split your checks into a few categories:

*   **Outcome goals:** Did the task complete? Does the app run?
*   **Process goals:** Did Codex invoke the skill and follow the tools and steps you intended?
*   **Style goals:** Does the output follow the conventions you asked for?
*   **Efficiency goals:** Did it get there without thrashing (for example, unnecessary commands or excessive token use)?

Keep this list small and focused on must-pass checks. The goal isn’t to encode every preference up front, but to capture the behaviors you care about most.

In this post, for example, the guide evaluates a skill that sets up a demo app. Some checks are concrete. Did it run `npm install`? Did it create `package.json`? The guide pairs those with a structured style rubric to evaluate conventions and layout.

This mix is intentional. You want fast, targeted signals that surface specific regressions early, rather than a single pass/fail verdict at the end.

## **2. Create the skill**

A Codex skill is a directory with a `SKILL.md` file that includes YAML front matter (`name`, `description`), followed by the Markdown instructions that define the skill’s behavior and optional resources and scripts. The name and description matter more than they might seem. They’re the primary signals Codex uses to decide _whether_ to invoke the skill at all, and _when_ to inject the rest of `SKILL.md` into the agent’s context. If these are vague or overloaded, the skill won’t trigger reliably.

The fastest way to get started is to use Codex’s built-in skill creator ([which itself is also a skill](https://github.com/openai/skills/tree/main/skills/.system/skill-creator)). It walks you through:

`$skill-creator`
The creator asks you what the skill does, when it should trigger, and whether it’s instruction-only or script-backed (instruction-only is the default recommendation). To learn more about creating a skill, [check out the documentation](https://developers.openai.com/codex/build-skills#create-a-skill).

### **A sample skill**

This post uses an intentionally minimal example: a skill that sets up a small React demo app in a predictable, repeatable way.

This skill will:

*   Scaffold a project using Vite’s React + TypeScript template
*   Configure Tailwind CSS using the official Vite plugin approach
*   Enforce a minimal, consistent file structure
*   Define a clear “definition of done” so success is straightforward to evaluate

Below is a compact draft you can paste either into:

*   `.codex/skills/setup-demo-app/SKILL.md` (repo-scoped), or
*   `~/.codex/skills/setup-demo-app/SKILL.md` (user-scoped).

```
---
name: setup-demo-app
description: Scaffold a Vite + React + Tailwind demo app with a small, consistent project structure.
---

## When to use this

Use when you need a fresh demo app for quick UI experiments or reproductions.

## What to build

Create a Vite React TypeScript app and configure Tailwind. Keep it minimal.

Project structure after setup:

- src/
  - main.tsx (entry)
  - App.tsx (root UI)
  - components/
    - Header.tsx
    - Card.tsx
  - index.css (Tailwind import)
- index.html
- package.json

Style requirements:

- TypeScript components
- Functional components only
- Tailwind classes for styling (no CSS modules)
- No extra UI libraries

## Steps

1. Scaffold with Vite using the React TS template:
   npm create vite@latest demo-app -- --template react-ts

2. Install dependencies:
   cd demo-app
   npm install

3. Install and configure Tailwind using the Vite plugin.
   - npm install tailwindcss @tailwindcss/vite
   - Add the tailwind plugin to vite.config.ts
   - In src/index.css, replace contents with:
     @import "tailwindcss";

4. Implement the minimal UI:
   - Header: app title and short subtitle
   - Card: reusable card container
   - App: render Header + 2 Cards with placeholder text

## Definition of done

- npm run dev starts successfully
- package.json exists
- src/components/Header.tsx and src/components/Card.tsx exist
```

This sample skill takes an opinionated stance on purpose. Without clear constraints, there’s nothing concrete to evaluate.

## **3. Manually trigger the skill to expose hidden assumptions**

Because skill invocation depends so much on the _name_ and _description_ in `SKILL.md`, the first thing to check is whether the `setup-demo-app` skill triggers when you expect it to.

Early on, explicitly activate the skill, either via the `/skills` slash command or by referencing it with the `$` prefix, in a real repository or a scratch directory, and watch where it breaks. This is where you surface the misses: cases where the skill doesn’t trigger at all, triggers too eagerly, or runs but deviates from the intended steps.

At this stage, you’re not optimizing for speed or polish. You’re looking for hidden assumptions the skill is making, such as:

*   **Triggering assumptions**: Prompts like “set up a quick React demo” that _should_ invoke `setup-demo-app` but don’t, or more generic prompts (“add Tailwind styling”) that unintentionally trigger it.

*   **Environment assumptions**: The skill assumes it’s running in an empty directory, or that `npm` is available and preferred over other package managers.

*   **Execution assumptions**: The agent skips `npm install` because it assumes dependencies are already installed, or configures Tailwind before the Vite project exists.

Once you’re ready to make these runs repeatable, switch to `codex exec`. It’s designed for automation and CI: it streams progress to `stderr` and writes only the final result to `stdout`, which makes runs easier to script, capture, and inspect.

By default, `codex exec` runs in a restricted sandbox. If your task needs to write files, run it with `--full-auto`. As a general rule, especially when automating, use the least permissions needed to get the job done.

A basic manual run might look like:

```
codex exec --full-auto \
  'Use the $setup-demo-app skill to create the project in this directory.'
```

This first hands-on pass is less about validating correctness and more about discovering edge cases. Every manual fix you make here, such as adding a missing `npm install`, correcting the Tailwind setup, or tightening the trigger description, is a candidate for a future eval, so you can lock in the intended behavior before evaluating at scale.

## **4. Use a small, targeted prompt set to catch regressions early**

You don’t need a large benchmark to get value from evals. For a single skill, a small set of 10–20 prompts is enough to surface regressions and confirm improvements early.

Start with a small CSV and grow it over time as you encounter real failures during development or usage. Each row should represent a situation where you care whether the `setup-demo-app` skill _does_ or _does not_ activate, and what success looks like when it does.

For example, an initial `evals/setup-demo-app.prompts.csv` might look like this:

```
id,should_trigger,prompt
test-01,true,"Create a demo app named `devday-demo` using the $setup-demo-app skill"
test-02,true,"Set up a minimal React demo app with Tailwind for quick UI experiments"
test-03,true,"Create a small demo app to showcase the Responses API"
test-04,false,"Add Tailwind styling to my existing React app"
```

Each of these cases is testing something slightly different:

*   **Explicit invocation (`test-01`)**

 This prompt names the skill directly. It ensures that Codex can invoke `setup-demo-app` when asked, and that changes to the skill’s name, description, or instructions don’t break direct usage.

*   **Implicit invocation (`test-02`)**

 This prompt describes _exactly_ the scenario the skill targets, setting up a minimal React + Tailwind demo, without mentioning the skill by name. It tests whether the name and description in `SKILL.md` are strong enough for Codex to select the skill on its own.

*   **Contextual invocation (`test-03`)**

 This prompt adds domain context (the Responses API) but still requires the same underlying setup. It checks that the skill triggers in realistic, slightly noisy prompts, and that the resulting app still matches the expected structure and conventions.

*   **Negative control (`test-04`)**

 This prompt should **not** invoke `setup-demo-app`. It’s a common adjacent request (“add Tailwind to an existing app”) that can unintentionally match the skill’s description (“React + Tailwind demo”). Including at least one `should_trigger=false` case helps catch **false positives**, where Codex selects the skill too eagerly and scaffolds a new project when the user wanted an incremental change to an existing one.

This mix is intentional. Some evals should confirm that the skill behaves correctly when invoked explicitly; others should check that it activates in real-world prompts where the user never mentions the skill at all.

As you discover misses, prompts that fail to trigger the skill, or cases where the output drifts from your expectations, add them as new rows. Over time, this small CSV becomes a living record of the scenarios the `setup-demo-app` skill must continue to get right.

Over time, this small dataset becomes a living record of what the skill must continue to get right.

## **5. Get started with lightweight deterministic graders**

This is the core of the evaluation step: use `codex exec --json` so your eval harness can score _what actually happened_, not just whether the final output looks right.

When you enable `--json`, `stdout` becomes a JSONL stream of structured events. That makes it straightforward to write deterministic checks tied directly to the behavior you care about, for example:

*   Did it run `npm install`?
*   Did it create `package.json`?
*   Did it invoke the expected commands, in the expected order?

These checks are intentionally lightweight. They give you fast, explainable signals before you add any model-based grading.

### **A minimal Node.js runner**

A “good enough” approach looks like this:

1.   For each prompt, run `codex exec --json --full-auto "<prompt>"`
2.   Save the JSONL trace to disk
3.   Parse the trace and run deterministic checks over the events

```
// evals/run-setup-demo-app-evals.mjs
import { spawnSync } from "node:child_process";
import { readFileSync, writeFileSync, existsSync, mkdirSync } from "node:fs";
import path from "node:path";

function runCodex(prompt, outJsonlPath) {
  const res = spawnSync(
    "codex",
    [
      "exec",
      "--json", // REQUIRED: emit structured events
      "--full-auto", // Allow file system changes
      prompt,
    ],
    { encoding: "utf8" }
  );

  mkdirSync(path.dirname(outJsonlPath), { recursive: true });

  // stdout is JSONL when --json is enabled
  writeFileSync(outJsonlPath, res.stdout, "utf8");

  return { exitCode: res.status ?? 1, stderr: res.stderr };
}

function parseJsonl(jsonlText) {
  return jsonlText
    .split("\n")
    .filter(Boolean)
    .map((line) => JSON.parse(line));
}

// deterministic check: did the agent run `npm install`?
function checkRanNpmInstall(events) {
  return events.some(
    (e) =>
      (e.type === "item.started" || e.type === "item.completed") &&
      e.item?.type === "command_execution" &&
      typeof e.item?.command === "string" &&
      e.item.command.includes("npm install")
  );
}

// deterministic check: did `package.json` get created?
function checkPackageJsonExists(projectDir) {
  return existsSync(path.join(projectDir, "package.json"));
}

// Example single-case run
const projectDir = process.cwd();
const tracePath = path.join(projectDir, "evals", "artifacts", "test-01.jsonl");

const prompt =
  "Create a demo app named demo-app using the $setup-demo-app skill";

runCodex(prompt, tracePath);

const events = parseJsonl(readFileSync(tracePath, "utf8"));

console.log({
  ranNpmInstall: checkRanNpmInstall(events),
  hasPackageJson: checkPackageJsonExists(path.join(projectDir, "demo-app")),
});
```

The value here is that everything is **deterministic and debuggable**.

If a check fails, you can open the JSONL file and see exactly what happened. Every command execution appears as an `item.*` event, in order. That makes regressions straightforward to explain and fix, which is exactly what you want at this stage.

## **6. Conduct qualitative checks with Codex and rubric-based grading**

Deterministic checks answer _“did it do the basics?”_ but they don’t answer _“did it do it the way you wanted?”_

For skills like `setup-demo-app`, many requirements are qualitative: component structure, styling conventions, or whether Tailwind follows the intended configuration. These are hard to capture with basic file existence checks or command counts alone.

A pragmatic solution is to add a second, model-assisted step to your eval pipeline:

1.   Run the setup skill (this writes code to disk)
2.   Run a **read-only style check** against the resulting repository
3.   Require a **structured response** that your harness can score consistently

Codex supports this directly via `--output-schema`, which constrains the final response to a JSON Schema you define.

### **A small rubric schema**

Start by defining a small schema that captures the checks you care about. For example, create `evals/style-rubric.schema.json`:

```
{
  "type": "object",
  "properties": {
    "overall_pass": { "type": "boolean" },
    "score": { "type": "integer", "minimum": 0, "maximum": 100 },
    "checks": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "id": { "type": "string" },
          "pass": { "type": "boolean" },
          "notes": { "type": "string" }
        },
        "required": ["id", "pass", "notes"],
        "additionalProperties": false
      }
    }
  },
  "required": ["overall_pass", "score", "checks"],
  "additionalProperties": false
}
```

This schema gives you stable fields (`overall_pass`, `score`, per-check results) that you can combine, diff, and track over time.

### **The style-check prompt**

Next, run a second `codex exec` that _only inspects the repository_ and emits a rubric-compliant JSON response:

```
codex exec \
  "Evaluate the demo-app repository against these requirements:
   - Vite + React + TypeScript project exists
   - Tailwind is configured via @tailwindcss/vite and CSS imports tailwindcss
   - src/components contains Header.tsx and Card.tsx
   - Components are functional and styled with Tailwind utility classes (no CSS modules)
   Return a rubric result as JSON with check ids: vite, tailwind, structure, style." \
  --output-schema ./evals/style-rubric.schema.json \
  -o ./evals/artifacts/test-01.style.json
```

This is where `--output-schema` is handy. Instead of free-form text that’s hard to parse or compare, you get a predictable JSON object that your eval harness can score across many runs.

If you later move this eval suite into CI, the Codex GitHub Action explicitly supports passing `--output-schema` through `codex-args`, so you can enforce the same structured output in automated workflows.

## **7. Extending your evals as the skill matures**

Once you have the core loop in place, you can extend your evals in the directions that matter most for your skill. Start small, then layer in deeper checks only where they add real confidence.

Some examples include:

*   **Command count and thrashing:** Count `command_execution` items in the JSONL trace to catch regressions where the agent starts looping or re-running commands. Token usage is also available in `turn.completed` events.

*   **Token budget:** Track `usage.input_tokens` and `usage.output_tokens` to spot accidental prompt bloat and compare efficiency across versions.

*   **Build checks:** Run `npm run build` after the skill completes. This acts as a stronger end-to-end signal and catches broken imports or incorrectly configured tooling.

*   **Runtime smoke checks:** Start `npm run dev` and hit the dev server with `curl`, or run a lightweight Playwright check if you already have one. Use this selectively. It adds confidence but costs time.

*   **Repository cleanliness:** Ensure the run generates no unwanted files and that `git status --porcelain` is empty (or matches an explicit allow list).

*   **Sandbox and permission regressions:** Verify the skill still works without escalating permissions beyond what you intended. Least-privilege defaults matter most once you automate.

The pattern is consistent: begin with fast checks that explain behavior, then add slower, heavier checks only when they reduce risk.

## **8. Key takeaways**

This small `setup-demo-app` example shows the shift from “it feels better” to “proof”: run the agent, record what happened, and grade it with a small set of checks. Once that loop exists, every tweak becomes easier to confirm, and every regression becomes clear. Here are the key takeaways:

*   **Measure what matters.** Good evals make regressions clear and failures explainable.
*   **Start from a checkable definition of done.** Use `$skill-creator` to bootstrap, then tighten the instructions until success is unambiguous.
*   **Ground evals in behavior.** Capture JSONL with `codex exec --json` and write deterministic checks against `command_execution` events.
*   **Use Codex where rules fall short.** Add a structured, rubric-based pass with `--output-schema` to grade style and conventions reliably.
*   **Let real failures drive coverage.** Every manual fix is a signal. Turn it into a test so the skill keeps getting it right.

Ask AI

## Docs agent

Loading docs agent...
