Skills Plugins MCP Prompt Model 博客 我的中心
开发编程 #ai #agent

agent-eval

Benchmark CodeGraph retrieval quality on a real codebase by comparing agent behavior with vs without CodeGraph. Use when the user runs /agent-eval or asks to test, benchmark, audit, or validate a codegraph version (the local dev build or a published npm version) against a language's repo.

DeepseekModel 官方收录技能 质量 优秀 · 90 v1.0.0

获取

https://deepseekmodel.com/api/download.php?id=colbymchenry-codegraph-claude-skills-agent-eval-skill-md&format=skill
下载 .skill 标准格式,含 system_prompt 与 model_config,导入任意 Agent 框架即可使用
.skill 文件中 system_prompt 字段的实际内容。
name agent-eval description Benchmark CodeGraph retrieval quality on a real codebase by comparing agent behavior with vs without CodeGraph. Use when the user runs /agent-eval or asks to test, benchmark, audit, or validate a codegraph version (the local dev build or a published npm version) against a language's repo. CodeGraph Quality Audit Measures how much CodeGraph helps an agent versus plain grep/read, for a chosen codegraph version on a chosen real-world repo. Drives the harness in scripts/agent-eval/ . Prerequisites tmux 3+, a logged-in claude CLI, node , git (macOS/Linux). Run from the codegraph repo root. Workflow Copy this checklist: - [ ] 1. Pick version (local or npm) - [ ] 2. Pick language - [ ] 3. Pick repo by size - [ ] 4. Pick harness (headless / tmux / both) - [ ] 5. Run audit.sh in the background - [ ] 6. Report results Step 1 — version. Ask with AskUserQuestion : which codegraph version to test. Offer "Local dev build" and "Latest published"; the free-text "Other" lets the user type a specific version (e.g. 0.7.10 ). Map the answer to a VERSION token: "Local dev build" → local "Latest published" → latest a typed version → that string (e.g. 0.7.10 ) Step 2 — language. Read .claude/skills/agent-eval/corpus.json . Ask with AskUserQuestion which language to test, listing the languages that have entries. Step 3 — repo. From the chosen language's entries, ask which repo. Label each option with its size and file count, e.g. excalidraw — Medium (~600 files) . Each entry carries the repo URL and a representative question . Step 4 — harness. Ask with AskUserQuestion which harness to run, and map the answer to a MODE token: "Headless" → headless — claude -p with stream-json: exact tokens/cost and a clean tool sequence (2 runs, fast, no TTY). "Interactive (tmux)" → tmux — drives the real Claude TUI in tmux: faithful Explore-subagent behavior, metrics from session logs (2 runs, slower). "Both" → all — headless + interactive (4 runs). Step 5 — run. Launch in the background (sets the version, clones if missing, wipes + re-indexes, runs the chosen arms — several minutes): scripts/agent-eval/audit.sh <VERSION> <repo-name> <repo-url> "<question>" <MODE> Step 6 — report. When the job finishes, read the log and report per arm: Headless ( parse-run.mjs ): total tool calls, file Read s, Grep/Bash, codegraph-tool calls, duration, total cost . Interactive ( parse-session.mjs ): the VERDICT: codegraph_explore used Nx | Read N | Grep/Bash N and TOKENS: lines. Both paths also print the three feedback metrics — residual context occupancy, explore sufficiency, allocation efficiency — and a headless A/B ends with a side-by-side ARM COMPARISON table. Report that table, and check its contamination row first: CLI calls that RETURNED output > 0 means the arm reached codegraph through Bash and its numbers are void. How to read the rest: docs/benchmarks/agent-eval-feedback-metrics.md . Lead with cost + tool/Read counts — they are the reliable signals; raw token in/out are confounded by subagent delegation and prompt caching. State whether codegraph reduced effort and whether both arms reached a correct answer. Notes The index is rebuilt every run ( audit.sh wipes .codegraph ) — different versions extract differently, so an index must be served by the same binary that built it. audit.sh temporarily mutates the global codegraph install for the test, then restores your dev link via local-install.sh . Corpus repos are cloned to /tmp/codegraph-corpus (reused if already present). Add or edit repos in corpus.json (fields: name , repo , size , files , question ).
Agent 识别该技能的关键词,点击任意一个即可复制。

该技能未提供触发词。

下载的 .skill 包内含以下字段。
字段 说明
format格式标识(skill/v1)
skill_id技能唯一 ID
name技能名称
version版本号
description技能描述
category所属分类(数组)
trigger_words触发词列表
tags标签列表
source来源标识
source_url来源链接(本页地址)
exported_at导出时间(每次下载生成)
system_prompt系统提示词正文
model_config模型参数:provider / model / temperature / max_tokens / top_p
examples示例
install_guide各平台导入说明(Coze / Dify / Claude / 自定义框架)
同一份技能可按不同平台格式导出。
.skill 标准格式,含 system_prompt 与 model_config,导入任意 Agent 框架即可使用 下载
.skillpro 增强格式,额外含脚本 / 工具 / 依赖 / 钩子占位 下载
.json 纯 JSON 导出,只含 system_prompt 与模型参数 下载
Coze 带 frontmatter 的 Markdown,Coze 平台导入用 下载
Dify Dify DSL,创建应用后直接导入 下载

每日精选 Skill 推荐,免费送到你邮箱

输入邮箱,每天接收一个精选 AI Agent 技能推荐。完全免费,持续更新。

验证码 --

提交后我们会发送一封确认邮件,点击邮件里的链接才会开始收信。

完全免费,取消任意时间。我们不会发送垃圾邮件。