Skills Plugins MCP Prompt Model 博客 我的中心
Development #ai #agent

agent-eval

Benchmark CodeGraph retrieval quality on a real codebase by comparing agent behavior with vs without CodeGraph. Use when the user runs /agent-eval or asks to test, benchmark, audit, or validate a codegraph version (the local dev build or a published npm version) against a language's repo.

DeepseekModel Curated skill Quality Excellent · 90 v1.0.0

Get

https://deepseekmodel.com/api/download.php?id=colbymchenry-codegraph-claude-skills-agent-eval-skill-md&format=skill
Download .skill Standard format with system_prompt and model_config, ready for any agent framework
The actual content of the system_prompt field in the .skill file.
name agent-eval description Benchmark CodeGraph retrieval quality on a real codebase by comparing agent behavior with vs without CodeGraph. Use when the user runs /agent-eval or asks to test, benchmark, audit, or validate a codegraph version (the local dev build or a published npm version) against a language's repo. CodeGraph Quality Audit Measures how much CodeGraph helps an agent versus plain grep/read, for a chosen codegraph version on a chosen real-world repo. Drives the harness in scripts/agent-eval/ . Prerequisites tmux 3+, a logged-in claude CLI, node , git (macOS/Linux). Run from the codegraph repo root. Workflow Copy this checklist: - [ ] 1. Pick version (local or npm) - [ ] 2. Pick language - [ ] 3. Pick repo by size - [ ] 4. Pick harness (headless / tmux / both) - [ ] 5. Run audit.sh in the background - [ ] 6. Report results Step 1 — version. Ask with AskUserQuestion : which codegraph version to test. Offer "Local dev build" and "Latest published"; the free-text "Other" lets the user type a specific version (e.g. 0.7.10 ). Map the answer to a VERSION token: "Local dev build" → local "Latest published" → latest a typed version → that string (e.g. 0.7.10 ) Step 2 — language. Read .claude/skills/agent-eval/corpus.json . Ask with AskUserQuestion which language to test, listing the languages that have entries. Step 3 — repo. From the chosen language's entries, ask which repo. Label each option with its size and file count, e.g. excalidraw — Medium (~600 files) . Each entry carries the repo URL and a representative question . Step 4 — harness. Ask with AskUserQuestion which harness to run, and map the answer to a MODE token: "Headless" → headless — claude -p with stream-json: exact tokens/cost and a clean tool sequence (2 runs, fast, no TTY). "Interactive (tmux)" → tmux — drives the real Claude TUI in tmux: faithful Explore-subagent behavior, metrics from session logs (2 runs, slower). "Both" → all — headless + interactive (4 runs). Step 5 — run. Launch in the background (sets the version, clones if missing, wipes + re-indexes, runs the chosen arms — several minutes): scripts/agent-eval/audit.sh <VERSION> <repo-name> <repo-url> "<question>" <MODE> Step 6 — report. When the job finishes, read the log and report per arm: Headless ( parse-run.mjs ): total tool calls, file Read s, Grep/Bash, codegraph-tool calls, duration, total cost . Interactive ( parse-session.mjs ): the VERDICT: codegraph_explore used Nx | Read N | Grep/Bash N and TOKENS: lines. Both paths also print the three feedback metrics — residual context occupancy, explore sufficiency, allocation efficiency — and a headless A/B ends with a side-by-side ARM COMPARISON table. Report that table, and check its contamination row first: CLI calls that RETURNED output > 0 means the arm reached codegraph through Bash and its numbers are void. How to read the rest: docs/benchmarks/agent-eval-feedback-metrics.md . Lead with cost + tool/Read counts — they are the reliable signals; raw token in/out are confounded by subagent delegation and prompt caching. State whether codegraph reduced effort and whether both arms reached a correct answer. Notes The index is rebuilt every run ( audit.sh wipes .codegraph ) — different versions extract differently, so an index must be served by the same binary that built it. audit.sh temporarily mutates the global codegraph install for the test, then restores your dev link via local-install.sh . Corpus repos are cloned to /tmp/codegraph-corpus (reused if already present). Add or edit repos in corpus.json (fields: name , repo , size , files , question ).
Keywords that activate this skill. Click one to copy it.

This skill does not provide trigger words.

The downloaded .skill package contains the following fields.
Field Description
formatFormat tag (skill/v1)
skill_idUnique skill ID
nameSkill name
versionVersion
descriptionDescription
categoryCategories (array)
trigger_wordsTrigger words
tagsTags
sourceSource
source_urlSource URL (this page)
exported_atExported at (set per download)
system_promptSystem prompt body
model_configModel config: provider / model / temperature / max_tokens / top_p
examplesExamples
install_guideImport guide for Coze / Dify / Claude / custom frameworks
The same skill can be exported in different platform formats.
.skill Standard format with system_prompt and model_config, ready for any agent framework Download
.skillpro Enhanced format with scripts, tools, dependencies and hooks Download
.json Plain JSON export with system_prompt and model parameters only Download
Coze Markdown with frontmatter, for Coze platform import Download
Dify Dify DSL, import directly after creating an app Download

每日精选 Skill 推荐,免费送到你邮箱

输入邮箱,每天接收一个精选 AI Agent 技能推荐。完全免费,持续更新。

验证码 --

提交后我们会发送一封确认邮件,点击邮件里的链接才会开始收信。

完全免费,取消任意时间。我们不会发送垃圾邮件。