Skills Plugins MCP Prompt Model 博客 我的中心
数据分析与咨询 #data #design #ai #agent

openjudge

Build custom LLM evaluation pipelines using the OpenJudge framework. Covers selecting and configuring graders (LLM-based, function-based, agentic), running batch evaluations with GradingRunner, combining scores with aggregators, applying evaluation strategies (voting, average), auto-generating graders from data, and analyzing results (pairwise win rates, statistics, validation metrics). Use when the user wants to evaluate LLM outputs, compare multiple models, design scoring criteria, or build an automated evaluation system.

DeepseekModel 官方收录技能 质量 优秀 · 90 v1.0.0

获取

https://deepseekmodel.com/api/download.php?id=agentscope-ai-openjudge-skills-openjudge-skill-md&format=skill
下载 .skill 标准格式,含 system_prompt 与 model_config,导入任意 Agent 框架即可使用
.skill 文件中 system_prompt 字段的实际内容。
name openjudge description Build custom LLM evaluation pipelines using the OpenJudge framework. Covers selecting and configuring graders (LLM-based, function-based, agentic), running batch evaluations with GradingRunner, combining scores with aggregators, applying evaluation strategies (voting, average), auto-generating graders from data, and analyzing results (pairwise win rates, statistics, validation metrics). Use when the user wants to evaluate LLM outputs, compare multiple models, design scoring criteria, or build an automated evaluation system. OpenJudge Skill Build evaluation pipelines for LLM applications using the openjudge library. When to Use This Skill User wants to evaluate LLM output quality (correctness, relevance, hallucination, etc.) User wants to compare two or more models and rank them User wants to design a scoring rubric and automate evaluation User wants to analyze evaluation results statistically User wants to build a reward model or quality filter Sub-documents — Read When Relevant Topic File Read when… Grader selection & configuration graders.md User needs to pick or configure an evaluator Batch evaluation pipeline pipeline.md User needs to run evaluation over a dataset Auto-generate graders from data generator.md No rubric yet; generate from labeled examples Analyze & compare results analyzer.md User wants win rates, statistics, or metrics Read the relevant sub-document before writing any code. Install pip install py-openjudge Architecture Overview Dataset (List[dict]) │ ▼ GradingRunner ← orchestrates everything │ ├─► Grader A ──► EvaluationStrategy ──► _aevaluate() ──► GraderScore / GraderRank ├─► Grader B ──► EvaluationStrategy ──► _aevaluate() ──► GraderScore / GraderRank └─► Grader C ... │ ├─► Aggregator (optional) ← combine multiple grader scores into one │ └─► RunnerResult ← {grader_name: [GraderScore, ...]} │ ▼ Analyzer ← statistics, win rates, validation metrics 5-Minute Quick Start Evaluate responses for correctness using a built-in grader: import asyncio from openjudge.models.openai_chat_model import OpenAIChatModel from openjudge.graders.common.correctness import CorrectnessGrader from openjudge.runner.grading_runner import GradingRunner # 1. Configure the judge model (OpenAI-compatible endpoint) model = OpenAIChatModel( model= "qwen-plus" , api_key= "sk-xxx" , base_url= "https://dashscope.aliyuncs.com/compatible-mode/v1" , ) # 2. Instantiate a grader grader = CorrectnessGrader(model=model) # 3. Prepare dataset dataset = [ { "query" : "What is the capital of France?" , "response" : "Paris is the capital of France." , "reference_response" : "Paris." , }, { "query" : "What is 2 + 2?" , "response" : "The answer is five." , "reference_response" : "4." , }, ] # 4. Run evaluation async def main (): runner = GradingRunner( grader_configs={ "correctness" : grader}, max_concurrency= 8 , ) results = await runner.arun(dataset) for i, result in enumerate (results[ "correctness" ]): print ( f"[ {i} ] score= {result.score} reason= {result.reason} " ) asyncio.run(main()) Expected output: [0] score=5 reason=The response accurately states Paris as capital... [1] score=1 reason=The response gives the wrong answer (five vs 4)... Key Data Types Type Description GraderScore Pointwise result: .score (float), .reason (str), .metadata (dict) GraderRank Listwise result: .rank (List[int]), .reason (str), .metadata (dict) GraderError Error during evaluation: .error (str), .reason (str) RunnerResult Dict[str, List[GraderResult]] — keyed by grader name Result Handling Pattern from openjudge.graders.schema import GraderScore, GraderRank, GraderError for grader_name, grader_results in results.items(): for i, result in enumerate (grader_results): if isinstance (result, GraderScore): print ( f" {grader_name} [ {i} ]: score= {result.score} " ) elif isinstance (result, GraderRank): print ( f" {grader_name} [ {i} ]: rank= {result.rank} " ) elif isinstance (result, GraderError): print ( f" {grader_name} [ {i} ]: ERROR — {result.error} " ) Model Configuration All LLM-based graders accept either a BaseChatModel instance or a dict config: # Option A: instance from openjudge.models.openai_chat_model import OpenAIChatModel model = OpenAIChatModel(model= "gpt-4o" , api_key= "sk-..." ) # Option B: dict (auto-creates OpenAIChatModel) model_cfg = { "model" : "gpt-4o" , "api_key" : "sk-..." } grader = CorrectnessGrader(model=model_cfg) # OpenAI-compatible endpoints (DashScope / local / etc.) model = OpenAIChatModel( model= "qwen-plus" , api_key= "sk-xxx" , base_url= "https://dashscope.aliyuncs.com/compatible-mode/v1" , )
Agent 识别该技能的关键词,点击任意一个即可复制。

该技能未提供触发词。

下载的 .skill 包内含以下字段。
字段 说明
format格式标识(skill/v1)
skill_id技能唯一 ID
name技能名称
version版本号
description技能描述
category所属分类(数组)
trigger_words触发词列表
tags标签列表
source来源标识
source_url来源链接(本页地址)
exported_at导出时间(每次下载生成)
system_prompt系统提示词正文
model_config模型参数:provider / model / temperature / max_tokens / top_p
examples示例
install_guide各平台导入说明(Coze / Dify / Claude / 自定义框架)
同一份技能可按不同平台格式导出。
.skill 标准格式,含 system_prompt 与 model_config,导入任意 Agent 框架即可使用 下载
.skillpro 增强格式,额外含脚本 / 工具 / 依赖 / 钩子占位 下载
.json 纯 JSON 导出,只含 system_prompt 与模型参数 下载
Coze 带 frontmatter 的 Markdown,Coze 平台导入用 下载
Dify Dify DSL,创建应用后直接导入 下载

每日精选 Skill 推荐,免费送到你邮箱

输入邮箱,每天接收一个精选 AI Agent 技能推荐。完全免费,持续更新。

验证码 --

提交后我们会发送一封确认邮件,点击邮件里的链接才会开始收信。

完全免费,取消任意时间。我们不会发送垃圾邮件。