{
    "format": "skill/v1",
    "skill_id": "nousresearch-hermes-agent-optional-skills-mlops-guidance-skill-md",
    "name": "guidance",
    "version": "1.0.0",
    "description": "Constrain LLM output with grammars; guarantee valid JSON.",
    "category": [
        "生活与工具"
    ],
    "trigger_words": [],
    "tags": [
        "ai"
    ],
    "source": "DeepseekModel",
    "source_url": "https://deepseekmodel.com/skill?id=nousresearch-hermes-agent-optional-skills-mlops-guidance-skill-md",
    "exported_at": "2026-09-17T02:26:27+08:00",
    "system_prompt": "name guidance description Constrain LLM output with grammars; guarantee valid JSON. version 1.0.1 author Orchestra Research license MIT dependencies [\"guidance\",\"transformers\"] platforms [\"linux\",\"macos\",\"windows\"] metadata {\"hermes\":{\"tags\":[\"Prompt Engineering\",\"Guidance\",\"Constrained Generation\",\"Structured Output\",\"JSON Validation\",\"Grammar\",\"Microsoft Research\",\"Format Enforcement\",\"Multi-Step Workflows\"]}} Guidance: Constrained LLM Generation When to Use This Skill Use Guidance when you need to: Control LLM output syntax with regex or grammars Guarantee valid JSON/XML/code generation Reduce latency vs traditional prompting approaches Enforce structured formats (dates, emails, IDs, etc.) Build multi-step workflows with Pythonic control flow Prevent invalid outputs through grammatical constraints GitHub Stars : 18,000+ | From : Microsoft Research Installation # Base installation pip install guidance # With specific backends pip install guidance[transformers] # Hugging Face models pip install guidance[llama_cpp] # llama.cpp models Quick Start Basic Example: Structured Generation from guidance import models, gen # Load model (supports OpenAI, Transformers, llama.cpp) lm = models.OpenAI( \"gpt-4\" ) # Generate with constraints result = lm + \"The capital of France is \" + gen( \"capital\" , max_tokens= 5 ) print (result[ \"capital\" ]) # \"Paris\" Chat format with a local model Constraint support requires local logit access. Regex, select() , and grammar-based constrained generation only work with local backends ( Transformers , LlamaCpp ). Remote API backends ( OpenAI , and Azure variants) support unconstrained gen() / chat only — they cannot enforce token-level constraints. guidance 0.3.x has no models.Anthropic class. from guidance import models, gen, system, user, assistant # Local model (supports constrained generation) lm = models.Transformers( \"microsoft/Phi-4-mini-instruct\" ) # Use context managers for chat format with system(): lm += \"You are a helpful assistant.\" with user(): lm += \"What is the capital of France?\" with assistant(): lm += gen(max_tokens= 20 ) Core Concepts 1. Context Managers Guidance uses Pythonic context managers for chat-style interactions. from guidance import system, user, assistant, gen lm = models.Transformers( \"microsoft/Phi-4-mini-instruct\" ) # System message with system(): lm += \"You are a JSON generation expert.\" # User message with user(): lm += \"Generate a person object with name and age.\" # Assistant response with assistant(): lm += gen( \"response\" , max_tokens= 100 ) print (lm[ \"response\" ]) Benefits: Natural chat flow Clear role separation Easy to read and maintain 2. Constrained Generation Guidance ensures outputs match specified patterns using regex or grammars. Regex Constraints from guidance import models, gen lm = models.Transformers( \"microsoft/Phi-4-mini-instruct\" ) # Constrain to valid email format lm += \"Email: \" + gen( \"email\" , regex= r\"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}\" ) # Constrain to date format (YYYY-MM-DD) lm += \"Date: \" + gen( \"date\" , regex= r\"\\d{4}-\\d{2}-\\d{2}\" ) # Constrain to phone number lm += \"Phone: \" + gen( \"phone\" , regex= r\"\\d{3}-\\d{3}-\\d{4}\" ) print (lm[ \"email\" ]) # Guaranteed valid email print (lm[ \"date\" ]) # Guaranteed YYYY-MM-DD format How it works: Regex converted to grammar at token level Invalid tokens filtered during generation Model can only produce matching outputs Selection Constraints from guidance import models, gen, select lm = models.Transformers( \"microsoft/Phi-4-mini-instruct\" ) # Constrain to specific choices lm += \"Sentiment: \" + select([ \"positive\" , \"negative\" , \"neutral\" ], name= \"sentiment\" ) # Multiple-choice selection lm += \"Best answer: \" + select( [ \"A) Paris\" , \"B) London\" , \"C) Berlin\" , \"D) Madrid\" ], name= \"answer\" ) print (lm[ \"sentiment\" ]) # One of: positive, negative, neutral print (lm[ \"answer\" ]) # One of: A, B, C, or D 3. Token Healing Guidance automatically \"heals\" token boundaries between prompt and generation. Problem: Tokenization creates unnatural boundaries. # Without token healing prompt = \"The capital of France is \" # Last token: \" is \" # First generated token might be \" Par\" (with leading space) # Result: \"The capital of France is Paris\" (double space!) Solution: Guidance backs up one token and regenerates. from guidance import models, gen lm = models.Transformers( \"microsoft/Phi-4-mini-instruct\" ) # Token healing enabled by default lm += \"The capital of France is \" + gen( \"capital\" , max_tokens= 5 ) # Result: \"The capital of France is Paris\" (correct spacing) Benefits: Natural text boundaries No awkward spacing issues Better model performance (sees natural token sequences) 4. Grammar-Based Generation Define complex structures by composing grammar functions. The template-string grammar= form is not part of current guidance — build grammars from composable functions, or use guidance.json() for JSON. from guidance import models, gen from guidance import json as gen_json from pydantic import BaseModel, Field lm = models.Transformers( \"microsoft/Phi-4-mini-instruct\" ) # JSON via a Pydantic schema (guidance.json compiles the schema to a grammar) class Person ( BaseModel ): name: str = Field(pattern= r\"[A-Za-z ]+\" ) age: int email: str = Field(pattern= r\"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}\" ) lm += gen_json(name= \"person\" , schema=Person) print (lm[ \"person\" ]) # Guaranteed valid JSON matching the schema # Or compose grammar functions directly: grammar = \"name=\" + gen( \"name\" , regex= r\"[A-Za-z ]+\" ) + \" age=\" + gen( \"age\" , regex= r\"[0-9]+\" ) lm += grammar Use cases: Complex structured outputs Nested data structures Programming language syntax Domain-specific languages 5. Guidance Functions Create reusable generation patterns with the @guidance decorator. from guidance import guidance, gen, models @guidance def generate_person ( lm ): \"\"\"Generate a person with name and age.\"\"\" lm += \"Name: \" + gen( \"name\" , max_tokens= 20 , stop= \"\\n\" ) lm += \"\\nAge: \" + gen( \"age\" , regex= r\"[0-9]+\" , max_tokens= 3 ) return lm # Use the function lm = models.Transformers( \"microsoft/Phi-4-mini-instruct\" ) lm = generate_person(lm) print (lm[ \"name\" ]) print (lm[ \"age\" ]) Stateful Functions: @guidance( stateless= False ) def react_agent ( lm, question, tools, max_rounds= 5 ): \"\"\"ReAct agent with tool use.\"\"\" lm += f\"Question: {question} \\n\\n\" for i in range (max_rounds): # Thought lm += f\"Thought {i+ 1 } : \" + gen( \"thought\" , stop= \"\\n\" ) # Action lm += \"\\nAction: \" + select( list (tools.keys()), name= \"action\" ) # Execute tool tool_result = tools[lm[ \"action\" ]]() lm += f\"\\nObservation: {tool_result} \\n\\n\" # Check if done lm += \"Done? \" + select([ \"Yes\" , \"No\" ], name= \"done\" ) if lm[ \"done\" ] == \"Yes\" : break # Final answer lm += \"\\nFinal Answer: \" + gen( \"answer\" , max_tokens= 100 ) return lm Backend Configuration OpenAI (remote — unconstrained only) Remote API backends cannot do constrained generation (regex/select/grammar); use them only for plain chat/ gen() . For constraints, use a local backend. from guidance import models lm = models.OpenAI( model= \"gpt-4o-mini\" , api_key= \"your-api-key\" # Or set OPENAI_API_KEY env var ) Local Models (Transformers) from guidance.models import Transformers lm = Transformers( \"microsoft/Phi-4-mini-instruct\" , device= \"cuda\" # Or \"cpu\" ) Local Models (llama.cpp) from guidance.models import LlamaCpp lm = LlamaCpp( model_path= \"/path/to/model.gguf\" , n_ctx= 4096 , n_gpu_layers= 35 ) Common Patterns Pattern 1: JSON Generation from guidance import models, gen, system, user, assistant lm = models.Transformers( \"microsoft/Phi-4-mini-instruct\" ) with system(): lm += \"You generate valid JSON.\" with user(): lm += \"Generate a user profile with name, age, and email.\" with assistant(): lm += \"\"\"{ \"name\": \"\"\" + gen( \"name\" , regex= r'\"[A-Za-z ]+\"' , max_tokens= 30 ) + \"\"\", \"age\": \"\"\" + gen( \"age\" , regex= r\"[0-9]+\" , max_tokens= 3 ) + \"\"\", \"email\": \"\"\" + gen( \"email\" , regex= r'\"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}\"' , max_tokens= 50 ) + \"\"\" }\"\"\" print (lm) # Valid JSON guaranteed Pattern 2: Classification from guidance import models, gen, select lm = models.Transformers( \"microsoft/Phi-4-mini-instruct\" ) text = \"This product is amazing! I love it.\" lm += f\"Text: {text} \\n\" lm += \"Sentiment: \" + select([ \"positive\" , \"negative\" , \"neutral\" ], name= \"sentiment\" ) lm += \"\\nConfidence: \" + gen( \"confidence\" , regex= r\"[0-9]+\" , max_tokens= 3 ) + \"%\" print ( f\"Sentiment: {lm[ 'sentiment' ]} \" ) print ( f\"Confidence: {lm[ 'confidence' ]} %\" ) Pattern 3: Multi-Step Reasoning from guidance import models, gen, guidance @guidance def chain_of_thought ( lm, question ): \"\"\"Generate answer with step-by-step reasoning.\"\"\" lm += f\"Question: {question} \\n\\n\" # Generate multiple reasoning steps for i in range ( 3 ): lm += f\"Step {i+ 1 } : \" + gen( f\"step_ {i+ 1 } \" , stop= \"\\n\" , max_tokens= 100 ) + \"\\n\" # Final answer lm += \"\\nTherefore, the answer is: \" + gen( \"answer\" , max_tokens= 50 ) return lm lm = models.Transformers( \"microsoft/Phi-4-mini-instruct\" ) lm = chain_of_thought(lm, \"What is 15% of 200?\" ) print (lm[ \"answer\" ]) Pattern 4: ReAct Agent from guidance import models, gen, select, guidance @guidance( stateless= False ) def react_agent ( lm, question ): \"\"\"ReAct agent with tool use.\"\"\" tools = { \"calculator\" : lambda expr: eval (expr), \"search\" : lambda query: f\"Search results for: {query} \" , } lm += f\"Question: {question} \\n\\n\" for round in range ( 5 ): # Thought lm += f\"Thought: \" + gen( \"thought\" , stop= \"\\n\" ) + \"\\n\" # Action selection",
    "model_config": {
        "provider": "deepseek",
        "model": "deepseek-chat",
        "temperature": 0.7,
        "max_tokens": 4096,
        "top_p": 0.9
    },
    "examples": [
        {
            "input": "请用guidance帮我处理问题",
            "output": "好的，我是guidance。Constrain LLM output with grammars; guarantee valid JSON. 我会根据你的需求提供专业帮助。"
        },
        {
            "input": "介绍一下你的能力",
            "output": "我是guidance，专注于生活与工具领域。Constrain LLM output with grammars; guarantee valid JSON."
        }
    ],
    "install_guide": {
        "coze": "在 Coze 平台创建 Bot -> 技能配置 -> 导入此 .skill 文件",
        "dify": "在 Dify 平台创建应用 -> 添加知识库 -> 导入此 .skill 配置",
        "claude": "将 system_prompt 字段内容复制到 Claude 自定义指令中",
        "custom": "将此 .skill 文件加载到你的 AI Agent 框架中，解析 system_prompt 和 model_config 即可使用"
    }
}