{
    "format": "skill/v1",
    "skill_id": "veniceai-skills-skills-venice-chat-skill-md",
    "name": "venice-chat",
    "version": "1.0.0",
    "description": "Call POST /chat/completions on Venice. Covers the OpenAI-compatible request shape, Venice-only venice_parameters (web search, E2EE, characters, thinking control, X search), multimodal inputs (images/audio/video), tool calls, reasoning controls, streaming, prompt caching, structured output, and model feature suffixes.",
    "category": [
        "内容创作"
    ],
    "trigger_words": [],
    "tags": [
        "image",
        "video",
        "ai",
        "web"
    ],
    "source": "DeepseekModel",
    "source_url": "https://deepseekmodel.com/skill?id=veniceai-skills-skills-venice-chat-skill-md",
    "exported_at": "2026-09-17T14:18:39+08:00",
    "system_prompt": "name venice-chat description Call POST /chat/completions on Venice. Covers the OpenAI-compatible request shape, Venice-only venice_parameters (web search, E2EE, characters, thinking control, X search), multimodal inputs (images/audio/video), tool calls, reasoning controls, streaming, prompt caching, structured output, and model feature suffixes. Venice Chat Completions POST /api/v1/chat/completions is Venice's main text endpoint. It's OpenAI-compatible, plus a venice_parameters object for Venice-only features. Use when You need LLM text generation, with or without tools, with or without streaming. You want multimodal inputs (images, audio, video) to a vision/audio-capable model. You want Venice-specific features: web search, E2EE, characters, xAI X/Twitter search, strip-thinking, web scraping. You need prompt caching for large system prompts or long documents. You need structured ( json_schema ) output. For the newer Alpha Responses API , see venice-responses . Minimal request curl https://api.venice.ai/api/v1/chat/completions \\ -H \"Authorization: Bearer $VENICE_API_KEY \" \\ -H \"Content-Type: application/json\" \\ -d '{ \"model\": \"zai-org-glm-5-1\", \"messages\": [{\"role\": \"user\", \"content\": \"Why is the sky blue?\"}] }' Response shape is the standard OpenAI chat.completion object ( id , object: \"chat.completion\" , choices[].message , usage ). With stream: true , responses come as SSE data: lines in chat.completion.chunk format. The request body Core fields (OpenAI-compatible) Field Notes model string — model ID, trait name, or compatibility mapping. Suffixes allowed (see below). Required. messages array of system / developer / user / assistant / tool messages. Required, min 1. temperature , top_p , top_k , min_p , min_temp , max_temp sampling controls repetition_penalty , frequency_penalty , presence_penalty repetition controls max_tokens (deprecated) / max_completion_tokens upper bound on output tokens n number of choices (keep 1 to minimize cost) seed integer for reproducibility stop / stop_token_ids up to 4 strings, or raw token IDs stream , stream_options.include_usage SSE streaming + include usage in the final chunk response_format {type:\"json_schema\", json_schema:{...}} (preferred), {type:\"json_object\"} , or {type:\"text\"} tools , tool_choice , parallel_tool_calls function calling / built-in tools logprobs , top_logprobs return token log-probabilities reasoning.effort / reasoning_effort none | minimal | low | medium | high | xhigh | max reasoning.summary auto | concise | detailed prompt_cache_key , prompt_cache_retention ( default / extended / 24h ) prompt caching hints. extended and 24h both extend retention to 24 hours on supported models verbosity , text.verbosity low / medium / high / auto . Also accepted as a root-level field, not only nested under text include array of extra fields to include in the response (OpenAI compat) fallbacks up to 10 entries. Anthropic beta parameter for Claude Fable 5 server-side refusal fallback. Forwarded only on direct Anthropic routes and ignored by every other provider metadata key/value strings for tracking user , store accepted but ignored (OpenAI compat) venice_parameters (Venice-only) All optional. Combined with model feature suffixes, these are how you enable Venice features. Field Type Default Effect character_slug string — Apply a published Venice character. Slug is the \"Public ID\" on the character page. See venice-characters . strip_thinking_response bool false Strip <think>...</think> from the assistant output on reasoning models. disable_thinking bool false Disable thinking entirely on supported reasoning models and strip tags. enable_e2ee bool true End-to-end encryption on E2EE-capable models when E2EE headers are present. Set to false to force TEE-only. enable_web_search \"off\" / \"auto\" / \"on\" \"off\" Venice server-side web search. Citations arrive in the first streamed chunk or the response. enable_web_scraping bool false Scrape any URLs found in the last user message (Firecrawl). enable_web_citations bool false Ask the LLM to cite sources with ^1^ / ^1,3^ superscripts. include_search_results_in_stream bool false Experimental — emit search results as the first stream chunk. return_search_results_as_documents bool — Also surface search results as a synthetic tool call venice_web_search_documents (LangChain-friendly). include_venice_system_prompt bool true Prepend Venice's curated system prompt. Turn off for full control. enable_x_search bool false xAI native web + X/Twitter search (Grok models with supportsXSearch ). Adds ~$0.01/search. Model feature suffixes Some venice_parameters can also be expressed as model feature suffixes on the model string — useful when the caller/library (OpenAI SDK, LangChain) can't set venice_parameters . Syntax: <model-id>:<key>=<value>[&<key>=<value>…] Values are URL-decoded. Supported keys (exact match): Key Type Maps to enable_web_search on / off / auto venice_parameters.enable_web_search enable_web_citations \"true\" / \"false\" venice_parameters.enable_web_citations enable_web_scraping \"true\" / \"false\" venice_parameters.enable_web_scraping include_venice_system_prompt \"true\" / \"false\" venice_parameters.include_venice_system_prompt include_search_results_in_stream \"true\" / \"false\" venice_parameters.include_search_results_in_stream return_search_results_as_documents \"true\" / \"false\" venice_parameters.return_search_results_as_documents character_slug string venice_parameters.character_slug strip_thinking_response \"true\" / \"false\" venice_parameters.strip_thinking_response disable_thinking \"true\" / \"false\" venice_parameters.disable_thinking Unknown keys are silently ignored. Examples: zai-org-glm-5-1:enable_web_search=on kimi-k2-6:strip_thinking_response=true&enable_web_search=auto zai-org-glm-5-1:character_slug=alan-watts Note: enable_e2ee and enable_x_search can only be set via venice_parameters , not as suffixes. Messages and modalities messages[].content is either a string or an array of typed parts. Roles: user , assistant , tool , system , developer (reasoning models like o-series / codex). Text + image ( image_url ) { \"model\" : \"zai-org-glm-5-1\" , \"messages\" : [ { \"role\" : \"user\" , \"content\" : [ { \"type\" : \"text\" , \"text\" : \"What's in this image?\" } , { \"type\" : \"image_url\" , \"image_url\" : { \"url\" : \"https://example.com/cat.jpg\" } } ] } ] } url accepts a public URL or data:image/png;base64,... . Models with model_spec.capabilities.supportsMultipleImages: true preserve images across the whole conversation; single-image vision models only keep images from the last user message. Check model_spec.capabilities.maxImages for the per-request cap. Audio input ( input_audio ) { \"role\" : \"user\" , \"content\" : [ { \"type\" : \"text\" , \"text\" : \"Transcribe this clip.\" } , { \"type\" : \"input_audio\" , \"input_audio\" : { \"data\" : \"<base64>\" , \"format\" : \"wav\" } } ] } Formats: wav , mp3 , aiff , aac , ogg , flac , m4a , pcm16 , pcm24 . Audio URLs are not supported — always inline base64. Video input ( video_url ) { \"role\" : \"user\" , \"content\" : [ { \"type\" : \"text\" , \"text\" : \"Summarize this.\" } , { \"type\" : \"video_url\" , \"video_url\" : { \"url\" : \"https://www.youtube.com/watch?v=...\" } } ] } Accepts public URLs (including YouTube for some providers) or data:video/mp4;base64,... . Supported formats: mp4 , mpeg , mov , webm . Prompt caching ( cache_control ) Any text / image_url / input_audio / video_url part can carry: { \"cache_control\" : { \"type\" : \"ephemeral\" , \"ttl\" : \"1h\" } } Combine with prompt_cache_key and prompt_cache_retention: \"24h\" on the root request for predictable cache routing. Cache read / write pricing is model-specific — check model_spec.pricing on /models . Tools & function calling Function tools { \"tools\" : [ { \"type\" : \"function\" , \"function\" : { \"name\" : \"get_weather\" , \"description\" : \"Get current weather for a city\" , \"parameters\" : { \"type\" : \"object\" , \"properties\" : { \"city\" : { \"type\" : \"string\" } } , \"required\" : [ \"city\" ] } , \"strict\" : true } } ] , \"tool_choice\" : \"auto\" } tool_choice can also be \"required\" , \"none\" , or {\"type\":\"function\",\"function\":{\"name\":\"get_weather\"}} . parallel_tool_calls: true (default) lets the model emit multiple calls at once. Respond by appending {\"role\":\"tool\",\"tool_call_id\":\"...\",\"content\":\"...\"} before the next call. Built-in tools \"tools\" : [ { \"type\" : \"web_search\" } , { \"type\" : \"x_search\" } ] Equivalent to toggling venice_parameters.enable_web_search / enable_x_search . x_search requires a model with supportsXSearch . Reasoning models On thinking models (GLM 5.1, Kimi K2.6, Claude Opus 4.7, GPT-5.4 Pro, …): { \"model\" : \"zai-org-glm-5-1\" , \"reasoning\" : { \"effort\" : \"medium\" , \"summary\" : \"auto\" } , \"venice_parameters\" : { \"strip_thinking_response\" : false } , \"messages\" : [ ... ] }",
    "model_config": {
        "provider": "deepseek",
        "model": "deepseek-chat",
        "temperature": 0.7,
        "max_tokens": 4096,
        "top_p": 0.9
    },
    "examples": [
        {
            "input": "请用venice-chat帮我处理问题",
            "output": "好的，我是venice-chat。Call POST /chat/completions on Venice. Covers the OpenAI-compatible request shape, Venice-only venice_parameters (web search, E2EE, characters, thinking control, X search), multimodal inputs (images/audio/video), tool calls, reasoning controls, streaming, prompt caching, structured output, and model feature suffixes. 我会根据你的需求提供专业帮助。"
        },
        {
            "input": "介绍一下你的能力",
            "output": "我是venice-chat，专注于内容创作领域。Call POST /chat/completions on Venice. Covers the OpenAI-compatible request shape, Venice-only venice_parameters (web search, E2EE, characters, thinking control, X search), multimodal inputs (images/audio/video), tool calls, reasoning controls, streaming, prompt caching, structured output, and model feature suffixes."
        }
    ],
    "install_guide": {
        "coze": "在 Coze 平台创建 Bot -> 技能配置 -> 导入此 .skill 文件",
        "dify": "在 Dify 平台创建应用 -> 添加知识库 -> 导入此 .skill 配置",
        "claude": "将 system_prompt 字段内容复制到 Claude 自定义指令中",
        "custom": "将此 .skill 文件加载到你的 AI Agent 框架中，解析 system_prompt 和 model_config 即可使用"
    }
}