生活与工具
#ai
embedding-strategies
Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains.
DeepseekModel
官方收录技能
质量 优秀 · 90
v1.0.0
获取
https://deepseekmodel.com/api/download.php?id=wshobson-agents-plugins-llm-application-dev-skills-embedding-strategies-skill-md&format=skill
下载 .skill
标准格式,含 system_prompt 与 model_config,导入任意 Agent 框架即可使用
.skill 文件中 system_prompt 字段的实际内容。
name embedding-strategies description Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains. Embedding Strategies Guide to selecting and optimizing embedding models for vector search applications. When to Use This Skill Choosing embedding models for RAG Optimizing chunking strategies Fine-tuning embeddings for domains Comparing embedding model performance Reducing embedding dimensions Handling multilingual content Core Concepts 1. Embedding Model Comparison (2026) Model Dimensions Max Tokens Best For voyage-3-large 1024 32000 Claude apps (Anthropic recommended) voyage-3 1024 32000 Claude apps, cost-effective voyage-code-3 1024 32000 Code search voyage-finance-2 1024 32000 Financial documents voyage-law-2 1024 32000 Legal documents text-embedding-3-large 3072 8191 OpenAI apps, high accuracy text-embedding-3-small 1536 8191 OpenAI apps, cost-effective bge-large-en-v1.5 1024 512 Open source, local deployment all-MiniLM-L6-v2 384 256 Fast, lightweight multilingual-e5-large 1024 512 Multi-language 2. Embedding Pipeline Document → Chunking → Preprocessing → Embedding Model → Vector ↓ [Overlap, Size] [Clean, Normalize] [API/Local] Templates and detailed worked examples Full template library and detailed worked examples live in references/details.md . Read that file when you need the concrete templates. Best Practices Do's Match model to use case : Code vs prose vs multilingual Chunk thoughtfully : Preserve semantic boundaries Normalize embeddings : For cosine similarity search Batch requests : More efficient than one-by-one Cache embeddings : Avoid recomputing for static content Use Voyage AI for Claude apps : Recommended by Anthropic Don'ts Don't ignore token limits : Truncation loses information Don't mix embedding models : Incompatible vector spaces Don't skip preprocessing : Garbage in, garbage out Don't over-chunk : Lose important context Don't forget metadata : Essential for filtering and debugging
Agent 识别该技能的关键词,点击任意一个即可复制。
该技能未提供触发词。
下载的 .skill 包内含以下字段。
| 字段 | 说明 |
|---|---|
| format | 格式标识(skill/v1) |
| skill_id | 技能唯一 ID |
| name | 技能名称 |
| version | 版本号 |
| description | 技能描述 |
| category | 所属分类(数组) |
| trigger_words | 触发词列表 |
| tags | 标签列表 |
| source | 来源标识 |
| source_url | 来源链接(本页地址) |
| exported_at | 导出时间(每次下载生成) |
| system_prompt | 系统提示词正文 |
| model_config | 模型参数:provider / model / temperature / max_tokens / top_p |
| examples | 示例 |
| install_guide | 各平台导入说明(Coze / Dify / Claude / 自定义框架) |