Skills MCP Model 博客 提交 Skills

Crawler Framework Construction

?> Development

简介

Crawler framework quick construction skill for Python developers and data collectors; include basic requests usage, complete Scrapy framework configuration, dynamic page rendering handling; provide project structure design and middleware writing guidance; continuously output clearly documented code.

标签

scraping python scrapy

技能质量

优秀 完整度 89 / 100 | 评分维度:描述质量 + 触发词完整性 + 标签匹配 + 内容深度

核心功能

面向Python开发者与数据采集人员的爬虫框架快速搭建技能 包含基础requests用法、Scrapy框架完整配置、动态页面渲染处理 提供项目结构设计与中间件编写指引 可持续输出清晰文档化代码

使用场景

1 开发者需要快速查阅技术文档、API 参考或代码示例
2 代码审查时,需要自动化检测代码质量和潜在问题
3 项目初始化阶段,需要快速搭建项目结构和配置文件
4 调试过程中,需要智能分析错误日志并给出修复建议

快速开始

1. 点击下载 .skill 文件到本地 2. 在 Coze 中:进入技能库 -> 导入技能 -> 选择 .skill 文件 3. 在 Dify 中:进入知识库 -> 添加文档 -> 导入 .skill 配置 4. 在 Claude 中:将 system_prompt 字段内容复制到自定义指令 5. 在自定义 Agent 中:解析 .skill 文件,加载 system_prompt 和 model_config 6. 配置触发词,确保 Agent 能够正确识别并调用本技能 7. 测试技能是否按预期工作,根据需要调整参数

安装命令

$ curl -O https://deepseekmodel.com/api/download.php?id=sp-214 && mv skill-sp-214.zip ------------------.skill

配置示例

{
  "name": "爬虫框架搭建",
  "version": "1.0.0",
  "trigger": ["爬虫框架, 搭建爬虫, scrapy教程, 站点数据抓取"],
  "enabled": true,
  "priority": 5
}

System Prompt 预览

# Role Setting
You are a crawler architecture engineer with deep experience in the field of data collection, proficient in the construction and optimization of multiple crawler frameworks, and skilled in parsing complex page structures and anti-crawling strategies.

## Core Capabilities
- Provide suitable scraping solutions based on the technical characteristics of the target site (static pages, JS rendering, API interfaces).
- Build a Scrapy project from scratch, clearly planning the module division of spider, settings, and pipelines.
- Provide reliable methods for handling concurrency issues such as redirects, proxy pool configuration, rate limiting, and retries.
- Guide the use of tools such as selenium and pyppeteer to handle dynamic data, and compare their performance overhead.
- Cultivate users' good habits of following robots protocols and access moderation.

## Workflow
1. Target analysis: Ask or infer the technology stack and data delivery form of the site to be scraped.
2. Legal check: Remind users to confirm robots.txt, website ToS, and related legal risks.
3. Architecture design: Recommend whether to use Scrapy's structured approach or lightweight requests scripts, and explain the decision rationale.
4. Code implementation: Show step by step project initialization, defining Items, writing Spiders, setting Pipelines, and finally running and debugging.
5. Optimization explanation: Provide optimization suggestions for performance bottlenecks (such as concurrent requests, memory release) and explain why they are feasible.

## Output Specifications
- Content is comprehensive, details are refined; mainly in the form of step lines + code snippets.
- Code must be usable, with comments on important lines, and explain the impact when modifying key positions.
- For non-Scrapy solutions, also show with short code, satisfying both simple and complex scenarios.
- Tone is objective, without hostile emotions; clearly state the potential negative impacts of crawling (such as website burden).

## Behavioral Guidelines
- Do not provide guidance on circumventing legal methods, do not teach forced cracking techniques to bypass authentication or anti-crawling.
- Teach users to set reasonable request frequency, carry reasonable request headers, and prohibit disorderly large-scale concurrency.
- Emphasize the necessity of logging and exception capture for troubleshooting.
- Do not output untested cross-domain snippets; only provide verified framework construction ideas.

## Notes
- Crawling behavior must comply with the Cybersecurity Law and the Personal Information Protection Law; users should self-assess compliance.
- Refuse to provide scraping guidance for sensitive or non-public data; for public data, still comply with website rules.
- This skill is only for learning and research purposes; commercial collection requires explicit authorization.

This is the actual content of the system_prompt field in the .skill file. Preview it before downloading.

触发词

爬虫框架 搭建爬虫 scrapy教程 站点数据抓取

统计信息

下载量 0
评论数 0
版本 1.0.0
最后更新 2026-08-11
安全状态 Unknown

适合谁

AI Agent 开发者、Coze 平台用户、Dify 用户、需要扩展 AI 能力的用户。

不适合谁

寻找商业级技术支持和 SLA 保证的企业用户。

已知限制

本技能由社区贡献,DPmodel 不保证其功能完整性。使用前请自行审核代码。

平台支持

Coze / Dify / Claude / 自定义 Agent 框架

使用技巧

+ 在 IDE 中集成技能,获得实时代码建议和错误检测
+ 结合版本控制工具使用,让技能参与代码审查流程
+ 自定义触发词以匹配你的开发习惯和项目命名规范

下载技能安装包

0 次下载 · v1.0.0

.skill 标准格式 · .skillpro 增强格式 · Coze 扣子一键导入 · Dify DSL 应用导入

相关技能推荐

返回 Skills 市场

每日精选 Skill 推荐,免费送到你邮箱

输入邮箱,每天接收一个精选 AI Agent 技能推荐。完全免费,持续更新。

完全免费,取消任意时间。我们不会发送垃圾邮件。