Skills Plugins MCP Prompt Model 博客 我的中心
开发编程 #github #ai

data-scraper-agent

构建一个全自动化的AI驱动数据收集代理,适用于任何公共来源——招聘网站、价格信息、新闻、GitHub、体育赛事等任何内容。按计划进行抓取,使用免费LLM(Gemini Flash)丰富数据,将结果存储在Notion/Sheets/Supabase中,并从用户反馈中学习。完全免费在GitHub Actions上运行。适用于用户希望自动监控、收集或跟踪任何公共数据的场景。

DeepseekModel 官方收录技能 质量 优秀 · 90 v1.0.0

获取

https://deepseekmodel.com/api/download.php?id=affaan-m-ecc-docs-zh-cn-skills-data-scraper-agent-skill-md&format=skill
下载 .skill 标准格式,含 system_prompt 与 model_config,导入任意 Agent 框架即可使用
.skill 文件中 system_prompt 字段的实际内容。
name data-scraper-agent description 构建一个全自动化的AI驱动数据收集代理,适用于任何公共来源——招聘网站、价格信息、新闻、GitHub、体育赛事等任何内容。按计划进行抓取,使用免费LLM(Gemini Flash)丰富数据,将结果存储在Notion/Sheets/Supabase中,并从用户反馈中学习。完全免费在GitHub Actions上运行。适用于用户希望自动监控、收集或跟踪任何公共数据的场景。 origin community 数据抓取代理 构建一个生产就绪、AI驱动的数据收集代理,适用于任何公共数据源。 按计划运行,使用免费LLM丰富结果,存储到数据库,并随时间推移不断改进。 技术栈:Python · Gemini Flash (免费) · GitHub Actions (免费) · Notion / Sheets / Supabase 何时激活 用户想要抓取或监控任何公共网站或API 用户说"构建一个检查...的机器人"、"为我监控X"、"从...收集数据" 用户想要跟踪工作、价格、新闻、仓库、体育比分、事件、列表 用户询问如何自动化数据收集而无需支付托管费用 用户想要一个能根据他们的决策随时间推移变得更智能的代理 核心概念 三层架构 每个数据抓取代理都有三层: COLLECT → ENRICH → STORE │ │ │ Scraper AI (LLM) Database runs on scores/ Notion / schedule summarises Sheets / & classifies Supabase 免费技术栈 层级 工具 原因 抓取 requests + BeautifulSoup 无成本,覆盖80%的公共网站 JS渲染的网站 playwright (免费) 当HTML抓取失败时使用 AI丰富 通过REST API的Gemini Flash 500次请求/天,100万令牌/天 — 免费 存储 Notion API 免费层级,用于审查的优秀UI 调度 GitHub Actions cron 对公共仓库免费 学习 仓库中的JSON反馈文件 零基础设施,在git中持久化 AI模型后备链 构建代理以在配额耗尽时自动在Gemini模型间回退: gemini-2.0-flash-lite (30 RPM) → gemini-2.0-flash (15 RPM) → gemini-2.5-flash (10 RPM) → gemini-flash-lite-latest (fallback) 批量API调用以提高效率 切勿为每个项目单独调用LLM。始终批量处理: # BAD: 33 API calls for 33 items for item in items: result = call_ai(item) # 33 calls → hits rate limit # GOOD: 7 API calls for 33 items (batch size 5) for batch in chunks(items, size= 5 ): results = call_ai(batch) # 7 calls → stays within free tier 工作流程 步骤 1: 理解目标 询问用户: 收集什么: "数据源是什么?URL / API / RSS / 公共端点?" 提取什么: "哪些字段重要?标题、价格、URL、日期、分数?" 如何存储: "结果应该存储在哪里?Notion、Google Sheets、Supabase,还是本地文件?" 如何丰富: "您希望AI对每个项目进行评分、总结、分类或匹配吗?" 频率: "应该多久运行一次?每小时、每天、每周?" 常见的提示示例: 招聘网站 → 根据简历评分相关性 产品价格 → 降价时发出警报 GitHub仓库 → 总结新版本 新闻源 → 按主题+情感分类 体育结果 → 提取统计数据到跟踪器 活动日历 → 按兴趣筛选 步骤 2: 设计代理架构 为用户生成以下目录结构: my-agent/ ├── config.yaml # 用户自定义此文件(关键词、过滤器、偏好设置) ├── profile/ │ └── context.md # AI 使用的用户上下文(简历、兴趣、标准) ├── scraper/ │ ├── __init__.py │ ├── main.py # 协调器:抓取 → 丰富 → 存储 │ ├── filters.py # 基于规则的预过滤器(快速,在 AI 处理之前) │ └── sources/ │ ├── __init__.py │ └── source_name.py # 每个数据源一个文件 ├── ai/ │ ├── __init__.py │ ├── client.py # Gemini REST 客户端,带模型回退 │ ├── pipeline.py # 批量 AI 分析 │ ├── jd_fetcher.py # 从 URL 获取完整内容(可选) │ └── memory.py # 从用户反馈中学习 ├── storage/ │ ├── __init__.py │ └── notion_sync.py # 或 sheets_sync.py / supabase_sync.py ├── data/ │ └── feedback.json # 用户决策历史(自动更新) ├── .env.example ├── setup.py # 一次性数据库/模式创建 ├── enrich_existing.py # 对旧行进行 AI 分数回填 ├── requirements.txt └── .github/ └── workflows/ └── scraper.yml # GitHub Actions 计划任务 步骤 3: 构建抓取器源 适用于任何数据源的模板: # scraper/sources/my_source.py """ [Source Name] — scrapes [what] from [where]. Method: [REST API / HTML scraping / RSS feed] """ import requests from bs4 import BeautifulSoup from datetime import datetime, timezone from scraper.filters import is_relevant HEADERS = { "User-Agent" : "Mozilla/5.0 (compatible; research-bot/1.0)" , } def fetch () -> list [ dict ]: """ Returns a list of items with consistent schema. Each item must have at minimum: name, url, date_found. """ results = [] # ---- REST API source ---- resp = requests.get( "https://api.example.com/items" , headers=HEADERS, timeout= 15 ) if resp.status_code == 200 : for item in resp.json().get( "results" , []): if not is_relevant(item.get( "title" , "" )): continue results.append(_normalise(item)) return results def _normalise ( raw: dict ) -> dict : """Convert raw API/HTML data to the standard schema.""" return { "name" : raw.get( "title" , "" ), "url" : raw.get( "link" , "" ), "source" : "MySource" , "date_found" : datetime.now(timezone.utc).date().isoformat(), # add domain-specific fields here } HTML抓取模式: soup = BeautifulSoup(resp.text, "lxml" ) for card in soup.select( "[class*='listing']" ): title = card.select_one( "h2, h3" ).get_text(strip= True ) link = card.select_one( "a" )[ "href" ] if not link.startswith( "http" ): link = f"https://example.com {link} " RSS源模式: import xml.etree.ElementTree as ET root = ET.fromstring(resp.text) for item in root.findall( ".//item" ): title = item.findtext( "title" , "" ) link = item.findtext( "link" , "" ) 步骤 4: 构建Gemini AI客户端 # ai/client.py import os, json, time, requests _last_call = 0.0 MODEL_FALLBACK = [ "gemini-2.0-flash-lite" , "gemini-2.0-flash" , "gemini-2.5-flash" , "gemini-flash-lite-latest" , ] def generate ( prompt: str , model: str = "" , rate_limit: float = 7.0 ) -> dict : """Call Gemini with auto-fallback on 429. Returns parsed JSON or {}.""" global _last_call api_key = os.environ.get( "GEMINI_API_KEY" , "" ) if not api_key: return {} elapsed = time.time() - _last_call if elapsed < rate_limit: time.sleep(rate_limit - elapsed) models = [model] + [m for m in MODEL_FALLBACK if m != model] if model else MODEL_FALLBACK _last_call = time.time() for m in models: url = f"https://generativelanguage.googleapis.com/v1beta/models/ {m} :generateContent?key= {api_key} " payload = { "contents" : [{ "parts" : [{ "text" : prompt}]}], "generationConfig" : { "responseMimeType" : "application/json" , "temperature" : 0.3 , "maxOutputTokens" : 2048 , }, } try : resp = requests.post(url, json=payload, timeout= 30 ) if resp.status_code == 200 : return _parse(resp) if resp.status_code in ( 429 , 404 ): time.sleep( 1 ) continue return {} except requests.RequestException: return {} return {} def _parse ( resp ) -> dict : try : text = ( resp.json() .get( "candidates" , [{}])[ 0 ] .get( "content" , {}) .get( "parts" , [{}])[ 0 ] .get( "text" , "" ) .strip() ) if text.startswith( "```" ): text = text.split( "\n" , 1 )[- 1 ].rsplit( "```" , 1 )[ 0 ] return json.loads(text) except (json.JSONDecodeError, KeyError): return {} 步骤 5: 构建AI管道(批量) # ai/pipeline.py import json import yaml from pathlib import Path from ai.client import generate def analyse_batch ( items: list [ dict ], context: str = "" , preference_prompt: str = "" ) -> list [ dict ]: """Analyse items in batches. Returns items enriched with AI fields.""" config = yaml.safe_load((Path(__file__).parent.parent / "config.yaml" ).read_text()) model = config.get( "ai" , {}).get( "model" , "gemini-2.5-flash" ) rate_limit = config.get( "ai" , {}).get( "rate_limit_seconds" , 7.0 ) min_score = config.get( "ai" , {}).get( "min_score" , 0 ) batch_size = config.get( "ai" , {}).get( "batch_size" , 5 ) batches = [items[i:i + batch_size] for i in range ( 0 , len (items), batch_size)] print ( f" [AI] { len (items)} items → { len (batches)} API calls" ) enriched = [] for i, batch in enumerate (batches): print ( f" [AI] Batch {i + 1 } / { len (batches)} ..." ) prompt = _build_prompt(batch, context, preference_prompt, config) result = generate(prompt, model=model, rate_limit=rate_limit) analyses = result.get( "analyses" , []) for j, item in enumerate (batch): ai = analyses[j] if j < len (analyses) else {} if ai: score = max ( 0 , min ( 100 , int (ai.get( "score" , 0 )))) if min_score and score < min_score: continue enriched.append({**item, "ai_score" : score, "ai_summary" : ai.get( "summary" , "" ), "ai_notes" : ai.get( "notes" , "" )}) else : enriched.append(item) return enriched def _build_prompt ( batch, context, preference_prompt, config ): priorities = config.get( "priorities" , []) items_text = "\n\n" .join( f"Item {i+ 1 } : {json.dumps({k: v for k, v in item.items() if not k.startswith( '_' )} )}" for i, item in enumerate (batch) ) return f"""Analyse these { len (batch)} items and return a JSON object. # Items {items_text} # User Context {context[: 800 ] if context else "Not provided" } # User Priorities { chr ( 10 ).join( f"- {p} " for p in priorities)} {preference_prompt} # Instructions Return: {{"analyses": [{{"score": <0-100>, "summary": "<2 sentences>", "notes": "<why this matches or doesn't>"}} for each item in order]}} Be concise. Score 90+=excellent match, 70-89=good, 50-69=ok, <50=weak.""" 步骤 6: 构建反馈学习系统 # ai/memory.py """Learn from user decisions to improve future scoring.""" import json from pathlib import Path FEEDBACK_PATH = Path(__file__).parent.parent / "data" / "feedback.json" def load_feedback () -> dict : if FEEDBACK_PATH.exists(): try : return json.loads(FEEDBACK_PATH.read_text()) except (json.JSONDecodeError, OSError): pass return { "positive" : [], "negative" : []} def save_feedback ( fb: dict ): FEEDBACK_PATH.parent.mkdir(parents= True , exist_ok= True ) FEEDBACK_PATH.write_text(json.dumps(fb, indent= 2 )) def build_preference_prompt ( feedback: dict , max_examples: int = 15 ) -> str : """Convert feedback history into a prompt bias section.""" lines = [] if feedback.get( "positive" ): lines.append( "# Items the user LIKED (positive signal):" ) for e in feedback[ "positive" ][-max_examples:]: lines.append( f"- {e} " ) if feedback.get( "negative" ): lines.append( "\n# Items the user SKIPPED/REJECTED (negative signal):" ) for e in feedback[ "negative" ][-max_examples:]: lines.append( f"- {e} " ) if lines: lines.append( "\nUse these patterns to bias scoring on new items." ) return "\n" .join(lines)
Agent 识别该技能的关键词,点击任意一个即可复制。

该技能未提供触发词。

下载的 .skill 包内含以下字段。
字段 说明
format格式标识(skill/v1)
skill_id技能唯一 ID
name技能名称
version版本号
description技能描述
category所属分类(数组)
trigger_words触发词列表
tags标签列表
source来源标识
source_url来源链接(本页地址)
exported_at导出时间(每次下载生成)
system_prompt系统提示词正文
model_config模型参数:provider / model / temperature / max_tokens / top_p
examples示例
install_guide各平台导入说明(Coze / Dify / Claude / 自定义框架)
同一份技能可按不同平台格式导出。
.skill 标准格式,含 system_prompt 与 model_config,导入任意 Agent 框架即可使用 下载
.skillpro 增强格式,额外含脚本 / 工具 / 依赖 / 钩子占位 下载
.json 纯 JSON 导出,只含 system_prompt 与模型参数 下载
Coze 带 frontmatter 的 Markdown,Coze 平台导入用 下载
Dify Dify DSL,创建应用后直接导入 下载

每日精选 Skill 推荐,免费送到你邮箱

输入邮箱,每天接收一个精选 AI Agent 技能推荐。完全免费,持续更新。

提交后我们会发送一封确认邮件,点击邮件里的链接才会开始收信。

完全免费,取消任意时间。我们不会发送垃圾邮件。