Skills Plugins MCP Prompt Model 博客 我的中心
Content Creation #data #design #api #web

extract

Crawl an existing website (capped, multi-page) and seed stardust/current/ with PRODUCT.md, DESIGN.md, DESIGN.json, a per-page inventory, and the consolidated brand surface — the captured design system, palette, typography, motifs, and voice of the live site. Use when the user wants to analyze an existing site's design, extract or reverse-engineer its design system or brand, capture design tokens from a live site, import a website as the starting point for a redesign, capture the current state before a migration, or invokes /stardust:extract. Trigger phrases include "analyze this site", "extract the design tokens", "capture the brand", "crawl the site", "reverse engineer the design". Not for scraping page data or content for its own sake (it captures design evidence, not datasets), and not for the redesign itself — extraction is descriptive; direction and prototyping happen downstream.

DeepseekModel Curated skill Quality Excellent · 78 v1.0.0

Get

https://deepseekmodel.com/api/download.php?id=adobe-skills-plugins-stardust-skills-extract-skill-md&format=skill
Download .skill Standard format with system_prompt and model_config, ready for any agent framework
The actual content of the system_prompt field in the .skill file.
name extract description Crawl an existing website (capped, multi-page) and seed stardust/current/ with PRODUCT.md, DESIGN.md, DESIGN.json, a per-page inventory, and the consolidated brand surface — the captured design system, palette, typography, motifs, and voice of the live site. Use when the user wants to analyze an existing site's design, extract or reverse-engineer its design system or brand, capture design tokens from a live site, import a website as the starting point for a redesign, capture the current state before a migration, or invokes /stardust:extract. Trigger phrases include "analyze this site", "extract the design tokens", "capture the brand", "crawl the site", "reverse engineer the design". Not for scraping page data or content for its own sake (it captures design evidence, not datasets), and not for the redesign itself — extraction is descriptive; direction and prototyping happen downstream. license Apache-2.0 stardust:extract Crawl an existing website, parse each page, extract the brand surface, and produce a stardust-formatted snapshot of the current state under stardust/current/ . The output describes what the site is ; later sub-commands consume it to decide what it should be . This skill is descriptive : it does not invent direction, it does not critique, and it does not modify the live site. It writes only under stardust/current/ and updates stardust/state.json . Inputs <url> — required. The origin to crawl. Examples: https://example.com , https://example.com/shop . A path narrows the same-origin crawl to that subtree. --cap <N> — optional. Override the default 5-page cap. The cap is intentionally small — a 5-page sample (home + four IA pillars/templates) is enough for cross-page brand aggregation, system-component detection, and the brand-review HTML; lift it (e.g. --cap 25 ) when a deeper crawl is genuinely needed. --all — optional. Lift the cap entirely; extract every discovered page after junk filtering. Equivalent to --cap 0 . Use when the user spontaneously asks for a full crawl. --pages <slug,slug,...> — optional. Restrict the crawl to specific paths (slugs derived per reference/ia-extraction.md ). Bypasses the cap. --refresh <slug> — optional. Re-extract one page that already exists in state.json . --single — optional. Equivalent to --cap 1 . Useful for testing. --wait <fast|medium|spec|auto> — optional. Wait strategy per page. Default medium . See reference/playwright-recipe.md § Wait modes. --no-junk-filter — optional. Disable the default junk-page filter in discovery (see reference/ia-extraction.md § Filtering). --no-consent-dismiss — optional. Skip the pre-flight consent / cookie banner dismissal (see reference/playwright-recipe.md § Pre-flight: consent dismissal). Use when the redesign scope includes the consent surface or the dismissal's side-effects (script activation that wouldn't otherwise run) must be avoided. Default is to dismiss, keeping screenshots, voice aggregation, and per-section style unpolluted by the banner. --concurrency <n> — optional. Parallel browser contexts for the per-page capture loop. Default 4; sane range 4–8. See § Concurrency. --brand-source <url> — optional, repeatable. An additional same-brand origin whose brand surface enriches the primary extraction (shallow capture: home + up to 2 nav-linked pages). See § Cross-site brand sources. --design-source <url> — optional. Design-donor origin: its design system is captured to stardust/canon-source/ and becomes the fixed redesign target while the primary origin supplies content. See § Cross-site brand sources. --prep — optional. Run in migrate-prep mode : lift the cap, type each page, detect module candidates, capture typed content slots, emit the prep summary. See § Prep mode below. Typically invoked via the prepare-migration orchestrator skill rather than directly. Setup Run the master skill's setup procedure first ( skills/stardust/SKILL.md § Setup): impeccable dep check, context loader, state read. Additional checks for this sub-command: Playwright availability. The extraction step needs a real browser. Detect Playwright in this order: a Playwright MCP server, then a project-importable playwright module. The npx playwright probe is NOT sufficient — it confirms the CLI (which resolves a global install) but the recipe and scripts/crawl.mjs do import { chromium } from 'playwright' , and ESM module resolution does not honour a global install or NODE_PATH — the import throws ERR_MODULE_NOT_FOUND even where npx playwright --version succeeds. Verify the module is import-resolvable from the project root (probe: node -e "import('playwright').then(()=>process.exit(0))" ); if it isn't, run npm i -D playwright --no-save --legacy-peer-deps (or use the Playwright MCP server) before crawling. The --legacy-peer-deps flag is required on aem-boilerplate targets (their pinned eslint@8 makes a plain npm i exit ERESOLVE before playwright is even considered). Don't trust the CLI probe alone. --no-save installs are ephemeral. Any later real npm i (e.g. a setup step adding a devDependency) prunes non-manifest packages, silently removing playwright mid-pipeline. Every downstream skill that renders (prototype, migrate, deploy, diff) must re-run the import-resolvability probe — and re-install on failure — at the start of its own run, not assume extract's install survived. Script location matters. ESM resolves import 'playwright' from the script's directory, and the plugin tree ships no node_modules — so running crawl.mjs from the plugin path throws ERR_MODULE_NOT_FOUND even when the project has playwright installed. Copy the script byte-identical into the project ( stardust/scripts/crawl.mjs ) and run the copy; it resolves against the project's node_modules . Bundled crawler. skills/extract/scripts/crawl.mjs is a runnable reference implementation of this whole sub-command — browser config + bot-management fallback, consent dismissal, wait scroll, the capture list, per-page full-page screenshots ( assets/screenshots/<slug>.png , consumed by the Phase 2.5 vision gate), response validation, and the § Capture-hygiene hardening (visibility filter, interstitial drop, SPA-shell flag, modal textContent capture, tracking-pixel discounting, cross-page duplicate detection). Prefer invoking it ( node skills/extract/scripts/crawl.mjs --url <origin> [--pages …] [--max N] [--concurrency N] ) over hand-rolling a Playwright script per run; extend its in-page capture() to cover any recipe field it doesn't yet emit. Origin collision. If stardust/state.json already records site.originUrl and the new <url> is a different origin, stop and ask before clobbering. Stardust does not silently mix two sites in one project. Browser contexts. Open a fresh BrowserContext per capture worker (§ Concurrency; default 4). Run the consent dismissal pre-flight per reference/playwright-recipe.md § Pre-flight: consent dismissal unless --no-consent-dismiss is set. Cookies persist within a context but not across contexts — re-establish consent state per context (re-run the dismissal on the worker's first page, or clone the probe context's storageState ). Record the resolved method in _crawl-log.json#consent.method . Bot-management probe. When the first navigation in the run returns ERR_HTTP2_PROTOCOL_ERROR or ERR_QUIC_PROTOCOL_ERROR , or hangs through the entire hard-cap on what should be a fast origin, do not retry headless — switch to headed Chrome per § Failure modes → Bot-management block and record the switch in _crawl-log.json#discovery.fetchTechnique so re-runs start in headed mode without rediscovering the issue. Procedure Phase 1 — Discovery Discover the page inventory before crawling. Procedure in reference/ia-extraction.md . In summary: Fetch <origin>/sitemap.xml , then <origin>/sitemap_index.xml , then check robots.txt for Sitemap: directives. If no sitemap is reachable, run a same-origin BFS crawl from <url> , depth-limited to 3, link-extracting from rendered HTML. Filter the discovered URL list: same origin only, exclude mailto: , tel: , anchor-only links, query-only variations, common asset paths ( .css , .js , .pdf , image extensions). De-duplicate trailing-slash variations. Apply the junk-page filter ( reference/ia-extraction.md § Junk-page filter) unless --no-junk-filter is set. Surface the filtered list to the user as overridable. Apply the cap (default 5, or --cap , or --all for no cap) and proceed silently . Print an informational summary of what was kept and what was cut — but do not gate on user confirmation. Users who want different scope set it spontaneously at command time: $stardust extract https://example.com # default 5 pages $stardust extract https://example.com --cap 25 # bump to 25 $stardust extract https://example.com --all # lift the cap $stardust extract https://example.com --pages home,about,pricing $stardust extract https://example.com --single # just the entry URL The agent reads spontaneous scope intent from the user's prompt (e.g. "extract all pages", "look at just the home and pricing", "do a full crawl") and applies the equivalent flag. No re-confirmation needed once intent is clear. Informational output (not a prompt — proceed immediately): Discovered 38 pages on https://example.com (sitemap.xml). Filtered as likely junk (5): /test/, /sample-page/, /holiday1/, ... Selecting 5 highest-priority pages: - / (home) - /about - /pricing - /products - /contact Cut (28 pages, --all to lift): /blog/post-1, /blog/post-2, ... Extracting... Selection heuristic: page-type checklist first, then score-based ranking (home + IA-pillar keywords + sitemap priority − archive / version markers). See reference/ia-extraction.md § Page selection and § Priority for the cap. The English-only keyword list is a known limitation for localized sites. Write the discovered list to stardust/current/_crawl-log.json (created if absent) with _provenance and the full discovery reasoning, including filteredAsJunk[] and userChoice . This is an audit trail, not a state file. Phase 2 — Per-page extraction For each page in the cap-respecting list, render with Playwright following reference/playwright-recipe.md . Captures run concurrently per § Concurrency. The recipe is mandatory per page — in particular, do not skip the wait, scroll, or capture-list steps: Viewport 1440 × 900 @ 2× DPR Wait per the configured wait mode (default medium ; see § Wait modes in reference/playwright-recipe.md ) Disable animations via prefers-reduced-motion: reduce After the wait resolves, scroll to bottom in 4 viewport-height steps with 300 ms pauses, then return to top — this is required to trigger lazy-load and IntersectionObserver-driven content Record waitMs and waitMode in the per-page _provenance Capture per page (full schema in reference/current-state-schema.md ): Page metadata (title, meta description, OG tags, theme-color) Semantic structure: heading outline, landmark roles, sections Hero headline + lede (resolved) — heroHeadline / heroLede picked by font-size × hero-region with a junk/hidden-state filter and a clean meta-description fallback (per reference/playwright-recipe.md § Capture list 5-bis). Required for JS-rendered sites whose document-order headings surface modal / promo / count junk instead of the real tagline. Content: visible text per section (full innerText, no truncation per reference/playwright-recipe.md § Capture list 7), structured paragraphs ( body[] ), lists, FAQ Q/A pairs, and review/testimonial quotes per § Capture list 7-bis. Without these structured fields, every body region under a heading falls back to placeholder signature at migrate time. CTA labels and href targets, link inventory (internal vs external) Per-section computed style summary: dominant colors, font families in use, spacing rhythm, border-radius, shadows Media inventory: img with currentSrc / srcset captured with query strings intact plus a resolves flag (HEAD/GET with browser UA + Referer), intrinsic dimensions, inline SVG count, video/iframe presence, cssBackgrounds[] (including pseudo-element ::before / ::after walks per § Capture list 11) so background-image heroes and motifs do not silently disappear and 404ing CDN images are flagged before migrate ships about:error . Font files captured via network-intercept (per § Capture list 16): every woff2 / woff / ttf / otf response saved under assets/fonts/ and recorded in _brand-extraction.json#type.files[] with licensing flag. Icon-font detection (per § Capture list 17): when the page uses [class^="icon-"] with non-default ::before font-family + codepoint, capture the family, save the file, and record the iconClass → codepoint table in _brand-extraction.json#iconFont . Interactive elements: forms (with field types), buttons, modals detected by ARIA roles Full-page screenshot to assets/screenshots/<slug>.png — script-captured by the bundled crawler after the wait/scroll settle (viewport-only fallback on extremely tall pages; mode in _signals.screenshotMode , relative path in the page JSON screenshot field) Save to stardust/current/pages/<slug>.json with _provenance as the first key. The bundled crawler also saves the settled rendered DOM verbatim as stardust/current/pages/<slug>.html ( page.content() after the wait/scroll settle; path in the record's renderedHtml field). Capture once, parse offline: every downstream importer or sibling generator iterates its extraction against this artifact — free, reproducible, and provenance — instead of re-running live probes per selector guess (recorded: 4+ live round-trips per page family before the switch). Live probes stay for what the static DOM cannot answer: geometry and computed styles. Save referenced media to stardust/current/assets/media/ preserving basename plus a short content hash. Live-render evidence (synthesis is forbidden). Refuse to mark a page extracted in state.json unless its _provenance contains renderedBy: "playwright" , an ISO-8601 fetchedAt , a positive integer waitMs , a waitMode from the recipe, and a final httpStatus in the 2xx/3xx range. These five fields are the contract enforced by reference/current-state-schema.md § Live-render evidence and read back by every downstream phase via validateProvenance() per skills/stardust/reference/state-machine.md § Provenance validation. Synthesizing a page record from _brand-extraction.json plus URL patterns plus captured photos — the 2026-04-30 e-commerce shortcut — is the failure mode this guard exists to prevent. When the agent (or a delegated sub- agent) cannot satisfy the contract for a page, treat the page as a Phase 2 failure: record under _crawl-log.json#crawl.failures[] with errorClass: "ProvenanceMissing" and continue. Mark the page extracted in state.json immediately after each successful page write. If a page fails, record the error in _crawl-log.json and continue — extraction is best-effort per page. Phase 2.5 — Vision verification Before anything downstream is authored, look at each captured page's screenshot ( assets/screenshots/<slug>.png ) — the multimodal model reads the image — and verify it against the extracted record: Does the recorded hero (headline + asset) match what the pixels show? Is the extracted palette plausible against the pixels? Does a cssBackgrounds: [] record look believable, or is imagery visibly present — a silent capture failure? Is the logo captured? Is the page actually rendered — not a consent wall, bot-block page, or blank SPA shell? On mismatch, re-run that page's capture with the escalation ladder before proceeding: bump the wait mode one step ( reference/playwright-recipe.md § Wait modes), then headed Chrome (§ Bot-management fallback), then a fresh browser context. Record the outcome per page in _crawl-log.json#visionCheck[] : { "slug" : "pricing" , "verdict" : "recaptured" , "notes" : "record said zero CSS backgrounds; screenshot shows a full-bleed photo hero" } verdict is "ok" | "recaptured" | "suspect" — suspect means the mismatch survived the ladder; downstream phases treat that record as unreliable. Vision is the authoritative capture check; the heuristic defenses (low-media flag, spaShellSuspect , duplicate hash) remain as cheap early signals but no longer gate alone. Phase 3 — Brand-surface extraction Run after the capture phases (2–2.5). Aggregation may proceed incrementally as concurrent page captures complete (§ Concurrency), but the written file must reflect every extracted page — including brand-source pages per § Cross-site brand sources. Produces stardust/current/_brand-extraction.json per reference/brand-surface.md . Some fields are home-only (logo, voice samples, register heuristic); the visual tokens that drive DESIGN.md (palette, radius, shadow, type) are aggregated across all extracted pages to avoid the home-page bias documented in brand-surface.md § Aggregation scope. Captures: Logo by the v1 priority chain: inline SVG → <img> with logo-ish class/id → apple-touch-icon → og:image → favicon → synthesized placeholder. Save to stardust/current/assets/logo.<ext> . Favicon — ALWAYS captured as its own asset (independent of the logo chain) to stardust/current/assets/favicon.<ext> , per reference/playwright-recipe.md § Favicon capture. Downstream, prototype embeds it in the proposed page head and deploy ships
Keywords that activate this skill. Click one to copy it.

This skill does not provide trigger words.

The downloaded .skill package contains the following fields.
Field Description
formatFormat tag (skill/v1)
skill_idUnique skill ID
nameSkill name
versionVersion
descriptionDescription
categoryCategories (array)
trigger_wordsTrigger words
tagsTags
sourceSource
source_urlSource URL (this page)
exported_atExported at (set per download)
system_promptSystem prompt body
model_configModel config: provider / model / temperature / max_tokens / top_p
examplesExamples
install_guideImport guide for Coze / Dify / Claude / custom frameworks
The same skill can be exported in different platform formats.
.skill Standard format with system_prompt and model_config, ready for any agent framework Download
.skillpro Enhanced format with scripts, tools, dependencies and hooks Download
.json Plain JSON export with system_prompt and model parameters only Download
Coze Markdown with frontmatter, for Coze platform import Download
Dify Dify DSL, import directly after creating an app Download

每日精选 Skill 推荐,免费送到你邮箱

输入邮箱,每天接收一个精选 AI Agent 技能推荐。完全免费,持续更新。

验证码 --

提交后我们会发送一封确认邮件,点击邮件里的链接才会开始收信。

完全免费,取消任意时间。我们不会发送垃圾邮件。