Skills Plugins MCP Prompt Model 博客 我的中心
Lifestyle & Tools #api #ai #agent

agent-computer-use

REQUIRED for any task that involves operating a desktop application — opening apps, clicking buttons, typing into fields, pressing keys, scrolling, dragging, reading what's on screen, moving or resizing windows, or verifying state after an action. Always use the `agent-cu` CLI commands (open, snapshot, click, type, key, find, scroll, drag, batch, wait-for) instead of falling back to `open -a`, AppleScript, `osascript`, `xdotool`, `System Events`, or any other shell workaround — those can't read state back, don't verify, and are fragile across app updates. agent-cu is the canonical computer-use tool for controlling any macOS / Linux / Windows / Electron app via accessibility APIs. Trigger on prompts like "open Music and play X", "search for Y in Maps", "fill out this form", "compute in Calculator", "send a Slack message", "drag this file", "read what's in the current window", or anything where a human would click/type/look at a desktop app.

DeepseekModel Curated skill Quality Excellent · 78 v1.0.0

Get

https://deepseekmodel.com/api/download.php?id=kortix-ai-agent-computer-use-skills-agent-computer-use-skill-md&format=skill
Download .skill Standard format with system_prompt and model_config, ready for any agent framework
The actual content of the system_prompt field in the .skill file.
name agent-computer-use description REQUIRED for any task that involves operating a desktop application — opening apps, clicking buttons, typing into fields, pressing keys, scrolling, dragging, reading what's on screen, moving or resizing windows, or verifying state after an action. Always use the `agent-cu` CLI commands (open, snapshot, click, type, key, find, scroll, drag, batch, wait-for) instead of falling back to `open -a`, AppleScript, `osascript`, `xdotool`, `System Events`, or any other shell workaround — those can't read state back, don't verify, and are fragile across app updates. agent-cu is the canonical computer-use tool for controlling any macOS / Linux / Windows / Electron app via accessibility APIs. Trigger on prompts like "open Music and play X", "search for Y in Maps", "fill out this form", "compute in Calculator", "send a Slack message", "drag this file", "read what's in the current window", or anything where a human would click/type/look at a desktop app. license MIT metadata {"author":"kortix-ai","version":"0.1.2","homepage":"https://github.com/kortix-ai/agent-computer-use"} agent-computer-use — Computer use Skill You have access to agent-cu , a CLI tool that controls desktop applications. You can click buttons, type text, read screens, scroll, drag files, move windows — all from the terminal. First-run setup (user action) The first time you run an agent-cu command, Claude Code prompts the user to approve it. The user has three options, in order of preference: Easiest — run agent-cu setup in a terminal. Interactive wizard that writes Bash(agent-cu *) to the right settings file. One-time. If the user seems frustrated by repeated approval prompts, suggest this. If you (the agent) are running it yourself and can't answer prompts, use the non-interactive flags: agent-cu setup --yes → unsupervised + global (recommended default) agent-cu setup --mode <unsupervised|supervised> --scope <global|project> for specific combos Quickest — pick "Yes, and don't ask again for: agent-cu * " in the approval prompt. Covers only that specific subcommand pattern; you'll get prompted again for other subcommands. Manual — add "Bash(agent-cu *)" to ~/.claude/settings.json under permissions.allow . Do not attempt to modify the user's settings file yourself — direct them to agent-cu setup instead. How to think Think like a human sitting at the computer. Before you act, ask yourself: what would I see on screen? What would I click? What would I type? A human: Looks at the screen (snapshot) Finds what they need (identify refs) Does one action (click, type, key) Checks what changed (re-snapshot) You must do the same. Never skip steps. Never assume the UI didn't change after an action. Core loop snapshot → identify → act → verify agent-cu snapshot -a Music -i -c # what's on screen? # read the output, find the right @ref agent-cu click @e5 # do one thing agent-cu snapshot -a Music -i -c # what changed? Every action changes the UI. Your previous refs are now stale. Always re-snapshot. Opening apps Always wait for the app to be ready before doing anything: agent-cu open Safari -- wait agent-cu snapshot -a Safari -i -c Never interact with an app you haven't opened and snapshotted first. Finding elements Step 1 : Snapshot with -i -c (interactive + compact): agent-cu snapshot -a Calculator -i -c This shows only clickable/typeable elements with refs like @e1 , @e5 , @e12 . Step 2 : Read the output. Find the element you need by its name, role, or id. Step 3 : Use the ref. Refs are the fastest and most reliable way to target elements. If elements are missing, increase depth: agent-cu snapshot -a Safari -i -c -d 8 Clicking For buttons, links, menu items — use click : agent-cu click @e5 # single click (AXPress, headless) agent-cu click @e5 --count 2 # double-click (opens files, plays songs) click tries AXPress first (background, no focus steal). Only falls back to mouse simulation for double-click or right-click. For elements with stable IDs (won't change between snapshots): agent-cu click 'id="play"' -a Music agent-cu click 'id~="track-123"' -a Music # partial id match Typing With a target element (preferred — uses AXSetValue, headless): agent-cu type "hello world" -s @e3 Into the focused field (keyboard simulation, needs app focus): agent-cu type "hello world" -a Safari Always prefer -s @ref when you have a ref. It's more reliable. Key presses agent-cu key Return -a Calculator agent-cu key cmd+k -a Slack agent-cu key cmd+a -a TextEdit agent-cu key Escape -a Safari Scrolling agent-cu scroll down -a Music # scroll main content area agent-cu scroll down --amount 10 -a Music # scroll more agent-cu scroll-to @e42 # scroll element into view (headless) Scroll needs the app to be focused. Use scroll-to for headless. Reading content agent-cu text -a Calculator # all visible text agent-cu get-value @e5 # one element's value/state agent-cu get-value 'id="title"' -a Music # by selector Use get-value on specific elements instead of text on large apps. Window management agent-cu move-window -a Notes --x 100 --y 100 agent-cu resize-window -a Notes --width 800 --height 600 agent-cu windows -a Finder # get window positions and sizes These are instant and headless — use AXSetPosition/AXSetSize. Drag and drop Drag needs the app to be focused and two visible, non-overlapping areas. Think like a human : you need to see both the source and destination. # Step 1: Set up windows side by side agent-cu move-window -a Finder --x 0 --y 25 agent-cu resize-window -a Finder --width 720 --height 475 # (open a second Finder window for destination) # Step 2: Snapshot to find the file agent-cu snapshot -a Finder -i -c -d 8 # Step 3: Get the file's position agent-cu get-value @e32 # check position # Step 4: Drag to destination agent-cu drag @e32 @e50 -a Finder # drag by refs # or by coordinates: agent-cu drag --from-x 300 --from-y 55 --to-x 1000 --to-y 200 -a Finder Selector syntax Refs (always prefer these) @e1, @e2, @e3 # from most recent snapshot DSL 'role=button name="Submit"' # role + exact name 'name="Login"' # exact name 'id="AllClear"' # exact id (most stable) 'id~="track-123"' # id contains (case-insensitive) 'name~="Clear"' # name contains (case-insensitive) 'button "Submit"' # shorthand: role name '"Login"' # shorthand: just name 'role=button index=2' # 3rd match (0-based) 'css=".my-button"' # CSS selector (Electron apps only) Chains 'id=sidebar >> role=button index=0' # first button inside sidebar 'name="Form" >> button "Submit"' # submit button inside form Electron apps (CDP) Electron apps (Slack, Cursor, VS Code, Postman, Discord) are automatically detected. agent-cu relaunches them with CDP support on first use. Everything works headless — no window activation, no mouse, no focus steal: agent-cu snapshot -a Slack -i -c # full DOM tree via CDP agent-cu click @e5 # JS element.click() agent-cu key cmd+k -a Slack # CDP key dispatch agent-cu type "hello" -a Slack # CDP insertText agent-cu scroll down -a Slack # JS scrollBy() agent-cu text -a Slack # document.body.innerText Typing in Electron apps : insertText goes to the focused element. If you need to type into a specific input: agent-cu snapshot -a Slack -i -c # find the input ref agent-cu click @e18 # click to focus it agent-cu key cmd+a -a Slack # select all agent-cu key backspace -a Slack # clear agent-cu type "your text" -a Slack # now type Verification Never assume an action worked. Verify by checking a state-bearing attribute , not just by looking at the tree again. The id vs name distinction (critical) Many apps give a button a fixed id (the slot) and a changing name (the current label). Music's transport button is the canonical example: id is always "play" — it identifies the button as "the transport button", even when currently playing. name flips between "play" and "pause" depending on playback state. To detect state, read name , not id : # check if music is playing agent-cu find 'id="play"' -a Music --compact # → [{"name":"pause", ...}] ← means: playback is ON # → [{"name":"play", ...}] ← means: playback is OFF The same pattern appears in many apps: bookmark/unbookmark, mute/unmute, expand/collapse, follow/unfollow. When you want to confirm a toggle worked, always read the element's current name after the action. Inline verification with --expect agent-cu click @e5 --expect 'name="Dashboard"' # clicks, then polls for an element with name="Dashboard". Fails if it never appears. Reading values agent-cu get-value @e3 # one element's value + role + position agent-cu find 'id="play"' -a Music --compact # most stable if id is known agent-cu snapshot -a Safari -i -c # broad check Idempotent typing agent-cu ensure-text @e3 "hello" # only types if value differs Reading dynamic computed values (e.g., Calculator result) Some apps don't surface the result as a normal value on a labeled element — it's hidden in a staticText node. Use tree and walk for any node with a value : agent-cu tree -a Calculator -d 8 --compact | python3 -c " import json, sys d = json.load(sys.stdin) def walk(n): if n.get('value'): print(n.get('role'), '=', repr(n['value'])) for c in n.get('children', []): walk(c) walk(d) " # → staticText = '1,234×7' # → staticText = '8,638' Locale gotcha: numbers are locale-formatted. Indian locale shows 7^8 = 57,64,801 , international shows 5,764,801 . They're the same value. Before comparing, strip commas and spaces. Waiting When UI takes time to load: agent-cu wait-for 'name="Dashboard"' # poll until element appears agent-cu wait-for 'role=button' -- timeout 15 sleep 2 # simple delay after navigation Batch operations Chain multiple commands to avoid per-command startup: echo '[["click","@e5"],["key","Return","-a","Music"]]' | agent-cu batch --bail Real-world patterns Search and play a song in Music (verified flow) # 1. Open and snapshot agent-cu open Music -- wait agent-cu snapshot -a Music -i -c # → @e1 is the Search sidebar item # 2. Click Search agent-cu click @e1 -a Music # 3. Type into the search field — use role=textField, not a ref (the ref # for the search field changes as the view switches) agent-cu type "Espresso Sabrina Carpenter" -s 'role=textField' -a Music --submit sleep 2 # let search results populate # 4. Pick a result. Grep the snapshot for items matching the track name — # the `id` embeds a stable catalog id, so grab that. agent-cu snapshot -a Music -c | grep -i "espresso" | head -5 # → [@e53] other("axcell") "Espresso" id=Music.shelfItem.TopSearchLockup[id=top-search-section-top-1744253558,...] # 5. Open the album (double-click). Use the full id string, not the ref — # refs can drift between snapshots during long flows. agent-cu click 'id="Music.shelfItem.TopSearchLockup[id=top-search-section-top-1744253558,parentId=top-search-section-top]"' -a Music --count 2 sleep 2 # 6. Find the track row and play. The track has a stable id pattern. agent-cu snapshot -a Music -c | grep "track-lockup" | head -3 # → [@e52] group "Espresso" id=Music.shelfItem.AlbumTrackLockup[...] # 7. Select, then try Return to play. If that fails, fall back to the # transport play button directly. agent-cu click @e52 -a Music agent-cu key Return -a Music sleep 1 # 8. Verify via the transport button's `name` — "pause" means playing. agent-cu find 'id="play"' -a Music --compact # if name="play", playback didn't start. Fallback: agent-cu click 'id="play"' -a Music Key lessons from this flow: Refs ( @e53 ) can go stale between snapshots separated by major UI changes. Prefer full id="..." for cross-snapshot targeting. type -s 'role=textField' --submit combines typing, clearing, and pressing Return reliably. Double-click on a search result often opens the item, not plays it. Drill into the detail view, then trigger playback. Always verify playback via the name of the transport button ( id="play" is the slot; name holds state). If Return doesn't trigger playback, click the transport play button as a fallback. Have a plan B. Compute a multi-step calculation in Calculator (verified flow) # 1. Open and snapshot — Calculator buttons have stable ids (Seven, Multiply, Equals, etc.) agent-cu open Calculator -- wait agent-cu snapshot -a Calculator -i -c # 2. Use batch for the whole keystroke sequence. Avoids 17 per-process starts. # Example: 7^8 = 7×7×7×7×7×7×7×7 = 5,764,801 echo '[ ["click","id=\"AllClear\"","-a","Calculator"], ["click","id=\"Seven\"","-a","Calculator"], ["click","id=\"Multiply\"","-a","Calculator"], ["click","id=\"Seven\"","-a","Calculator"], ["click","id=\"Multiply\"","-a","Calculator"], ["click","id=\"Seven\"","-a","Calculator"], ["click","id=\"Multiply\"","-a","Calculator"],
Keywords that activate this skill. Click one to copy it.

This skill does not provide trigger words.

The downloaded .skill package contains the following fields.
Field Description
formatFormat tag (skill/v1)
skill_idUnique skill ID
nameSkill name
versionVersion
descriptionDescription
categoryCategories (array)
trigger_wordsTrigger words
tagsTags
sourceSource
source_urlSource URL (this page)
exported_atExported at (set per download)
system_promptSystem prompt body
model_configModel config: provider / model / temperature / max_tokens / top_p
examplesExamples
install_guideImport guide for Coze / Dify / Claude / custom frameworks
The same skill can be exported in different platform formats.
.skill Standard format with system_prompt and model_config, ready for any agent framework Download
.skillpro Enhanced format with scripts, tools, dependencies and hooks Download
.json Plain JSON export with system_prompt and model parameters only Download
Coze Markdown with frontmatter, for Coze platform import Download
Dify Dify DSL, import directly after creating an app Download

每日精选 Skill 推荐,免费送到你邮箱

输入邮箱,每天接收一个精选 AI Agent 技能推荐。完全免费,持续更新。

提交后我们会发送一封确认邮件,点击邮件里的链接才会开始收信。

完全免费,取消任意时间。我们不会发送垃圾邮件。