Skills Plugins MCP Prompt Model 导航 博客 资讯 我的中心
工具与能力 #dsh-plugin#data-usage-analytics#deepseek-harness-plugin#ai-dataset-tracker

datatally

DataTally — usage records for data assets. The first data-asset plugin for DeepSeek Harness. Official domain: datatally.xyz

daamaao @daamaao ⬇ 1 ★ 0 main

安装

dsh plugin --profile web add github:daamaao/datatally
下载安装清单

需要可复现安装时,可在仓库后追加 #commit 固定提交。

DataTally — usage records for data assets. The first data-asset plugin for DeepSeek Harness. Official domain: datatally.xyz

该插件未提供要点说明,请参考仓库 README。

dsh-plugindata-usage-analyticsdeepseek-harness-pluginai-dataset-tracker
  1. 安装并启动 DeepSeek Harness:npx @deepseek-ai/dsh web
  2. 在终端执行上面的安装命令(CLI 会解析插件并核验来源)
  3. 用 dsh plugins list 确认已安装,必要时重启 Harness 生效

插件以当前 dsh 进程的权限运行,安装时可能执行代码。请先通读仓库源码与许可证,确认无破坏性命令与越权访问;本站只做索引,不对第三方插件安全性作担保。

代码仓库github.com/daamaao/datatally
许可证MIT
主要语言main
下载量1
GitHub 星标0
最近推送2026-09-08
收录日期2026-09-19
分类工具与能力

事实信息来自公开插件目录快照(2026-10-01),介绍文案由本站再加工。

以下为插件仓库 README 全文(原始内容,由公开目录抓取整理)。

# DataTally

**Official domain: https://datatally.xyz**

**Help AI find the most valuable data.**
帮助 AI 找到最值钱的数据。

**DataTally records. AI judges.**

*For AI, "valuable" means worth the compute — the cost of using data is time and tokens, not money. DataTally helps AI find data worth using.*

DataTally is a [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) plugin that aggregates public usage signals of data assets — downloads, citations, stars — into a single, verifiable profile.

---

## What it does

Data has no intrinsic properties. A byte count says nothing about value. The only real signal of a data asset's value is **how it has been used** — by whom, how often, and with what results. DataTally collects those usage signals from public sources and presents them in one place:

| Tool | What it does |
|------|-------------|
| `search_assets` | Search data asset profiles by keyword or domain |
| `get_asset_profile` | Get the full usage profile of one asset: per-source metrics, timeline, citations |
| `compare_assets` | Compare two assets across dimensions (within the same signal layer) |

Every metric carries its source and fetch time. Nothing is interpreted. Nothing is ranked. **DataTally records; you judge.**

DataTally **never**: computes value, weighs signals, ranks assets, recommends "which is better", sells data, or fabricates numbers. The record is DataTally's. The interpretation is the user's.

---

## Signal layers

- **Deep signals** — actual use: downloads, citations, forks, commits
- **Shallow signals** — interest only: stars, likes ("bookmarked" ≠ "used")

Layers describe behavior versus interest — **not a strength scale**. Comparisons are only made within the same layer. A star is not a download. Missing provenance is marked explicitly — a single-source asset says so in its profile summary; a missing license surfaces as `null`, never a fabricated default.

---

## Install

Requires DeepSeek Harness installed (`dsh` on your PATH) and Node.js ≥ 22.19.

```bash
dsh plugin --profile web add datatally
dsh web
```

### From a checkout (development)

```bash
git clone https://github.com/DAAMAAO/datatally.git
cd datatally-plugin
npm install
npm run typecheck
npm test
```

Load it into a running dsh for end-to-end testing:

```bash
dsh web --patch ./cordis.patch.yml
# (replace the `name` in cordis.patch.yml with the absolute path to src/index.ts)
```

### Use it

Ask your agent:

> "Find datasets for sentiment analysis and compare the two most used ones."

The agent will call `search_assets`, then `compare_assets`, and answer with sourced usage data. An actual session transcript:

**User**:帮我找情感分析的数据集,然后对比两个使用最多的

**Agent**(调 `search_assets("sentiment")`):

> - imdb (stanfordnlp/imdb) | domain: nlp | deep: downloads 195,669 | shallow: likes 725 | fetched_at: 2026-09-08T14:35:50Z | snapshot: v2
> - glue (nyu-mll/glue) | deep: downloads 826,859 … (共 8 个命中)

**Agent**(调 `compare_assets("nyu-mll/glue", "stanfordnlp/imdb")`):

> - downloads [deep]: 791,429 vs 197,595 | same_source: true | same_access: true
> - forks [deep]: 178 vs 0 | same_source: false | same_access: true
> - likes [shallow]: 741 vs 722 | same_source: true | same_access: true
> note: cross-source metrics are not directly comparable

**Agent** 的最终回答只陈述记录(引用次数、来源、抓取时间),**不下"哪个更好"的结论**——判断留给用户。

---

## Configuration

The plugin ships with defaults; override per row in your profile patch:

| Field | Default | Meaning |
|-------|---------|---------|
| `snapshotPath` | `./data/snapshot_v2.json` | Local snapshot file. Relative paths resolve against the process cwd first, then against the package's bundled `data/` seed. |
| `maxResults` | `20` | Search result cap (the `limit` parameter is clamped to it). |
| `enableCitationFields` | `true` | Include the `citations` field in profiles. |

---

## Snapshot format

DataTally reads usage data from a **local, self-owned** JSON file (version 2). Replace the seed by pointing `snapshotPath` at your own file:

```json
{
  "version": "2",
  "generated_at": "2026-09-06T12:00:00Z",
  "assets": [
    {
      "asset_id": "stanfordnlp/imdb",
      "id_type": "hf",
      "name": "imdb",
      "domain": "nlp",
      "access": "open",
      "license": "other",
      "verification": "public_api",
      "sources": [
        {
          "source": "huggingface",
          "metrics": {
            "downloads": { "value": 191564, "signal_type": "deep" },
            "likes": { "value": 709, "signal_type": "shallow" }
          },
          "fetched_at": "2026-09-08T12:32:11.463Z"
        }
      ],
      "timeline": [],
      "citations": []
    }
  ]
}
```

Schema principles: structured from day one · every field has a source · evidence attributes are facts, not interpretations · snapshots are self-owned · machine-readable first. Unknown fields (including the legacy `asset_class`) are rejected loudly; v1 snapshots are rejected with a migration pointer.

---

## CLI

The same core, in a terminal (the thin-wrapper form over the plugin core):

```bash
datatally profile stanfordnlp/imdb
datatally search sentiment --domain nlp --limit 5
datatally compare HuggingFaceFW/fineweb allenai/c4
datatally refresh                        # re-fetch the four public sources into a new snapshot
datatally refresh --query protein        # domain-focused catalog: any keyword, no code change
datatally refresh --filter task_ids:sentiment-classification --limit 30
# snapshot location: --snapshot  or DATATALLY_SNAPSHOT env
```

`--query ` / `--filter ` are repeatable and replace the shipped default queries; when given, the queried candidates lead the catalog.

---

## Development

```
src/
├── index.ts              # plugin entry: Config + apply + tool registration
├── core/                 # pure query logic (no harness deps) + text rendering
├── tools/                # one defineTool per file: schema + execute + render + UI cards
├── snapshot/             # loader (strict validation, fail-loud), types, AssetId brand
│   └── fetchers/         # four-source refresh pipeline (HF/ModelScope/DataCite/GitHub)
├── schema/               # strict snapshot validator
cli/main.ts               # CLI (same core + refresh)
test/                     # 51 tests: unit + fetchers + pipeline + keyless drive + goldens
data/snapshot_v2.json     # seed snapshot (10 real HF datasets, multi-source)
cordis.patch.yml          # bundle patch (installed) / dev patch (checkout)
```

- `npm run typecheck` — strict TypeScript, no errors
- `npm test` — builds, then runs 34 tests including keyless drive tests through the real `ctx.tools.execute` pipeline
- Peer packages (`@deepseek-ai/cordis`, `dsh-tools`, `dsh-llm`) are provided by the DeepSeek Harness deployment, exactly like the official dsh tool plugins.

---

## Data provenance

The seed snapshot carries **real public usage data** for **25 open Hugging Face datasets** — the famous list, 7 sentiment-classification datasets (glue, rotten_tomatoes, sst2, …), and the top-downloads sweep — aggregated from up to four public sources (fetched 2026-09-08):

| Source | Signals |
|--------|---------|
| Hugging Face Hub | downloads, likes, model uses (deep / shallow / deep) |
| ModelScope | downloads, likes (mirrored datasets, probed by short name) |
| DataCite | citation counts (when the dataset card carries a DOI) |
| GitHub | stars (shallow), forks/commits (deep) — only for **curated** dataset→repo mappings whose repo is the dataset's canonical release home (see `DEFAULT_GITHUB_MAP` in `src/snapshot/fetchers/pipeline.ts`) |

Provenance discipline: every metric carries its `source` + `fetched_at`; `model_uses` is exact below the scan cap and recorded as `model_uses_min` (an honest lower bound) at the cap; single-source assets are marked explicitly ("single source only — multi-source aggregation not met"); missing provenance is never fabricated. Hugging Face numbers were fetched through the hf-mirror.com mirror (counts are the mirror's index, which can differ from hf.co's counters); set `DATATALLY_HF_BASE` to refresh from a different channel.

Refresh the seed yourself (four-source pipeline, same core as the plugin):

```bash
datatally refresh --snapshot ./data/snapshot_v2.json
# env: DATATALLY_HF_BASE (default https://huggingface.co),
#      DATATALLY_MODELSCOPE_BASE (default https://modelscope.cn),
#      DATATALLY_GITHUB_TOKEN (optional, raises the GitHub rate limit)
```

---

## License

MIT

## Contributing

Issues and pull requests welcome. Please keep contributions within the stated scope: **recording usage signals, not interpreting them.**

数据来源:公开的 DeepSeek Harness 插件目录与各插件 GitHub 仓库。本站为独立第三方目录,与 DeepSeek、幻方(High-Flyer)及插件作者均无隶属或背书关系。

每日精选 Skill 推荐,免费送到你邮箱

输入邮箱,每天接收一个精选 AI Agent 技能推荐。完全免费,持续更新。

提交后我们会发送一封确认邮件,点击邮件里的链接才会开始收信。

完全免费,取消任意时间。我们不会发送垃圾邮件。