A skill your agent uses when a broad paper candidate pool needs deterministic deduplication and a stable core set.

No licenceAuto-check passedData & Analytics

Install Dedupe Rank

skills CLI
$ npx skills add WILLOSCAR/research-units-pipeline-skills --skill dedupe-rank -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install WILLOSCAR/research-units-pipeline-skills dedupe-rank --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/WILLOSCAR/research-units-pipeline-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/dedupe-rank .claude/skills/dedupe-rank && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dedupe-rank
GitHub stars
513
Token cost
~438 tokens
SKILL.md length
159 words
Files
5 (incl. scripts, references, assets)
Skills in repo
107
Repo updated
First seen
Licence
None found

At a glance

A skill your agent uses when a broad paper candidate pool needs deterministic deduplication and a stable core set.

  • A broad paper candidate pool needs deterministic deduplication and a stable core set
  • SKILL.md covers Triggers & routing, Input, Outputs and Script boundary, plus 3 more sections
  • Runs Python scripts from its folder
  • Tasks that involve Data cleaning

What it does

Dedupe Rank is an agent skill from WILLOSCAR/research-units-pipeline-skills. Use when a broad paper candidate pool needs deterministic deduplication and a stable core set.

Its SKILL.md is about 440 tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts, reference files and assets (for example `assets/domain_packs/embodied_ai.json`, `assets/domain_packs/llm_agents.json` and `references/domain_pack_overview.md`).

It sits in Data & Analytics, covering Data cleaning. The repository describes itself as: Research pipelines as semantic execution units: each skill declares inputs/outputs, acceptance criteria, and guardrails. Evidence-first methodology prevents hollow writing…

When your agent uses it

  • A broad paper candidate pool needs deterministic deduplication and a stable core set
  • Tasks that involve Data cleaning

Example prompts

  • “/dedupe-rank”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit c92912a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dedupe Rank loads about 438 tokens when it runs, and up to ~696 if it reads all its reference files. Until then it costs about 27 tokens; SKILL.md has 159 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~27
When it runs · the whole SKILL.md, loaded when a task matches
~438
With references · SKILL.md plus every file in references/, read only if the agent opens them
~696

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 159 words (~438 tokens).

“Turns a raw candidate pool into a deduped pool and a stable core set.”

— opening of SKILL.md by WILLOSCAR
name
dedupe-rank

Read the full SKILL.md on GitHub

Files

SKILL.md and 4 other files (scripts, references, assets) in .codex/skills/dedupe-rank of WILLOSCAR/research-units-pipeline-skills.

  • SKILL.md
  • assets/domain_packs/embodied_ai.json
  • assets/domain_packs/llm_agents.json
  • references/domain_pack_overview.md
  • scripts/run.py

Open the folder on GitHubat commit c92912a

Compare with similar skills

Dedupe Rank next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dedupe Rank compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dedupe Rank this skillWILLOSCAR/research-units-pipeline-skills513—~438Automated safety check: PassNone
Question2reportrefraction-ray/xalpha2.7k—~3.2kAutomated safety check: PassMIT
Dingo VerifyMigoXLab/dingo757—~741Automated safety check: NotesApache-2.0
Data Validationplatonai/Browser41.2k—~896Automated safety check: PassApache-2.0
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT
Issues DeduplicationJetBrains/ideavim10k—~1.3kAutomated safety check: PassMIT

Similar skills

  • Question2report

    refraction-ray/xalpha

    Turn a natural-language financial question into a polished, self-contained HTML report.

    2.7k GitHub stars~3.2k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~741 tokensUpdated 11 days ago
    Data & AnalyticsAuto-check: notes
  • Data Validation

    platonai/Browser4

    Validates data against common and custom rules (required fields, formats, ranges).

    1.2k GitHub stars~896 tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Issues Deduplication

    JetBrains/ideavim

    Official

    Handles deduplication of YouTrack issues. An agent skill from JetBrains/ideavim.

    10k GitHub stars~1.3k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Openbb Data Fetcher

    monarchjuno/vibe-investing

    Fetch financial, market, economic, fundamental, news, options, crypto, ETF, index, and macro data through the OpenBB Python interface instead of the OpenBB MCP server.

    299 GitHub stars~2.9k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check: notes

More from WILLOSCAR/research-units-pipeline-skills

All 107 skills in this repo
  • Appendix Table Writer

    WILLOSCAR/research-units-pipeline-skills

    Curate reader-facing survey tables for the Appendix (clean layout + high information density), using only in-scope evidence and existing citation keys.

    513 GitHub stars~1.8k tokensUpdated 4 days ago
    Auto-check passed
  • Artifact Contract Auditor

    WILLOSCAR/research-units-pipeline-skills

    Audit one research Workspace for declared Unit outputs and Pipeline target Artifacts, writing output/CONTRACTREPORT.md; use for mid-Run coverage snapshots or final delivery completeness, not deep…

    513 GitHub stars~918 tokensUpdated 4 days ago
    Auto-check passed
  • Arxiv Search

    WILLOSCAR/research-units-pipeline-skills

    Retrieve arXiv paper metadata with keyword queries or import an offline arXiv export, and save results as JSONL (papers/papersraw.jsonl).

    513 GitHub stars~1.5k tokensUpdated 4 days ago
    Auto-check passed
  • Chapter Lead Writer

    WILLOSCAR/research-units-pipeline-skills

    Write H2 chapter lead blocks (sections/S<secidlead.md) that preview the chapter's comparison lens and connect its H3 subsections, without adding new facts.

    513 GitHub stars~1.5k tokensUpdated 4 days ago
    Auto-check passed
  • Chapter Skeleton

    WILLOSCAR/research-units-pipeline-skills

    Build a retrieval-informed chapter skeleton (outline/chapterskeleton.yml) from taxonomy/core scope before stable H3 decomposition.

    513 GitHub stars~475 tokensUpdated 4 days ago
    Auto-check passed
  • Evaluation Anchor Checker

    WILLOSCAR/research-units-pipeline-skills

    Audit and rewrite evaluation/numeric claims to ensure they carry minimal protocol context (task + metric + constraint) and avoid underspecified model naming.

    513 GitHub stars~1.4k tokensUpdated 4 days ago
    Auto-check passed

Questions about Dedupe Rank

What does Dedupe Rank do?

A skill your agent uses when a broad paper candidate pool needs deterministic deduplication and a stable core set. Dedupe Rank is an agent skill from WILLOSCAR/research-units-pipeline-skills. Use when a broad paper candidate pool needs deterministic deduplication and a stable core set.

When should I use Dedupe Rank?

Dedupe Rank fits situations like: A broad paper candidate pool needs deterministic deduplication and a stable core set; tasks that involve Data cleaning.

How do I install Dedupe Rank in Claude Code?

Run `npx skills add WILLOSCAR/research-units-pipeline-skills --skill dedupe-rank -a claude-code`. Or copy the skill folder (.codex/skills/dedupe-rank in WILLOSCAR/research-units-pipeline-skills) into .claude/skills/dedupe-rank in your project. Claude Code loads it when a task matches its description.

How do I install Dedupe Rank in Codex?

Run `npx skills add WILLOSCAR/research-units-pipeline-skills --skill dedupe-rank -a codex`. Or copy the skill folder (.codex/skills/dedupe-rank in WILLOSCAR/research-units-pipeline-skills) into .agents/skills/dedupe-rank in your project. Codex loads it when a task matches its description.

Can I use Dedupe Rank in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add WILLOSCAR/research-units-pipeline-skills --skill dedupe-rank -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dedupe-rank, .gemini/skills/dedupe-rank, .github/skills/dedupe-rank and .opencode/skills/dedupe-rank in your project.

What does Dedupe Rank need to run?

Going by SKILL.md and its folder, Dedupe Rank needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Dedupe Rank access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Dedupe Rank safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Dedupe Rank use?

No licence was found for Dedupe Rank or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Dedupe Rank use?

About 438 tokens (SKILL.md is roughly 1.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 258 tokens, read only when the agent opens those files.

What are the alternatives to Dedupe Rank?

Skills that share tags, products or a category with Dedupe Rank: Question2report (refraction-ray/xalpha, 2.7k stars), Dingo Verify (MigoXLab/dingo, 757 stars), Data Validation (platonai/Browser4, 1.2k stars) and Pandas Pro (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dedupe Rank?

WILLOSCAR (a GitHub user) maintains it in WILLOSCAR/research-units-pipeline-skills, which has 513 GitHub stars. The repository holds 107 skills in this directory. The repository was last updated on October 5, 2026.

Source: WILLOSCAR/research-units-pipeline-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.