Agent skill

Big Data Labeling Variable Construction

by Drchronx in Drchronx/ai-agent-research-starter-kit

Route text-labeling requests across LDA topic modeling, sklearn baselines, pretrained transformer models, and OpenAI-compatible LLM labeling.

Custom licenceAuto-check passedData & Analytics

Install Big Data Labeling Variable Construction

skills CLI
$ npx skills add Drchronx/ai-agent-research-starter-kit --skill big-data-labeling-variable-construction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Drchronx/ai-agent-research-starter-kit big-data-labeling-variable-construction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Drchronx/ai-agent-research-starter-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'综合学术部署包/05_文本挖掘与NLP Skills/big-data-labeling-variable-construction' .claude/skills/big-data-labeling-variable-construction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
big-data-labeling-variable-construction
GitHub stars
134
Token cost
~1.1k tokens
SKILL.md length
318 words
Files
23 (incl. scripts, assets)
Skills in repo
53
Repo updated
First seen
Licence
Custom licence

At a glance

Route text-labeling requests across LDA topic modeling, sklearn baselines, pretrained transformer models, and OpenAI-compatible LLM labeling.

  • Works in 8 steps: Recommend a route → Prepare a confirmation plan from the… → Wait for user confirmation → …
  • Tasks that involve Data cleaning
  • SKILL.md covers Mandatory interaction pattern, What the user usually does not…, What the user still needs to… and Recommended flow, plus 3 more sections
  • Runs Python scripts from its folder; calls python; reaches api.openai.com

What it does

Big Data Labeling Variable Construction is an agent skill from Drchronx/ai-agent-research-starter-kit. Route text-labeling requests across LDA topic modeling, sklearn baselines, pretrained transformer models, and OpenAI-compatible LLM labeling. First inspect the dataset and infer columns/defaults automatically, then wait for explicit user confirmation before execution.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 25 other files, including scripts and assets (for example `agents/openai.yaml`, `scripts/bert_labeling.py` and `scripts/build_empirical_variables.py`).

It sits in Data & Analytics, covering Data cleaning, Machine learning and Natural language processing. It works with scikit-learn, OpenAI and PyTorch. The repository describes itself as: AI Agent 科研全流程教学包,能教学生从零部署 Codex、Claude Code、OpenClaw、Hermes 等 Agent,学会使用 Skills、飞书、AMiner、AI4Scholar、Zotero、Obsidian…

When your agent uses it

  • Tasks that involve Data cleaning
  • Tasks that involve Machine learning
  • Tasks that involve Natural language processing

Example prompts

  • “/big-data-labeling-variable-construction”

Requirements

  • Python 3
  • A credential in YOUR_KEY

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Recommend a route
  2. Prepare a confirmation plan from the dataset
  3. Wait for user confirmation
  4. Run the actual method script
  5. LDA主题建模
  6. sklearn常见机器学习方法
  7. 预训练模型
  8. 调用大语言模型标注数据

What it can do on your machine

Read from SKILL.md and the folder at commit aab1133. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 13 files in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.openai.com

    Also links to:

    • pytorch.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Big Data Labeling Variable Construction loads about 1.1k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 318 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 318 words (~1,147 tokens).

name
big-data-labeling-variable-construction

Read the full SKILL.md on GitHub

Files

SKILL.md and 22 other files (scripts, assets) in 综合学术部署包/05_文本挖掘与NLP Skills/big-data-labeling-variable-construction of Drchronx/ai-agent-research-starter-kit.

  • SKILL.md
  • agents/openai.yaml
  • assets/sample_labeled_text_dataset.csv
  • assets/sample_text_dataset.csv
  • assets/sample_topic_corpus.csv
  • scripts/bert_labeling.py
  • scripts/build_empirical_variables.py
  • scripts/check_pretrained_env.py
  • scripts/diagnose_torch_runtime.py
  • scripts/ernie_labeling.py
  • scripts/labeling_common.py
  • scripts/labeling_workflow.py
  • scripts/lda_topic_model.py
  • scripts/openai_llm_labeling.py
  • scripts/prepare_labeling_plan.py
  • scripts/requirements.txt
  • scripts/sklearn_knn_labeling.py
  • scripts/sklearn_labeling_common.py
  • … and 5 more

Open the folder on GitHubat commit aab1133

Compare with similar skills

Big Data Labeling Variable Construction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Big Data Labeling Variable Construction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Big Data Labeling Variable Construction this skillDrchronx/ai-agent-research-starter-kit134—~1.1kAutomated safety check: PassCustom licence
Longbridge Quanthelsome/folio2691 repos~1.6kAutomated safety check: PassMIT
ML EngineerRightNow-AI/openfang18k—~987Automated safety check: PassApache-2.0
ML Model Trainingsecondsky/claude-skills2271 repos~1.7kAutomated safety check: PassMIT
Scikit Learn Machine Learningjaechang-hits/SciAgent-Skills3701 repos~4kAutomated safety check: PassBSD-3-Clause
Deep Learningericrisco/rsc-harness156—~3.4kAutomated safety check: PassMIT

Similar skills

  • Longbridge Quant

    helsome/folio

    Quantitative strategy frameworks: pairs trading/cointegration, volatility regime strategies, seasonality/calendar effects, multi-factor models (IC/IR), factor research and screening, correlation…

    269 GitHub starsUsed in 1 repo~1.6k tokens
    Data & AnalyticsAuto-check passed
  • ML Engineer

    RightNow-AI/openfang

    Machine learning engineer expert for PyTorch, scikit-learn, model evaluation, and MLOps

    18k GitHub stars~987 tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • ML Model Training

    secondsky/claude-skills

    Train ML models with scikit-learn, PyTorch, TensorFlow. An agent skill from secondsky/claude-skills.

    227 GitHub starsUsed in 1 repo~1.7k tokens
    Data & AnalyticsAuto-check passed
  • Scikit Learn Machine Learning

    jaechang-hits/SciAgent-Skills

    Classical ML in Python: classification, regression, clustering, dim reduction, evaluation, tuning, preprocessing pipelines.

    370 GitHub starsUsed in 1 repo~4k tokens
    Data & AnalyticsAuto-check passed
  • Deep Learning

    ericrisco/rsc-harness

    A skill your agent uses when training or debugging a neural net in PyTorch — the forward/loss/backward/step loop and its silent bugs, mixed precision (AMP), AdamW/LR schedules, DDP/FSDP/ZeRO…

    156 GitHub stars~3.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Machine Learning

    ericrisco/rsc-harness

    A skill your agent uses when predicting a column from rows of tabular features with classic models — scikit-learn pipelines, RandomForest, XGBoost/LightGBM, leak-free cross-validation, metrics for…

    156 GitHub stars~4.2k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from Drchronx/ai-agent-research-starter-kit

All 53 skills in this repo
  • Literature Reviewer Skill

    Drchronx/ai-agent-research-starter-kit

    Build high-quality literature reviews from a research topic using a 10-phase workflow.

    134 GitHub stars~2.5k tokensUpdated 4 mo ago
    Auto-check passed
  • Scenario Experiment Analysis

    Drchronx/ai-agent-research-starter-kit

    Analyze scenario/vignette experiment datasets for behavioral research.

    134 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Scenario Experiment Benchmark Mining

    Drchronx/ai-agent-research-starter-kit

    Mine and synthesize real top-journal scenario/vignette experiment patterns for behavioral research.

    134 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Web Browsing

    Drchronx/ai-agent-research-starter-kit

    Browse and summarize websites, extract content from URLs, search the web for information.

    134 GitHub stars~593 tokensUpdated 4 mo ago
    Auto-check passed
  • Multi Source Data Integration Extraction

    Drchronx/ai-agent-research-starter-kit

    Automatically merge scattered Excel and CSV files, normalize column names, and extract structured tables from PDF, HTML, TXT, or Markdown documents.

    134 GitHub stars~671 tokensUpdated 4 mo ago
    Auto-check passed
  • Scientific Slides

    Drchronx/ai-agent-research-starter-kit

    Build slide decks and presentations for research talks using Nano Banana Pro AI.

    134 GitHub stars~9.8k tokensUpdated 4 mo ago
    Auto-check: notes

Questions about Big Data Labeling Variable Construction

What does Big Data Labeling Variable Construction do?

Route text-labeling requests across LDA topic modeling, sklearn baselines, pretrained transformer models, and OpenAI-compatible LLM labeling. Big Data Labeling Variable Construction is an agent skill from Drchronx/ai-agent-research-starter-kit. Route text-labeling requests across LDA topic modeling, sklearn baselines, pretrained transformer models, and OpenAI-compatible LLM labeling.

When should I use Big Data Labeling Variable Construction?

Big Data Labeling Variable Construction fits situations like: tasks that involve Data cleaning; tasks that involve Machine learning; tasks that involve Natural language processing.

How do I install Big Data Labeling Variable Construction in Claude Code?

Run `npx skills add Drchronx/ai-agent-research-starter-kit --skill big-data-labeling-variable-construction -a claude-code`. Or copy the skill folder (综合学术部署包/05_文本挖掘与NLP Skills/big-data-labeling-variable-construction in Drchronx/ai-agent-research-starter-kit) into .claude/skills/big-data-labeling-variable-construction in your project. Claude Code loads it when a task matches its description.

How do I install Big Data Labeling Variable Construction in Codex?

Run `npx skills add Drchronx/ai-agent-research-starter-kit --skill big-data-labeling-variable-construction -a codex`. Or copy the skill folder (综合学术部署包/05_文本挖掘与NLP Skills/big-data-labeling-variable-construction in Drchronx/ai-agent-research-starter-kit) into .agents/skills/big-data-labeling-variable-construction in your project. Codex loads it when a task matches its description.

Can I use Big Data Labeling Variable Construction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Drchronx/ai-agent-research-starter-kit --skill big-data-labeling-variable-construction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/big-data-labeling-variable-construction, .gemini/skills/big-data-labeling-variable-construction, .github/skills/big-data-labeling-variable-construction and .opencode/skills/big-data-labeling-variable-construction in your project.

What does Big Data Labeling Variable Construction need to run?

Going by SKILL.md and its folder, Big Data Labeling Variable Construction needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3; A credential in YOUR_KEY.

Does Big Data Labeling Variable Construction access the network?

SKILL.md names 2 domains. In commands or code: api.openai.com; the agent is likely to contact it when it follows the instructions. As links in the text: pytorch.org. This is read from the text; nothing was executed.

Is Big Data Labeling Variable Construction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Big Data Labeling Variable Construction use?

Big Data Labeling Variable Construction has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Big Data Labeling Variable Construction use?

About 1.1k tokens (SKILL.md is roughly 4.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Big Data Labeling Variable Construction?

Skills that share tags, products or a category with Big Data Labeling Variable Construction: Longbridge Quant (helsome/folio, 269 stars), ML Engineer (RightNow-AI/openfang, 18k stars), ML Model Training (secondsky/claude-skills, 227 stars) and Scikit Learn Machine Learning (jaechang-hits/SciAgent-Skills, 370 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Big Data Labeling Variable Construction?

Drchronx (a GitHub user) maintains it in Drchronx/ai-agent-research-starter-kit, which has 134 GitHub stars. The repository holds 53 skills in this directory. The repository was last updated on May 19, 2026.

Source: Drchronx/ai-agent-research-starter-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.