A skill your agent uses when planning or auditing the analysis of a Language (LSA) manuscript so the evidence credibly supports the theoretical claim.

MITAuto-check passedData & Analytics

Install Lang Data Analysis

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill lang-data-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills lang-data-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/Language-Linguistic-Society-Skills/skills/lang-data-analysis .claude/skills/lang-data-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
lang-data-analysis
GitHub stars
1.2k
Token cost
~1.9k tokens
SKILL.md length
871 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when planning or auditing the analysis of a Language (LSA) manuscript so the evidence credibly supports the theoretical claim.

  • Auditing the analysis of a Language (LSA) manuscript so the evidence credibly supports the theoretical claim
  • SKILL.md covers When to trigger, Analysis norms (by mode), Convergent evidence (a… and Referee-pushback patterns on…, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Data analysis

What it does

Lang Data Analysis is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when planning or auditing the analysis of a Language (LSA) manuscript so the evidence credibly supports the theoretical claim. Covers quantitative modeling (mixed-effects in R), phonetic measurement, corpus statistics, and the analytic trail from glossed data or judgments to the generalization. Improves the analysis chain; it does not fabricate results.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data analysis and Statistics. The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Auditing the analysis of a Language (LSA) manuscript so the evidence credibly supports the theoretical claim
  • Tasks that involve Data analysis
  • Tasks that involve Statistics

Example prompts

  • “/lang-data-analysis”

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Lang Data Analysis loads about 1.9k tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 871 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~95
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 871 words, ~1,935 tokens.

Download SKILL.mdSave it as .claude/skills/lang-data-analysis/SKILL.md (or your agent's skills folder).
name
lang-data-analysis
description
Use when planning or auditing the analysis of a Language (LSA) manuscript so the evidence credibly supports the theoretical claim. Covers quantitative modeling (mixed-effects in R), phonetic measurement, corpus statistics, and the analytic trail from glossed data or judgments to the generalization. Improves the analysis chain; it does not fabricate results.

Data Analysis (lang-data-analysis)

At Language the analysis exists to make the theoretical claim credible — not to display technique or notation. A cross-subfield, double-anonymous reviewer will ask whether the evidence actually warrants the generalization and whether uncertainty is handled honestly. Where the work is quantitative, Language now expects properly specified models (typically mixed-effects models in R) rather than by-subject t-tests or raw counts; where it is analytic, it expects the pattern to be demonstrable from the glossed data. This skill stress-tests the analysis chain in the idiom of your work.

When to trigger

  • Planning the analysis, or auditing it before writing up
  • A reader doubts the statistics, the evidence-to-claim link, or the treatment of variability
  • Reconciling multiple data sources (corpus + experiment, judgments + text) into one argument
  • Deciding which analyses are confirmatory vs. exploratory

Analysis norms (by mode)

Quantitative (experiment / corpus)
  • Fit mixed-effects models with the random-effects structure the design justifies (crossed by-subject and by-item random effects; random slopes for within-cluster predictors). Report the model, not just p-values.
  • Report effect sizes and intervals, not stars alone; state the coding/contrasts and the convergence status; keep seeds and pinned package versions.
  • Distinguish preregistered/confirmatory from exploratory analyses where applicable.
Phonetic
  • State measurement settings (windowing, formant ceilings, alignment) and how outliers/mis-tracks were handled; show the effect is not an artifact of the measurement pipeline.
Analytic (formal)
  • Demonstrate the generalization directly from numbered, glossed examples; show the analysis derives the attested cases and blocks the unattested ones.
Historical / typological
  • Make the inferential logic explicit (implicational universals, reconstruction, statistical tendencies); guard against non-independence of the sample.

Convergent evidence (a Language strength)

Language rewards a generalization shown through more than one window — e.g., an experimental effect corroborated by a corpus trend, or judgments backed by text frequencies. When windows disagree, say so and explain the discrepancy rather than hiding the inconvenient one.

Referee-pushback patterns on the evidence chain (Language fixes)

Referee writes…The Language-specific fix
"No random effects / pseudoreplication."fit the justified mixed model; cluster by subject and item
"Significance without effect size."report estimates + intervals in interpretable units
"The stat model doesn't match the design."align random-effects structure with the sampling
"Analysis doesn't rule out the alternative."show it derives attested and blocks unattested cases

Calibration (Language appetite, hedged)

Orienting heuristics; confirm against the current author pages. Language increasingly expects that a quantitative claim rests on a model appropriate to the clustered, repeated-measures nature of linguistic data — the modal avoidable failure is pseudoreplication (ignoring by-speaker or by-item structure). Illustrative: a paper claims a durational contrast "is significant (p < .01)" from 1,200 tokens produced by 8 speakers, analyzed as if independent. A referee writes "pseudoreplication." The fix refits a mixed-effects model with by-speaker and by-word random intercepts and slopes, reports the estimate (an illustrative 12 ms, 95% CI ~4–20), and notes two speakers who show no effect — turning a fragile claim into a credible, bounded one.

Show full SKILL.md (394 more words)Show less

Execution bridge (StatsPAI / Stata MCP)

Language asks for a properly specified model, which is a claim about a fitted object, not about a paragraph. Fit it and report from it. Full map: execution-with-mcp.

  • Mixed-effects: mixed (continuous responses) and melogit / meglm (binary and categorical ones) carry the crossed by-subject and by-item structure the design justifies; icc states how much clustering there actually is, which is the number a reviewer needs when the sample's non-independence is the objection.
  • Uncertainty: bootstrap for intervals where the asymptotics are thin — small fieldwork samples and unbalanced cells, both routine here.
  • Multiple comparisons: holm or benjamini_hochberg across a family of contrasts. A typological or corpus paper testing many predictors at once needs this stated, not assumed.
  • Reporting: etable / margins so the effect sizes and intervals in the prose are the fitted ones.

Where a server is not connected, adapt the ../../resources/code/ skeleton and say so — never report a number you did not compute. This bridge touches the quantitative strand only; analytic and historical arguments are made from the glossed examples themselves.

Anti-patterns

  • Treating repeated measures as independent (pseudoreplication); stars-only reporting
  • A statistical model whose random structure ignores the sampling design
  • Phonetic effects that are artifacts of measurement settings, not language
  • Cherry-picked examples that ignore counterexamples in the same corpus/elicitation
  • Presenting exploratory results as if confirmatory
  • Notation or technique foregrounded over the generalization it is meant to support

Evidence pass for Language

Treat this skill as an executable review pass, not a prose hint. First lock the empirical generalization, evidence base, warrant, and theoretical payoff; then judge whether the manuscript answers the venue's real reader: linguists across subfields who value grounded analysis, transparent and checkable evidence, and careful, appropriately scoped generalizations.

  • Do the pass: audit the analysis before polishing prose — unit of analysis, random-effects structure, effect sizes, measurement pipeline, exclusions, and reproducibility must be visible.
  • Return a ledger: give claim / evidence / risk / manuscript location rows so the next agent can edit rather than rediscover the issue.
  • Sibling guard: compare against Laboratory Phonology, Journal of Memory and Language, Language Variation and Change; if a sibling owns the contribution, recommend re-routing before polishing.
  • Stop condition: do not give submission-ready advice until resources/official-source-map.md has been checked and the manuscript has one concrete fix for the largest venue-specific risk.

Output format

【Claim under test】from theory-building
【Primary evidence】the analysis that carries the claim
【Model】mixed-effects structure matches the design? [Y/N/NA]
【Uncertainty】effect sizes + intervals reported? [Y/N]
【Convergence】corroborated across windows? [Y/N/NA]
【Confirmatory vs. exploratory】labeled where relevant? [Y/N]
【Next】lang-data-and-transparency

Supplementary resources

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in Language-Linguistic-Society-Skills/skills/lang-data-analysis of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Lang Data Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Lang Data Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Lang Data Analysis this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.9kAutomated safety check: PassMIT
MatlabzLanqing/codex-claude-academic-skills4.6k9 repos~2.3kAutomated safety check: NotesGPL-3.0
Eqtl Catalogue Region FetchClawBio/ClawBio1.2k1 repos~4.3kAutomated safety check: PassMIT
CSV Data Analysis5zjk5/prompt-engineering127—~2.6kAutomated safety check: PassNone
Meridian MMM Model Buildinggoogle/meridian1.6k—~2.5kAutomated safety check: PassApache-2.0
Gwas Catalog Region FetchClawBio/ClawBio1.2k1 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • Matlab

    zLanqing/codex-claude-academic-skills

    MATLAB and GNU Octave numerical computing for matrix operations, data analysis, visualization, and scientific computing.

    4.6k GitHub starsUsed in 9 repos~2.3k tokens
    Data & AnalyticsAuto-check: notes
  • Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~4.3k tokens
    Data & AnalyticsAuto-check passed
  • CSV Data Analysis

    5zjk5/prompt-engineering

    This skill should be used when users need to analyze CSV or Excel files, understand data patterns, generate statistical summaries, or create data visualizations.

    127 GitHub stars~2.6k tokensUpdated 23 days ago
    Data & AnalyticsAuto-check passed
  • Official

    Takes a user through building a Meridian marketing mix model, from loading CSV data and mapping columns to running EDA, fitting and saving the model.

    1.6k GitHub stars~2.5k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Data & AnalyticsAuto-check passed
  • Statistical Data Analysis

    lingzhi227/agent-research-skills

    Writes statistical analysis code for experimental data, runs it through a four-round review, and reports effect sizes, p-values and confidence intervals.

    384 GitHub stars~886 tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 11 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 11 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 11 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 11 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 11 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 11 days ago
    Auto-check passed

Questions about Lang Data Analysis

What does Lang Data Analysis do?

A skill your agent uses when planning or auditing the analysis of a Language (LSA) manuscript so the evidence credibly supports the theoretical claim. Lang Data Analysis is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when planning or auditing the analysis of a Language (LSA) manuscript so the evidence credibly supports the theoretical claim.

When should I use Lang Data Analysis?

Lang Data Analysis fits situations like: auditing the analysis of a Language (LSA) manuscript so the evidence credibly supports the theoretical claim; tasks that involve Data analysis; tasks that involve Statistics.

How do I install Lang Data Analysis in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill lang-data-analysis -a claude-code`. Or copy the skill folder (Language-Linguistic-Society-Skills/skills/lang-data-analysis in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/lang-data-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Lang Data Analysis in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill lang-data-analysis -a codex`. Or copy the skill folder (Language-Linguistic-Society-Skills/skills/lang-data-analysis in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/lang-data-analysis in your project. Codex loads it when a task matches its description.

Can I use Lang Data Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill lang-data-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/lang-data-analysis, .gemini/skills/lang-data-analysis, .github/skills/lang-data-analysis and .opencode/skills/lang-data-analysis in your project.

What does Lang Data Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: Lang Data Analysis is instructions for the agent only.

Does Lang Data Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Lang Data Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Lang Data Analysis use?

Lang Data Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Lang Data Analysis use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Lang Data Analysis?

Skills that share tags, products or a category with Lang Data Analysis: Matlab (zLanqing/codex-claude-academic-skills, 4.6k stars), Eqtl Catalogue Region Fetch (ClawBio/ClawBio, 1.2k stars), CSV Data Analysis (5zjk5/prompt-engineering, 127 stars) and Meridian MMM Model Building (google/meridian, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Lang Data Analysis?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,219 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.