Agent skill

2026 Legal Research Agent

by curiositech in curiositech/some_claude_skills

Expert legal research agent for finding and scraping expungement data state by state.

MITAuto-check passedLegal & Compliance

Install 2026 Legal Research Agent

skills CLI
$ npx skills add curiositech/some_claude_skills --skill 2026-legal-research-agent -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install curiositech/some_claude_skills 2026-legal-research-agent --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/curiositech/some_claude_skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/2026-legal-research-agent .claude/skills/2026-legal-research-agent && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
2026-legal-research-agent
GitHub stars
244
Token cost
~2.6k tokens
SKILL.md length
678 words
Files
5 (incl. scripts, references)
Skills in repo
95
Repo updated
First seen
Licence
MIT

At a glance

Expert legal research agent for finding and scraping expungement data state by state.

  • Works in 6 steps: Authoritative Source Hierarchy → URL Pattern Knowledge by State Type → 2026 Legal Landscape Awareness → …
  • Tasks that involve Web scraping
  • SKILL.md covers When to Use This Skill, Core Instructions, Anti-Patterns and Project-Specific Context, plus 2 more sections
  • Runs TypeScript scripts from its folder; calls npx; needs FIRECRAWL_API_KEY

What it does

2026 Legal Research Agent is an agent skill from curiositech/some_claude_skills. Expert legal research agent for finding and scraping expungement data state by state. Knows authoritative sources, URL patterns, Firecrawl configuration, and 2026 legal landscape. Activate on "find expungement data", "scrape state laws", "legal research", "court URLs", "statute sources", "Clean Slate laws", "automatic expungement research". NOT for interpreting laws (use national-expungement-expert), building UI, or legal advice.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `.claude-plugin/plugin.json`, `references/clean-slate-timeline.md` and `references/url-patterns-by-state.md`).

It sits in Legal & Compliance, covering Web scraping and Legal research. It works with Firecrawl. The repository describes itself as: Claude skills that make my life easier. The licence is MIT.

When your agent uses it

  • Tasks that involve Web scraping
  • Tasks that involve Legal research

Example prompts

  • “find expungement data”
  • “scrape state laws”
  • “legal research”
  • “/2026-legal-research-agent”

Requirements

  • Node.js
  • A credential in FIRECRAWL_API_KEY

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Authoritative Source Hierarchy
  2. URL Pattern Knowledge by State Type
  3. 2026 Legal Landscape Awareness
  4. Firecrawl Configuration Expertise
  5. Data Validation Checklist
  6. Gap Analysis Process

What it can do on your machine

Read from SKILL.md and the folder at commit 6713fc7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (TypeScript), which the agent can run.

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • FIRECRAWL_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

2026 Legal Research Agent loads about 2.6k tokens when it runs, and up to ~7.6k if it reads all its reference files. Until then it costs about 115 tokens; SKILL.md has 678 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~115
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from curiositech/some_claude_skills at commit 6713fc7, republished under its MIT licence (© curiositech). 678 words, ~2,591 tokens.

Download SKILL.mdSave it as .claude/skills/2026-legal-research-agent/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
2026-legal-research-agent
description
Expert legal research agent for finding and scraping expungement data state by state. Knows authoritative sources, URL patterns, Firecrawl configuration, and 2026 legal landscape. Activate on "find expungement data", "scrape state laws", "legal research", "court URLs", "statute sources", "Clean Slate laws", "automatic expungement research". NOT for interpreting laws (use national-expungement-expert), building UI, or legal advice.
metadata.gated
true
metadata.category
Research & Analysis
metadata.tags
imported, needs-review

When to Use This Skill

Use this skill when you need to:

  • Find authoritative legal sources for a specific state's expungement laws
  • Configure Firecrawl jobs to scrape court systems, legislatures, or legal aid sites
  • Validate scraped data for accuracy and completeness
  • Research 2026 law changes including Clean Slate acts and marijuana expungement
  • Build URL patterns for systematic state-by-state data collection
  • Identify gaps in existing scraped data coverage

Do NOT use this skill for:

  • Interpreting what laws mean (use national-expungement-expert)
  • Building user interfaces or components
  • Providing legal advice to users
  • General web scraping unrelated to legal data

Core Instructions

1. Authoritative Source Hierarchy

When researching expungement laws, prioritize sources in this order:

Tier 1 (Primary Authority):
├── State Legislature websites (statute text)
├── State Court Administrative Office
└── State Attorney General publications

Tier 2 (Official Secondary):
├── State Bar Association guides
├── Court self-help centers
└── Public law databases (public.law, justia.com)

Tier 3 (Tertiary but Valuable):
├── Legal aid organizations (LSC grantees)
├── Law school clinics
└── Reentry organizations (CCRC, NACDL)

Tier 4 (Verification Only):
├── Commercial legal databases
├── News articles about law changes
└── Attorney blog posts

Shibboleth: A novice scrapes the first Google result. An expert knows that courts.{state}.gov contains the self-help forms while legislature.{state}.gov contains the statute text—and both are needed.

2. URL Pattern Knowledge by State Type

States organize their legal resources differently. Know the patterns:

Unified Court Systems (courts own everything):

California: courts.ca.gov/selfhelp-expungement.htm
Oregon: courts.oregon.gov/programs/exp/Pages/default.aspx
Washington: courts.wa.gov/forms/?fa=forms.contribute&formID=101

Split Systems (legislature + court separate):

Texas: txcourts.gov (forms) + texas.public.law (statutes)
New York: nycourts.gov (forms) + nysenate.gov/legislation/laws (statutes)
Florida: flcourts.gov (forms) + leg.state.fl.us/statutes (statutes)

Public.law States (excellent statute hosting):

oregon.public.law, california.public.law, texas.public.law
michigan.public.law, washington.public.law

Shibboleth: Knowing that apps.leg.wa.gov/RCW/ is Washington's statute database while leg.wa.gov is the general legislature site—the RCW subdomain is where the actual law text lives.

As of 2026, these major changes affect research:

Clean Slate States (automatic expungement passed):

  • Pennsylvania (2018), Utah (2019), New Jersey (2019), Michigan (2020)
  • California (2020), Connecticut (2021), Delaware (2021), Virginia (2021)
  • Oklahoma (2022), Colorado (2022), New York (2023), Minnesota (2023)
  • Maryland (2024), Illinois (2024), Oregon (2025)

Marijuana Expungement (specific statutes):

  • Most states now have separate marijuana expungement provisions
  • Search for "cannabis conviction" alongside "expungement"
  • Check for retroactive application dates

2025-2026 Law Changes to Verify:

  • Oregon HB 2316 (expanded eligibility)
  • California AB 1076 (automatic relief expansion)
  • Check CCRC's Restoration of Rights Project for current status

Shibboleth: Knowing that "automatic expungement" doesn't mean immediate—Pennsylvania's Clean Slate has a 10-year waiting period for arrests and varies by offense. Research must capture these nuances.

4. Firecrawl Configuration Expertise

When configuring scrape jobs:

Extraction Schema Design:

typescript
// For statute pages, extract:
{
  statuteCitation: "string",   // e.g., "ORS 137.225"
  title: "string",             // e.g., "Setting aside conviction"
  fullText: "string",          // Complete statute text
  effectiveDate: "string",     // When current version took effect
  lastAmended: "string",       // Most recent amendment date
  subsections: "array",        // Parsed subsections
}

// For court self-help pages, extract:
{
  stateName: "string",
  expungementPageUrl: "string",
  formsLibraryUrl: "string",
  selfHelpUrl: "string",
  contactPhone: "string",
  feeScheduleUrl: "string",
}

// For forms, extract:
{
  formNumber: "string",        // e.g., "MC-440"
  formTitle: "string",
  pdfUrl: "string",
  applicableTo: "array",       // ["misdemeanor", "arrest"]
  lastUpdated: "string",
}

Rate Limiting for Government Sites:

typescript
rateLimit: 2,  // 2 requests/second max for .gov sites
timeout: 90000,  // Government sites can be slow
maxRetries: 3,  // Retry on timeout
waitFor: 3000,  // Wait for JavaScript on modern court sites

Shibboleth: Knowing to set onlyMainContent: true for statute pages (to skip navigation chrome) but onlyMainContent: false for forms pages (where the form links are often in sidebars).

5. Data Validation Checklist

After scraping, validate:

□ Statute citations match official format (e.g., "ORS" not "Or. Rev. Stat.")
□ Effective dates are parseable and reasonable (not future, not too old)
□ URLs are live and return 200 status
□ PDF form links actually download PDFs (not HTML error pages)
□ Phone numbers are in consistent format
□ Fee amounts are numeric and reasonable ($0-$500 typical range)
□ State code extracted correctly (watch for ambiguous URLs)

Common Extraction Errors:

  • "oregon.public.law" matching "la" (Louisiana) instead of "or" (Oregon)
  • Statute text truncated at 10,000 characters (increase limit)
  • Form "last updated" dates in inconsistent formats
  • County-specific URLs mistaken for state-level
Show full SKILL.md (275 more words)Show less
6. Gap Analysis Process

To identify missing data for a state:

bash
# Check what we have
ls src/data/scraped/states/{state}/

# Expected files for complete coverage:
# - statutes.json (eligibility rules from statute text)
# - court-system.json (court URLs, contacts, forms links)
# - forms/ (actual PDF forms)
# - fees.json (filing fee amounts)
# - counties/ (county-specific court data)

# Cross-reference with state data file
grep -l "waitingPeriods\|eligibilityRules" src/data/states/{state}.ts

Priority order for filling gaps:

  1. Statutes (foundation for all rules)
  2. Court forms (what users actually need to file)
  3. Fee information (users need to budget)
  4. County contacts (where to file)
  5. Spanish resources (accessibility)

Anti-Patterns

Never Do These:
  1. Scrape without rate limiting - Government sites will block you
  2. Trust secondary sources for statute text - Always verify against primary
  3. Assume URL patterns are consistent - Each state is different
  4. Ignore effective dates - Laws change; scraped data needs timestamps
  5. Scrape county sites without state context - County rules supplement, not replace, state law
  6. Skip the self-help sections - Often have the clearest eligibility summaries
  7. Treat all states the same - Clean Slate states have fundamentally different processes
Common Mistakes:
❌ Scraping Wikipedia for statute text
✅ Scraping the state legislature's official code

❌ Using findlaw.com as primary source
✅ Using findlaw.com to find the citation, then scraping the official source

❌ Assuming "expungement" is the only term
✅ Searching for: expungement, sealing, set-aside, dismissal, destruction, pardons

❌ Treating waiting periods as simple numbers
✅ Capturing offense-specific waiting periods (felonies vs misdemeanors vs arrests)

Project-Specific Context

This skill is designed for the National Expungement Guide project:

Existing Infrastructure
  • Firecrawl scripts: scripts/firecrawl/
  • Job definitions: scripts/firecrawl/jobs.ts (P0-P4 priority jobs)
  • URL config: scripts/firecrawl/config.ts (all 50 states)
  • Output path: src/data/scraped/states/{state}/
  • State data: src/data/states/ (TypeScript files per state)
Running Scrapes
bash
# Set API key first
export FIRECRAWL_API_KEY=your_key

# Run P0 (state statutes + courts) - ~$0.20 cost
npx tsx scripts/firecrawl/run-p0.ts

# Dry run to preview
npx tsx scripts/firecrawl/run-p0.ts --dry-run

# Check reports
cat scripts/firecrawl/reports/p0-*.json
Data Flow
Firecrawl scrape → src/data/scraped/{state}/*.json
       ↓
Manual review + cleanup
       ↓
Integrated into src/data/states/{state}.ts
       ↓
Used by eligibility wizard + PDF generator

References

See references/ folder for:

  • url-patterns-by-state.md - Complete URL patterns for all 50 states
  • clean-slate-timeline.md - When each Clean Slate law passed and took effect
  • firecrawl-schemas.md - All extraction schemas used

Example Workflow

User request: "Research California's 2026 expungement laws and scrape the latest data"

Agent workflow:

  1. Check existing data: ls src/data/scraped/states/ca/
  2. Verify current statute version at california.public.law
  3. Check for 2025-2026 law changes via CCRC or news search
  4. Update scripts/firecrawl/config.ts if URLs changed
  5. Run targeted scrape: add CA-specific URLs to P0 job
  6. Validate extracted data against known statute citations
  7. Document any gaps or changes found

© curiositech, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in .claude/skills/2026-legal-research-agent of curiositech/some_claude_skills.

  • SKILL.md
  • .claude-plugin/plugin.json
  • references/clean-slate-timeline.md
  • references/url-patterns-by-state.md
  • scripts/validate-scraped-data.ts

Open the folder on GitHubat commit 6713fc7

Compare with similar skills

2026 Legal Research Agent next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

2026 Legal Research Agent compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
2026 Legal Research Agent this skillcuriositech/some_claude_skills244—~2.6kAutomated safety check: PassMIT
Firecrawl Build Onboardingfirecrawl/firecrawl190k1 repos~1.4kAutomated safety check: NotesISC
Firecrawl Page Scrape Integrationfirecrawl/firecrawl190k1 repos~944Automated safety check: PassISC
Firecrawl App Integrationfirecrawl/firecrawl190k—~2.1kAutomated safety check: NotesISC
Firecrawl Interact Integrationfirecrawl/firecrawl190k1 repos~731Automated safety check: PassISC
Firecrawl Search Integrationfirecrawl/firecrawl190k1 repos~1.1kAutomated safety check: PassISC

Similar skills

  • Firecrawl Build Onboarding

    firecrawl/firecrawl

    Gets Firecrawl working in a project: signs you in through the browser, saves FIRECRAWL_API_KEY to .env and picks the first SDK or REST path.

    190k GitHub starsUsed in 1 repo~1.4k tokens
    Backend & APIsAuto-check: notes
  • Adds Firecrawl's /scrape endpoint to application code to pull markdown, HTML, links, screenshots or structured data from a single known URL.

    190k GitHub starsUsed in 1 repo~944 tokens
    Data & AnalyticsAuto-check passed
  • Firecrawl App Integration

    firecrawl/firecrawl

    Adds web search, scraping, structured extraction and browser interaction to application code using Firecrawl's scrape, search and interact endpoints.

    190k GitHub stars~2.1k tokensUpdated today
    Data & AnalyticsAuto-check: notes
  • Guides adding Firecrawl's /interact endpoint to product code for pages that need clicks, forms, pagination or logged-in flows beyond plain scraping.

    190k GitHub starsUsed in 1 repo~731 tokens
    Data & AnalyticsAuto-check passed
  • Firecrawl Search Integration

    firecrawl/firecrawl

    Guidance for adding Firecrawl's /search endpoint to product code and agent workflows when a feature starts from a query rather than a URL.

    190k GitHub starsUsed in 1 repo~1.1k tokens
    Backend & APIsAuto-check passed
  • Power Design

    ItsssssJack/power-design

    Generate beautiful, on-brand HTML — presentation decks or full responsive websites — in any brand's design language, combining brand DNA extracted via Firecrawl with codified, research-backed design…

    722 GitHub stars~2.8k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed

More from curiositech/some_claude_skills

All 95 skills in this repo
  • Crisis Detection Intervention AI

    curiositech/some_claude_skills

    Detect crisis signals in user content using NLP, mental health sentiment analysis, and safe intervention protocols.

    244 GitHub starsUsed in 2 repos~3.8k tokens
    Auto-check passed
  • Form Validation Architect

    curiositech/some_claude_skills

    End-to-end form handling with react-hook-form, Zod schemas, validation patterns, error messaging, field arrays, and multi-step wizards.

    244 GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • GitHub Actions Pipeline Builder

    curiositech/some_claude_skills

    Build production CI/CD pipelines with GitHub Actions. An agent skill from curiositech/some_claude_skills.

    244 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check: notes
  • Background Job Orchestrator

    curiositech/some_claude_skills

    Expert in background job processing with Bull/BullMQ (Redis), Celery, and cloud queues.

    244 GitHub stars~3.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Competitive Cartographer

    curiositech/some_claude_skills

    Strategic analyst that maps competitive landscapes, identifies white space opportunities, and provides positioning recommendations.

    244 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Computer Vision Pipeline

    curiositech/some_claude_skills

    Build production computer vision pipelines for object detection, tracking, and video analysis.

    244 GitHub stars~4k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about 2026 Legal Research Agent

What does 2026 Legal Research Agent do?

Expert legal research agent for finding and scraping expungement data state by state. 2026 Legal Research Agent is an agent skill from curiositech/some_claude_skills. Expert legal research agent for finding and scraping expungement data state by state.

When should I use 2026 Legal Research Agent?

2026 Legal Research Agent fits situations like: tasks that involve Web scraping; tasks that involve Legal research.

How do I install 2026 Legal Research Agent in Claude Code?

Run `npx skills add curiositech/some_claude_skills --skill 2026-legal-research-agent -a claude-code`. Or copy the skill folder (.claude/skills/2026-legal-research-agent in curiositech/some_claude_skills) into .claude/skills/2026-legal-research-agent in your project. Claude Code loads it when a task matches its description.

How do I install 2026 Legal Research Agent in Codex?

Run `npx skills add curiositech/some_claude_skills --skill 2026-legal-research-agent -a codex`. Or copy the skill folder (.claude/skills/2026-legal-research-agent in curiositech/some_claude_skills) into .agents/skills/2026-legal-research-agent in your project. Codex loads it when a task matches its description.

Can I use 2026 Legal Research Agent in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add curiositech/some_claude_skills --skill 2026-legal-research-agent -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/2026-legal-research-agent, .gemini/skills/2026-legal-research-agent, .github/skills/2026-legal-research-agent and .opencode/skills/2026-legal-research-agent in your project.

What does 2026 Legal Research Agent need to run?

Going by SKILL.md and its folder, 2026 Legal Research Agent needs TypeScript for the scripts in its folder, the command-line tools its instructions call (npx) and credentials named FIRECRAWL_API_KEY. Our summary lists: Node.js; A credential in FIRECRAWL_API_KEY.

Does 2026 Legal Research Agent access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is 2026 Legal Research Agent safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does 2026 Legal Research Agent use?

2026 Legal Research Agent is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does 2026 Legal Research Agent use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5k tokens, read only when the agent opens those files.

What are the alternatives to 2026 Legal Research Agent?

Skills that share tags, products or a category with 2026 Legal Research Agent: Firecrawl Build Onboarding (firecrawl/firecrawl, 190k stars), Firecrawl Page Scrape Integration (firecrawl/firecrawl, 190k stars), Firecrawl App Integration (firecrawl/firecrawl, 190k stars) and Firecrawl Interact Integration (firecrawl/firecrawl, 190k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains 2026 Legal Research Agent?

curiositech (a GitHub organization) maintains it in curiositech/some_claude_skills, which has 244 GitHub stars. The repository holds 95 skills in this directory. The repository was last updated on September 6, 2026.

Source: curiositech/some_claude_skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.