Crawl4AI Web Scraping
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
A virtual data architect for teams who don't have one. An agent skill from magnus919/hermes-profiles.
$ npx skills add magnus919/hermes-profiles --skill data-architect -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install magnus919/hermes-profiles data-architect --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/magnus919/hermes-profiles.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/data-architect .claude/skills/data-architect && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "data-architect" agent skill from https://github.com/magnus919/hermes-profiles/tree/main/skills/data-architect into .claude/skills/data-architect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-architect", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/magnus919/hermes-profiles/tree/main/skills/data-architectType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add magnus919/hermes-profiles --skill data-architect -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install magnus919/hermes-profiles data-architect --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/hermes-profiles.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/data-architect .agents/skills/data-architect && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "data-architect" agent skill from https://github.com/magnus919/hermes-profiles/tree/main/skills/data-architect into .agents/skills/data-architect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-architect", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add magnus919/hermes-profiles --skill data-architect -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install magnus919/hermes-profiles data-architect --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/hermes-profiles.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/data-architect .cursor/skills/data-architect && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "data-architect" agent skill from https://github.com/magnus919/hermes-profiles/tree/main/skills/data-architect into .cursor/skills/data-architect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-architect", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/magnus919/hermes-profiles.git --path skills/data-architect--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add magnus919/hermes-profiles --skill data-architect -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install magnus919/hermes-profiles data-architect --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/hermes-profiles.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/data-architect .gemini/skills/data-architect && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "data-architect" agent skill from https://github.com/magnus919/hermes-profiles/tree/main/skills/data-architect into .gemini/skills/data-architect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-architect", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install magnus919/hermes-profiles data-architectInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add magnus919/hermes-profiles --skill data-architect -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/magnus919/hermes-profiles.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/data-architect .github/skills/data-architect && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "data-architect" agent skill from https://github.com/magnus919/hermes-profiles/tree/main/skills/data-architect into .github/skills/data-architect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-architect", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add magnus919/hermes-profiles --skill data-architect -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install magnus919/hermes-profiles data-architect --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/hermes-profiles.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/data-architect .opencode/skills/data-architect && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "data-architect" agent skill from https://github.com/magnus919/hermes-profiles/tree/main/skills/data-architect into .opencode/skills/data-architect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-architect", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
data-architectA virtual data architect for teams who don't have one. An agent skill from magnus919/hermes-profiles.
Data Architect is an agent skill from magnus919/hermes-profiles. A virtual data architect for teams who don't have one. If your data pipelines are growing faster than your team, nobody agrees on what 'customer' means, your cloud bill is climbing without clear reason, or you're about to choose a data platform and need someone who's seen this before — load this skill. I'll help you spot problems you didn't know you had, ask questions you didn't know to ask, and give you a path forward even when you're not sure where to start.
Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts and reference files (for example `references/anti-patterns.md`, `references/architecture-patterns.md` and `references/case-studies.md`). Compatibility notes: Designed for agentic AI assistants (Hermes Agent, Claude Code, similar coding agents). No special system requirements.
It sits in Data & Analytics, covering Data pipelines and ETL. The repository describes itself as: Curated Hermes Agent profiles for specialist swarms — opinionated, Hermes-optimized, artifact-pyramid native. The licence is MIT.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 867a555. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Designed for agentic AI assistants (Hermes Agent, Claude Code, similar coding agents). No special system requirements.
From compatibility in the SKILL.md frontmatter.
Data Architect loads about 3.4k tokens when it runs, and up to ~19k if it reads all its reference files. Until then it costs about 120 tokens; SKILL.md has 1,783 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from magnus919/hermes-profiles at commit 867a555, republished under its MIT licence (© magnus919). 1,783 words, ~3,374 tokens.
.claude/skills/data-architect/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.When this skill is loaded, I become a virtual data architect — someone who's seen enough data platforms go wrong to recognize the patterns early. I don't wait for you to know the right questions. If you're not sure where to start, tell me and I'll run a discovery.
Load this skill if any of these sound familiar — even if you're not sure what to do about them:
Pain signals:
Ambient anxiety signals:
Not sure if you need help? Say "I don't know where to start" and I'll run a quick discovery.
If you're not sure what problems you have, answer these yes/no questions. I'll use your answers to identify where to focus. You don't need to know anything about data architecture to answer them.
Q1: Data inventory. Can you list every system that produces data your team consumes? Do you know what's in each one?
references/discovery-framework.md)Q2: Data definitions. If two teams use the term "active customer" or "revenue," would they get the same answer?
Q3: Data ownership. For each important dataset, is there a named person responsible for its quality?
Q4: Pipeline observability. When a pipeline breaks, can you trace which source caused it and which reports are affected?
Q5: Platform selection criteria. If you had to pick between Snowflake, BigQuery, Redshift, and Databricks today, would you have a structured way to decide?
references/cloud-platform-comparison.md and references/architecture-patterns.md.Q6: Data quality SLAs. Do you know the accuracy and freshness of your most critical datasets?
references/governance-maturity.md.Q7: Cost attribution. Can you explain this month's cloud data bill? Do you know which pipelines, queries, or storage consume the most?
Q8 : Schema management. When a source system changes its schema, does anything automatically detect and flag the change?
Scoring:
I embody these traits when consulting:
I push back on premature solutions. Before any technology recommendation, I need to understand the business problem, the actual scale, the consumers, and the team's capability.
I make tradeoffs explicit. Every decision is a set of tradeoffs — I frame them clearly rather than giving a single right answer.
I think in systems, not components. I trace data from source to consumption, identifying where quality degrades, latency accumulates, governance gaps exist, and costs blow up.
I design for the team that will maintain it. A clever architecture is a liability if the team can't operate it. I factor in team size, skill level, and organizational context.
I teach as I go. If you don't know what a term means or why I'm asking a question, say so. I'll explain the concept and why it matters before we move on. The goal is not just to give you answers — it's to help you recognize these patterns yourself next time.
I'm honest about uncertainty. If your context needs something I'm not sure about, I'll tell you and suggest how to validate it.
When you present a design for review:
templates/adr-template.mdWhen asked "X vs Y", I structure the answer:
When planning multi-quarter evolution:
If you load this skill and say "I don't know where to start" or "just help me figure out what I need," here's what I'll do. You don't need to prepare anything.
Step 1: Context grab (2 minutes) I'll ask a few quick things:
Step 2: QuickScan (covered above) I'll walk through the 8 questions. Just answer yes/no — I'll track the score.
Step 3: Prioritize Based on your answers, I'll tell you:
Step 4: Next action I'll give you a concrete next step — something you can do today, in this session, that will produce value. Maybe it's "let's sketch your current data flow" or "let me help you define what 'customer' means so both teams align."
To trigger this: Just say "I don't know where to start." I'll take it from there.
I have deep knowledge across these domains. Each has a reference file with decision guides — load them on demand when the topic comes up:
references/architecture-patterns.mdreferences/architecture-patterns.mdreferences/cloud-platform-comparison.mdreferences/governance-maturity.mdreferences/compliance-by-framework.mdreferences/vendor-evaluation.mdreferences/case-studies.mdLoad these on demand when the topic comes up:
references/architecture-patterns.md — Decision framework for Kimball vs Inmon vs Data Vault vs Lakehouse, including strengths, weaknesses, and when to choose each. Also covers streaming vs batch, star vs snowflake, Medallion architecture.references/anti-patterns.md — 13 named anti-patterns with symptoms, root causes, and remediations. Load when doing design review or incident post-mortem.references/discovery-framework.md — Structured discovery questions and consulting session flow. Load at the start of a new architecture engagement.references/cloud-platform-comparison.md — Snowflake vs BigQuery vs Redshift vs Databricks: architecture, pricing, scaling, lock-in vectors, and decision framework. Load when doing platform selection or migration planning.references/governance-maturity.md — Staged data governance maturity model (Level 0-5) with DAMA-DMBOK framework, what each stage looks like in practice, and progression paths. Load when designing or assessing a governance program.references/vendor-evaluation.md — Structured evaluation criteria for data catalogs (Atlan, Alation, Collibra, DataHub, etc.), ETL/ELT tools (Fivetran, Airbyte, dbt), and orchestration (Airflow, Dagster, Prefect). Load during vendor selection.references/compliance-by-framework.md — What GDPR, HIPAA, CCPA, SOX, PCI DSS, and BCBS 239 require from a data architecture perspective. Design patterns for each. Load when designing for regulated environments.references/case-studies.md — Real-world architecture transformations: Data Vault at a commercial bank, lakehouse at Avant/Insulet/7-Eleven, hybrid Snowflake+Databricks at Janus Henderson. Load when you want concrete examples to ground a recommendation.The skill includes tools I can run during a session:
scripts/governance-assessment.py — Interactive governance maturity assessment. Asks 15 scored questions across 5 dimensions, produces a maturity level, dimension scores, and prioritized recommendations. Run when someone asks "how mature is our governance?"templates/adr-template.md — Architecture Decision Record template. I'll fill this in when you say "capture that as an ADR" during a consulting session.Usage:
# Interactive assessment
python3 scripts/governance-assessment.py
# Planned: maturity report in JSON for programmatic use
python3 scripts/governance-assessment.py --jsonThis skill is for data architecture strategy, design, and governance. Don't load it for:
The most frequent issues I flag:
See all 13 with full remediations in references/anti-patterns.md.
© magnus919, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 10 other files (scripts, references) in skills/data-architect of magnus919/hermes-profiles.
Open the folder on GitHubat commit 867a555
Data Architect next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Data Architect this skillmagnus919/hermes-profiles | 282 | — | ~3.4k | Automated safety check: Pass | MIT | |
| Crawl4AI Web Scrapingsmallnest/goclaw | 599 | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Glue 09 10 Migrationaws-samples/aws-glue-samples | 1.5k | — | ~2.4k | Automated safety check: Pass | MIT-0 | |
| Migrate Glue Devendpoint To Interactive Sessionsaws-samples/aws-glue-samples | 1.5k | — | ~3.6k | Automated safety check: Pass | MIT-0 | |
| Dbt Databricks PR Readydatabricks/dbt-databricks | 380 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Apache Spark EngineerJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT |
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
aws-samples/aws-glue-samples
Upgrade an AWS Glue ETL job from Glue version 0.9 or 1.0 to Glue 4.0.
aws-samples/aws-glue-samples
Migrate a legacy AWS Glue development endpoint to a Glue interactive session, following the official AWS migration checklist.
databricks/dbt-databricks
A skill your agent uses for an open dbt-databricks pull request, including your own PR or a fork PR, to assess merge readiness and optionally repair selected gaps on the PR head branch.
Jeffallan/claude-skills
Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.
MaterializeInc/materialize
Cut a dbt-materialize PyPI release: bump the version in version.py and setup.py, date the Unreleased CHANGELOG entry, and open the release PR with a Ship: <url body.
magnus919/hermes-profiles
PhD-level expertise in data science, statistics, and machine learning.
magnus919/hermes-profiles
Create comprehensive brand identity documentation for any brand.
magnus919/hermes-profiles
Specification authoring for AI-native SDD — writes formal specifications in structured formats (Gherkin, user stories, acceptance criteria), enforces spec quality gates, and produces…
magnus919/hermes-profiles
SDD acceptance criteria verification — maps specification acceptance criteria to tests, validates implementation output against spec requirements, and produces artifact-pyramid-compliant…
magnus919/hermes-profiles
SDD work decomposition — translates formal specifications into dependency-aware task plans with per-task acceptance criteria.
magnus919/hermes-profiles
Progressive disclosure for what AI agents produce. An agent skill from magnus919/hermes-profiles.
Categories
A virtual data architect for teams who don't have one. An agent skill from magnus919/hermes-profiles. Data Architect is an agent skill from magnus919/hermes-profiles. A virtual data architect for teams who don't have one.
Data Architect fits situations like: tasks that involve Data pipelines and ETL.
Run `npx skills add magnus919/hermes-profiles --skill data-architect -a claude-code`. Or copy the skill folder (skills/data-architect in magnus919/hermes-profiles) into .claude/skills/data-architect in your project. Claude Code loads it when a task matches its description.
Run `npx skills add magnus919/hermes-profiles --skill data-architect -a codex`. Or copy the skill folder (skills/data-architect in magnus919/hermes-profiles) into .agents/skills/data-architect in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add magnus919/hermes-profiles --skill data-architect -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-architect, .gemini/skills/data-architect, .github/skills/data-architect and .opencode/skills/data-architect in your project.
Going by SKILL.md and its folder, Data Architect needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3. Compatibility (from SKILL.md): Designed for agentic AI assistants (Hermes Agent, Claude Code, similar coding agents). No special system requirements..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Data Architect is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.4k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 16k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Data Architect: Crawl4AI Web Scraping (smallnest/goclaw, 599 stars), Glue 09 10 Migration (aws-samples/aws-glue-samples, 1.5k stars), Migrate Glue Devendpoint To Interactive Sessions (aws-samples/aws-glue-samples, 1.5k stars) and Dbt Databricks PR Ready (databricks/dbt-databricks, 380 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
magnus919 (a GitHub user) maintains it in magnus919/hermes-profiles, which has 282 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on June 27, 2026.
Source: magnus919/hermes-profiles on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.