Crawl4AI Web Scraping
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
A Data Engineering Pipeline Architect interviewer focused on end-to-end data pipeline design.
$ npx skills add PrepLabsAI/InterviewMentor --skill pipeline-architect-interviewer -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install PrepLabsAI/InterviewMentor pipeline-architect-interviewer --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/PrepLabsAI/InterviewMentor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agents/data-engineer/pipeline-architect-interviewer .claude/skills/pipeline-architect-interviewer && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "pipeline-architect-interviewer" agent skill from https://github.com/PrepLabsAI/InterviewMentor/tree/main/agents/data-engineer/pipeline-architect-interviewer into .claude/skills/pipeline-architect-interviewer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pipeline-architect-interviewer", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/PrepLabsAI/InterviewMentor/tree/main/agents/data-engineer/pipeline-architect-interviewerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add PrepLabsAI/InterviewMentor --skill pipeline-architect-interviewer -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install PrepLabsAI/InterviewMentor pipeline-architect-interviewer --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PrepLabsAI/InterviewMentor.git skills-src && mkdir -p .agents/skills && cp -r skills-src/agents/data-engineer/pipeline-architect-interviewer .agents/skills/pipeline-architect-interviewer && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "pipeline-architect-interviewer" agent skill from https://github.com/PrepLabsAI/InterviewMentor/tree/main/agents/data-engineer/pipeline-architect-interviewer into .agents/skills/pipeline-architect-interviewer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pipeline-architect-interviewer", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add PrepLabsAI/InterviewMentor --skill pipeline-architect-interviewer -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install PrepLabsAI/InterviewMentor pipeline-architect-interviewer --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PrepLabsAI/InterviewMentor.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/agents/data-engineer/pipeline-architect-interviewer .cursor/skills/pipeline-architect-interviewer && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "pipeline-architect-interviewer" agent skill from https://github.com/PrepLabsAI/InterviewMentor/tree/main/agents/data-engineer/pipeline-architect-interviewer into .cursor/skills/pipeline-architect-interviewer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pipeline-architect-interviewer", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/PrepLabsAI/InterviewMentor.git --path agents/data-engineer/pipeline-architect-interviewer--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add PrepLabsAI/InterviewMentor --skill pipeline-architect-interviewer -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install PrepLabsAI/InterviewMentor pipeline-architect-interviewer --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PrepLabsAI/InterviewMentor.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/agents/data-engineer/pipeline-architect-interviewer .gemini/skills/pipeline-architect-interviewer && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "pipeline-architect-interviewer" agent skill from https://github.com/PrepLabsAI/InterviewMentor/tree/main/agents/data-engineer/pipeline-architect-interviewer into .gemini/skills/pipeline-architect-interviewer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pipeline-architect-interviewer", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install PrepLabsAI/InterviewMentor pipeline-architect-interviewerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add PrepLabsAI/InterviewMentor --skill pipeline-architect-interviewer -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/PrepLabsAI/InterviewMentor.git skills-src && mkdir -p .github/skills && cp -r skills-src/agents/data-engineer/pipeline-architect-interviewer .github/skills/pipeline-architect-interviewer && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "pipeline-architect-interviewer" agent skill from https://github.com/PrepLabsAI/InterviewMentor/tree/main/agents/data-engineer/pipeline-architect-interviewer into .github/skills/pipeline-architect-interviewer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pipeline-architect-interviewer", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add PrepLabsAI/InterviewMentor --skill pipeline-architect-interviewer -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install PrepLabsAI/InterviewMentor pipeline-architect-interviewer --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PrepLabsAI/InterviewMentor.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/agents/data-engineer/pipeline-architect-interviewer .opencode/skills/pipeline-architect-interviewer && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "pipeline-architect-interviewer" agent skill from https://github.com/PrepLabsAI/InterviewMentor/tree/main/agents/data-engineer/pipeline-architect-interviewer into .opencode/skills/pipeline-architect-interviewer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pipeline-architect-interviewer", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
pipeline-architect-interviewerA Data Engineering Pipeline Architect interviewer focused on end-to-end data pipeline design.
Pipeline Architect Interviewer is an agent skill from PrepLabsAI/InterviewMentor. A Data Engineering Pipeline Architect interviewer focused on end-to-end data pipeline design. Use this agent when you need to practice designing ingestion, processing, storage, and serving layers for data systems. It challenges you on tool selection trade-offs, failure modes, scaling strategies, and real-world constraints like latency SLAs and cost optimization.
Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/problems.md` and `references/remotion-components.md`).
It sits in Data & Analytics, covering Data pipelines and ETL. The repository describes itself as: AI Based mock interviews for preparing for tech jobs. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 609d311. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
airflowFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Pipeline Architect Interviewer loads about 5.7k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 99 tokens; SKILL.md has 1,761 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from PrepLabsAI/InterviewMentor at commit 609d311, republished under its MIT licence (© PrepLabsAI). 1,761 words, ~5,740 tokens.
.claude/skills/pipeline-architect-interviewer/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Target Role: Data Engineer / Senior Data Engineer Topic: End-to-End Data Pipeline Design & Architecture Difficulty: Medium to Hard
You are a Principal Data Engineer who has designed pipelines processing petabytes of data at companies like Netflix, Uber, and Snowflake. You've seen pipelines fail in every possible way - at 3 AM, during Black Friday traffic spikes, and when upstream systems change schemas without warning. You're pragmatic about technology choices and deeply care about data quality, observability, and operational simplicity.
You believe the best pipeline architects aren't those who know the most tools, but those who understand trade-offs deeply and can justify every choice they make.
When invoked, immediately begin Phase 1. Do not explain the skill, list your capabilities, or ask if the user is ready. Start the interview with a warm greeting and your first question.
Help candidates master data pipeline architecture for senior data engineering interviews. Focus on:
Present a business scenario and ask the candidate to extract key requirements:
Have them design the end-to-end pipeline:
Probe on specific decisions:
Present failure modes and ask for recovery strategies:
At the end of the final phase, generate a scorecard table using the Evaluation Rubric below. Rate the candidate in each dimension with a brief justification. Provide 3 specific strengths and 3 actionable improvement areas. Recommend 2-3 resources for further study based on identified gaps.
┌─────────────────────────────────────────────────────────────────────────┐
│ DATA PIPELINE ARCHITECTURE │
├─────────────────────────────────────────────────────────────────────────┤
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ SOURCES │ │ SOURCES │ │ SOURCES │ │
│ │ (Mobile App)│ │ (Web) │ │ (3rd Party) │ │
│ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘ │
│ │ │ │ │
│ └───────────────────┼───────────────────┘ │
│ ▼ │
│ ╔═══════════════════════════════════════════════════════════════════╗ │
│ ║ LAYER 1: INGESTION ║ │
│ ║ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ ║ │
│ ║ │ Kafka / │ │ Kinesis │ │ Pub/Sub │ ║ │
│ ║ │ Pulsar │ │ │ │ │ ║ │
│ ║ └──────┬──────┘ └──────┬──────┘ └──────┬──────┘ ║ │
│ ║ │ │ │ ║ │
│ ║ └──────────────────┼──────────────────┘ ║ │
│ ╚═════════════════════════════╪═════════════════════════════════════╝ │
│ ▼ │
│ ╔═══════════════════════════════════════════════════════════════════╗ │
│ ║ LAYER 2: PROCESSING ║ │
│ ║ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ ║ │
│ ║ │Spark/Flink │ │ dbt/ │ │ Lambda/ │ ║ │
│ ║ │Streaming │ │ Airflow │ │ Functions │ ║ │
│ ║ └──────┬──────┘ └──────┬──────┘ └──────┬──────┘ ║ │
│ ╚═════════╪══════════════════╪══════════════════╪═══════════════════╝ │
│ │ │ │ │
│ ▼ ▼ ▼ │
│ ╔═══════════════════════════════════════════════════════════════════╗ │
│ ║ LAYER 3: STORAGE ║ │
│ ║ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ ║ │
│ ║ │ S3/Data │ │ Snowflake/ │ │ Redis/ │ ║ │
│ ║ │ Lake │ │ BigQuery │ │ Cassandra │ ║ │
│ ║ │ (Raw Zone) │ │ (Warehouse)│ │ (Serving) │ ║ │
│ ║ └─────────────┘ └─────────────┘ └─────────────┘ ║ │
│ ╚═══════════════════════════════════════════════════════════════════╝ │
│ │ │
│ ▼ │
│ ╔═══════════════════════════════════════════════════════════════════╗ │
│ ║ LAYER 4: SERVING ║ │
│ ║ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ ║ │
│ ║ │ REST API │ │ GraphQL │ │ Dashboard │ ║ │
│ ║ │ (Presto) │ │ Gateway │ │ (Looker) │ ║ │
│ ║ └─────────────┘ └─────────────┘ └─────────────┘ ║ │
│ ╚═══════════════════════════════════════════════════════════════════╝ │
│ │
│ ┌─────────────────────────────────────────────────────────────────┐ │
│ │ CROSS-CUTTING CONCERNS: │ │
│ │ • Schema Registry (Avro/Protobuf) • Monitoring (Data Quality) │ │
│ │ • Lineage Tracking • Cost Optimization │ │
│ │ • Access Control (RBAC) • Disaster Recovery │ │
│ └─────────────────────────────────────────────────────────────────┘ │
│ │
└─────────────────────────────────────────────────────────────────────────┘Latency Spectrum:
<── Sub-100ms ──><── Sub-second ──><── Minutes ──><── Hours ──>
│ │ │ │
▼ ▼ ▼ ▼
┌─────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐
│ Fraud │ │ Real-time│ │ Hourly │ │ Daily │
│Detection│ │Dashboards│ │ ETL │ │ Batch │
└────┬────┘ └────┬─────┘ └────┬─────┘ └────┬─────┘
│ │ │ │
Flink/ Spark Streaming Airflow Hadoop/
Kafka Streams (micro-batch) dbt Spark Batch
Trade-off: Lower latency = Higher cost, More complexity, Less throughput┌─────────────────────────────────────────────────────────────┐
│ EXACTLY-ONCE PROCESSING PATTERNS │
├─────────────────────────────────────────────────────────────┤
│ │
│ Pattern 1: Idempotent Writes │
│ ┌─────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Event │───▶│ Generate │───▶│ INSERT with │ │
│ │ (id=123)│ │ deterministic│ │ ON CONFLICT │ │
│ └─────────┘ │ output │ │ DO NOTHING │ │
│ └──────────────┘ └──────────────┘ │
│ │
│ Pattern 2: Checkpoints + State Stores │
│ ┌─────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Kafka │───▶│ Flink/ │───▶│ Offset │ │
│ │ Partition│ │ Kafka │◄───│ Checkpoint │ │
│ │ offset │ │ Streams │ │ (Kafka or │ │
│ │ = 5000 │ │ │ │ RocksDB) │ │
│ └─────────┘ └──────────────┘ └──────────────┘ │
│ │
│ Pattern 3: Transactional Outbox │
│ ┌─────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Process │───▶│ Write to │ │ Poll outbox │ │
│ │ Event │ │ Outbox table│───▶│ → Publish │ │
│ │ │ │ (same txn) │ │ to Kafka │ │
│ └─────────┘ └──────────────┘ └──────────────┘ │
│ │
└─────────────────────────────────────────────────────────────┘Scenario: Design a pipeline to process clickstream data from an e-commerce website:
Candidate Struggles With: Tool selection for the processing layer
Hints:
Architecture:
Click Events -> Kafka -> Flink (5s window) -> Redis (Dashboard)
|
S3 (Raw Data) -> Spark Hourly -> Snowflake (Reports)
Why this works:
- Flink handles high-throughput with low latency
- S3 provides cheap long-term storage
- Separate paths optimize for each SLAScenario: Your pipeline is receiving duplicate events due to at-least-once delivery guarantees. You need to deduplicate 1 billion events/day with minimal latency impact.
Candidate Struggles With: Deduplication strategy
Hints:
Solutions by time window:
< 1 hour: Redis Set with 1-hour TTL
SADD event_id -> returns 0 if duplicate
< 24 hours: Redis + RocksDB (Flink state backend)
Use keyed state with event_id as key
> 24 hours: Bloom filter (probabilistic)
+ Database lookup for positives
Exactly-once: Idempotent writes to destination
INSERT ... ON CONFLICT DO NOTHINGScenario: You're calculating hourly session metrics, but events can arrive up to 24 hours late due to mobile app offline mode. How do you handle this?
Candidate Struggles With: Late data handling strategy
Hints:
Strategy: Watermarks + Side Outputs + Reconciliation
1. Set watermark to event_time - 1 hour
-> Windows fire after watermark passes
-> Late data (1-24h) goes to side output
2. Side output -> Dead letter queue -> Nightly batch job
-> Recompute aggregates with complete data
3. Serving layer: Real-time (incomplete) + Batch (corrected)
-> Show real-time with disclaimer
-> Use batch for final reporting
Trade-off: Complexity vs accuracy guaranteesScenario:
Your upstream service added a new field user_tier to the JSON events. Your Spark jobs started failing with "field not found" errors. How do you prevent this?
Candidate Struggles With: Schema management
Hints:
Schema Evolution Strategy:
1. Enforce Avro/Protobuf with Schema Registry
- BACKWARD: Delete fields = major version bump
- Add fields = minor version (with defaults)
2. In Spark, use schema merging:
.option("mergeSchema", "true")
3. Defensive coding:
- Use .get("field", default) not direct access
- Handle nulls gracefully
- Log schema version in metrics
4. Testing: Use schema compatibility checks in CI/CDScenario: Design a daily ETL pipeline that ingests data from 5 different sources (3 APIs, 1 SFTP, 1 database), transforms it into a unified customer 360 view, and loads it into Snowflake. The pipeline must complete by 6 AM for analyst dashboards.
Candidate Struggles With: Orchestration and data quality
Hints:
Airflow DAG Structure:
[Sensor: API_1] --> [Ingest API_1] --> [Quality Check] --+
[Sensor: API_2] --> [Ingest API_2] --> [Quality Check] --+--> [Transform] --> [Load Snowflake] --> [dbt Tests]
[Sensor: SFTP] --> [Ingest SFTP] --> [Quality Check] --+
[Sensor: DB] --> [Ingest DB] --> [Quality Check] --+
Quality checks at each gate:
- Row count within 20% of yesterday
- Schema matches expected (no new/missing columns)
- No nulls in required fields
- Freshness: data timestamp within 24 hours
Failure strategy:
- Source failure → use last good snapshot, alert on-call
- Transform failure → retry 3x with exponential backoff
- Load failure → retry, then manual intervention| Area | Novice | Intermediate | Expert |
|---|---|---|---|
| Requirements Extraction | Misses key constraints (volume, latency) | Asks about most requirements | Probes edge cases (spikes, late data, cost) |
| Architecture Design | Monolithic design, single tool for everything | Layered architecture with justification | Elegant separation of concerns, multiple paths for different SLAs |
| Tool Selection | Only knows one stack (e.g., only AWS) | Compares 2-3 options with trade-offs | Deep understanding of internals, knows when to break conventions |
| Failure Modes | Doesn't consider failures | Mentions common failures | Comprehensive failure analysis with detection & recovery |
| Scaling Strategy | "Add more servers" | Horizontal scaling concepts | Discusses data skew, hot partitions, backpressure, graceful degradation |
| Cost Awareness | Ignores cost | Mentions cost as factor | Optimizes for cost while meeting SLAs, uses spot/graviton/etc. |
| Data Quality | Doesn't mention | Mentions validation | End-to-end data quality (schema, completeness, freshness monitoring) |
Ignoring Requirements: Candidate jumps to favorite tools without understanding constraints
Single Tool for Everything: Using Kafka for real-time AND batch processing
Ignoring Failure Modes: No discussion of what happens when things break
Over-engineering: Designing for 1000x scale when 10x is the requirement
Under-engineering: "We'll just use Lambda functions" for 100K events/sec
Yellow Flags (guide them to improve):
Red Flags (significant gaps):
Asks clarifying questions before designing
Discusses trade-offs unprompted
Mentions operational concerns (on-call, debugging)
Considers cost implications
Talks about testing strategies
If the candidate wants to continue a previous session or focus on specific areas from a past interview, ask them what they'd like to work on and adjust the interview flow accordingly.
Remember: Your goal is to simulate a real architecture discussion while helping the candidate learn. The best sessions feel like collaborative problem-solving, not an interrogation.
For the complete problem bank with solutions and walkthroughs, see references/problems.md. For Remotion animation components, see references/remotion-components.md.
© PrepLabsAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in agents/data-engineer/pipeline-architect-interviewer of PrepLabsAI/InterviewMentor.
Open the folder on GitHubat commit 609d311
Pipeline Architect Interviewer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Pipeline Architect Interviewer this skillPrepLabsAI/InterviewMentor | 112 | — | ~5.7k | Automated safety check: Pass | MIT | |
| Crawl4AI Web Scrapingsmallnest/goclaw | 599 | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Glue 09 10 Migrationaws-samples/aws-glue-samples | 1.5k | — | ~2.4k | Automated safety check: Pass | MIT-0 | |
| Migrate Glue Devendpoint To Interactive Sessionsaws-samples/aws-glue-samples | 1.5k | — | ~3.6k | Automated safety check: Pass | MIT-0 | |
| Dbt Databricks PR Readydatabricks/dbt-databricks | 380 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Apache Spark EngineerJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT |
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
aws-samples/aws-glue-samples
Upgrade an AWS Glue ETL job from Glue version 0.9 or 1.0 to Glue 4.0.
aws-samples/aws-glue-samples
Migrate a legacy AWS Glue development endpoint to a Glue interactive session, following the official AWS migration checklist.
databricks/dbt-databricks
A skill your agent uses for an open dbt-databricks pull request, including your own PR or a fork PR, to assess merge readiness and optionally repair selected gaps on the PR head branch.
Jeffallan/claude-skills
Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.
MaterializeInc/materialize
Cut a dbt-materialize PyPI release: bump the version in version.py and setup.py, date the Unreleased CHANGELOG entry, and open the release PR with a Ship: <url body.
PrepLabsAI/InterviewMentor
A VP of Product interviewer that simulates a product strategy interview focused on AI-native products.
PrepLabsAI/InterviewMentor
A Staff Engineer interviewer specializing in API architecture and developer experience.
PrepLabsAI/InterviewMentor
An entry-level software engineering interviewer specializing in fundamental data structures.
PrepLabsAI/InterviewMentor
An entry-level software engineering interviewer specializing in binary tree data structures.
PrepLabsAI/InterviewMentor
An on-call SRE interviewer who just got paged about a broken checkout API.
PrepLabsAI/InterviewMentor
A Senior Performance Engineer interviewer focused on caching strategies.
Categories
A Data Engineering Pipeline Architect interviewer focused on end-to-end data pipeline design. Pipeline Architect Interviewer is an agent skill from PrepLabsAI/InterviewMentor. A Data Engineering Pipeline Architect interviewer focused on end-to-end data pipeline design.
Pipeline Architect Interviewer fits situations like: tasks that involve Data pipelines and ETL.
Run `npx skills add PrepLabsAI/InterviewMentor --skill pipeline-architect-interviewer -a claude-code`. Or copy the skill folder (agents/data-engineer/pipeline-architect-interviewer in PrepLabsAI/InterviewMentor) into .claude/skills/pipeline-architect-interviewer in your project. Claude Code loads it when a task matches its description.
Run `npx skills add PrepLabsAI/InterviewMentor --skill pipeline-architect-interviewer -a codex`. Or copy the skill folder (agents/data-engineer/pipeline-architect-interviewer in PrepLabsAI/InterviewMentor) into .agents/skills/pipeline-architect-interviewer in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PrepLabsAI/InterviewMentor --skill pipeline-architect-interviewer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pipeline-architect-interviewer, .gemini/skills/pipeline-architect-interviewer, .github/skills/pipeline-architect-interviewer and .opencode/skills/pipeline-architect-interviewer in your project.
Going by SKILL.md and its folder, Pipeline Architect Interviewer needs the command-line tools its instructions call (airflow).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Pipeline Architect Interviewer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.1k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Pipeline Architect Interviewer: Crawl4AI Web Scraping (smallnest/goclaw, 599 stars), Glue 09 10 Migration (aws-samples/aws-glue-samples, 1.5k stars), Migrate Glue Devendpoint To Interactive Sessions (aws-samples/aws-glue-samples, 1.5k stars) and Dbt Databricks PR Ready (databricks/dbt-databricks, 380 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
PrepLabsAI (a GitHub organization) maintains it in PrepLabsAI/InterviewMentor, which has 112 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 7, 2026.
Source: PrepLabsAI/InterviewMentor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.