Clickhouse Io
hellangleZ/burn-in-cceverywhere-ralph
ClickHouse database patterns, query optimization, analytics, and data engineering best practices for high-performance analytical workloads.
Expert guidance for working with the Apache Spark Catalyst query optimisation framework.
$ npx skills add aehrc/pathling --skill spark-catalyst -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install aehrc/pathling spark-catalyst --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/aehrc/pathling.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/spark-catalyst .claude/skills/spark-catalyst && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "spark-catalyst" agent skill from https://github.com/aehrc/pathling/tree/main/.claude/skills/spark-catalyst into .claude/skills/spark-catalyst/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-catalyst", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/aehrc/pathling/tree/main/.claude/skills/spark-catalystType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add aehrc/pathling --skill spark-catalyst -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install aehrc/pathling spark-catalyst --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aehrc/pathling.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/spark-catalyst .agents/skills/spark-catalyst && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "spark-catalyst" agent skill from https://github.com/aehrc/pathling/tree/main/.claude/skills/spark-catalyst into .agents/skills/spark-catalyst/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-catalyst", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aehrc/pathling --skill spark-catalyst -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install aehrc/pathling spark-catalyst --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aehrc/pathling.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/spark-catalyst .cursor/skills/spark-catalyst && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "spark-catalyst" agent skill from https://github.com/aehrc/pathling/tree/main/.claude/skills/spark-catalyst into .cursor/skills/spark-catalyst/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-catalyst", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/aehrc/pathling.git --path .claude/skills/spark-catalyst--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add aehrc/pathling --skill spark-catalyst -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install aehrc/pathling spark-catalyst --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aehrc/pathling.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/spark-catalyst .gemini/skills/spark-catalyst && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "spark-catalyst" agent skill from https://github.com/aehrc/pathling/tree/main/.claude/skills/spark-catalyst into .gemini/skills/spark-catalyst/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-catalyst", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install aehrc/pathling spark-catalystInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add aehrc/pathling --skill spark-catalyst -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/aehrc/pathling.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/spark-catalyst .github/skills/spark-catalyst && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "spark-catalyst" agent skill from https://github.com/aehrc/pathling/tree/main/.claude/skills/spark-catalyst into .github/skills/spark-catalyst/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-catalyst", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aehrc/pathling --skill spark-catalyst -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install aehrc/pathling spark-catalyst --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aehrc/pathling.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/spark-catalyst .opencode/skills/spark-catalyst && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "spark-catalyst" agent skill from https://github.com/aehrc/pathling/tree/main/.claude/skills/spark-catalyst into .opencode/skills/spark-catalyst/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-catalyst", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
spark-catalystExpert guidance for working with the Apache Spark Catalyst query optimisation framework.
Spark Catalyst is an agent skill from aehrc/pathling. Expert guidance for working with the Apache Spark Catalyst query optimisation framework. Use this skill when working with Spark SQL internals, creating custom expressions, implementing query optimisations, working with logical/physical plans, or extending Catalyst. Trigger keywords include "catalyst", "spark sql", "expression", "logical plan", "physical plan", "tree node", "query optimisation", "rule executor", "analyzer", "optimizer", "code generation".
Its SKILL.md is about 5.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Databases, covering Query optimization. It works with Apache Spark. The repository describes itself as: Tools that make it easier to use FHIR and clinical terminology within data analytics, built on Apache Spark. The licence is Apache-2.0.
10 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 56a3b4a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are scala).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Spark Catalyst loads about 5.4k tokens when it runs. Until then it costs about 118 tokens; SKILL.md has 759 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from aehrc/pathling at commit 56a3b4a, republished under its Apache-2.0 licence (© aehrc). 759 words, ~5,398 tokens.
.claude/skills/spark-catalyst/SKILL.md (or your agent's skills folder).You are an expert in the Apache Spark Catalyst query optimisation framework. This skill provides comprehensive guidance on using the Catalyst API for query processing, optimisation, and code generation.
Apache Spark Catalyst is a query optimisation framework that powers Spark SQL and DataFrames. It provides:
The Catalyst module is located at sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/ and contains:
catalyst/
├── analysis/ # Query analysis and resolution
├── catalog/ # Catalog management and metadata
├── expressions/ # Expression definitions and evaluation
├── optimizer/ # Query optimisation rules
├── parser/ # SQL parsing
├── planning/ # Query planning strategies
├── plans/ # Query plan representations
│ ├── logical/ # Logical plan operators
│ └── physical/ # Physical plan operators
├── rules/ # Rule execution framework
├── trees/ # Tree node infrastructure
├── types/ # Data type utilities
└── util/ # Utility classesTreeNode is the base class for all tree structures in Catalyst, including expressions and query plans.
abstract class TreeNode[BaseType <: TreeNode[BaseType]] extends Product {
// Children of this node.
def children: Seq[BaseType]
// Transform this tree by applying a function to all nodes.
def transform(rule: PartialFunction[BaseType, BaseType]): BaseType
// Transform all nodes bottom-up.
def transformDown(rule: PartialFunction[BaseType, BaseType]): BaseType
// Transform all nodes top-down.
def transformUp(rule: PartialFunction[BaseType, BaseType]): BaseType
// Fast equality check.
def fastEquals(other: TreeNode[_]): Boolean
// Tree pattern matching for efficient traversal.
def treePatternBits: BitSet
}Bottom-up transformation:
plan.transformUp {
case Filter(condition, child) if isAlwaysTrue(condition) =>
child
}Top-down transformation:
plan.transformDown {
case Project(projectList, child) =>
// Transform project first, then children
optimiseProject(projectList, child)
}Collect nodes matching a pattern:
val filters = plan.collect {
case f @ Filter(_, _) => f
}Expression is the base class for all expression trees in Catalyst.
Literal, AttributeReference)Cast, Not, IsNull)Add, EqualTo, And)If, Substring)abstract class Expression extends TreeNode[Expression] {
// Can this expression be evaluated at query planning time?
def foldable: Boolean
// Does this expression always return the same result for fixed inputs?
def deterministic: Boolean
// Can this expression evaluate to null?
def nullable: Boolean
// Data type of the expression result.
def dataType: DataType
// Evaluate this expression given an input row.
def eval(input: InternalRow): Any
// Generate code for this expression.
def doGenCode(ctx: CodegenContext, ev: ExprCode): ExprCode
// Attributes referenced by this expression.
def references: AttributeSet
}Literal values:
Literal(42) // Integer literal
Literal("hello") // String literal
Literal(null, StringType) // Null literalAttribute references:
AttributeReference("name", StringType, nullable = true)()
AttributeReference("age", IntegerType, nullable = false)()Predicates:
EqualTo(left, right) // left = right
GreaterThan(left, right) // left > right
LessThanOrEqual(left, right) // left <= right
And(left, right) // left AND right
Or(left, right) // left OR right
Not(child) // NOT childArithmetic:
Add(left, right) // left + right
Subtract(left, right) // left - right
Multiply(left, right) // left * right
Divide(left, right) // left / rightString operations:
Substring(str, pos, len) // substring(str, pos, len)
Upper(child) // upper(child)
Lower(child) // lower(child)
Concat(children) // concat(child1, child2, ...)Type conversion:
Cast(child, targetType) // cast(child as targetType)QueryPlan is the base class for both logical and physical query plans.
abstract class QueryPlan[PlanType <: QueryPlan[PlanType]]
extends TreeNode[PlanType] {
// Output schema of this plan node.
def output: Seq[Attribute]
// Set of output attributes.
def outputSet: AttributeSet
// Set of attributes from all children.
def inputSet: AttributeSet
// Attributes produced by this node.
def producedAttributes: AttributeSet
// Attributes referenced by expressions.
def references: AttributeSet
// Attributes referenced but not provided by children.
def missingInput: AttributeSet
// All expressions in this plan node.
def expressions: Seq[Expression]
// Transform expressions in this plan.
def transformExpressions(rule: PartialFunction[Expression, Expression]): PlanType
}Logical plans represent query semantics without execution strategy.
Data sources:
// Read from a relation.
LogicalRelation(relation, output, catalogTable)
// Local in-memory data.
LocalRelation(output, data)
// Empty relation.
EmptyRelation(output)Projections and filters:
// Project specific columns.
Project(projectList: Seq[NamedExpression], child: LogicalPlan)
// Filter rows.
Filter(condition: Expression, child: LogicalPlan)
// Select distinct rows.
Distinct(child: LogicalPlan)Aggregations:
// Group by and aggregate.
Aggregate(
groupingExpressions: Seq[Expression],
aggregateExpressions: Seq[NamedExpression],
child: LogicalPlan
)Joins:
// Join two relations.
Join(
left: LogicalPlan,
right: LogicalPlan,
joinType: JoinType,
condition: Option[Expression],
hint: JoinHint
)Sorting:
// Sort rows.
Sort(
order: Seq[SortOrder],
global: Boolean,
child: LogicalPlan
)Limits:
// Limit number of rows.
Limit(limitExpr: Expression, child: LogicalPlan)Set operations:
// Union of two relations.
Union(children: Seq[LogicalPlan])
// Intersection.
Intersect(left: LogicalPlan, right: LogicalPlan, isAll: Boolean)
// Difference.
Except(left: LogicalPlan, right: LogicalPlan, isAll: Boolean)InternalRow is the internal representation of a row in Catalyst.
abstract class InternalRow extends SpecializedGetters {
// Number of fields in this row.
def numFields: Int
// Check if field at ordinal is null.
def isNullAt(ordinal: Int): Boolean
// Get value at ordinal.
def get(ordinal: Int, dataType: DataType): Any
// Specialised getters for primitive types.
def getBoolean(ordinal: Int): Boolean
def getByte(ordinal: Int): Byte
def getShort(ordinal: Int): Short
def getInt(ordinal: Int): Int
def getLong(ordinal: Int): Long
def getFloat(ordinal: Int): Float
def getDouble(ordinal: Int): Double
def getDecimal(ordinal: Int, precision: Int, scale: Int): Decimal
def getUTF8String(ordinal: Int): UTF8String
// Update value at ordinal.
def update(ordinal: Int, value: Any): Unit
// Set field to null.
def setNullAt(ordinal: Int): Unit
// Create a copy of this row.
def copy(): InternalRow
// Convert to Scala sequence.
def toSeq(schema: StructType): Seq[Any]
}Rules define tree transformations, and RuleExecutor applies them in batches.
// Basic rule.
object MyOptimisationRule extends Rule[LogicalPlan] {
def apply(plan: LogicalPlan): LogicalPlan = plan.transformUp {
case Filter(condition, child) if isAlwaysTrue(condition) =>
child
}
}
// Configurable rule.
case class MyParameterisedRule(conf: SQLConf) extends Rule[LogicalPlan] {
def apply(plan: LogicalPlan): LogicalPlan = {
if (conf.myFeatureEnabled) {
optimisePlan(plan)
} else {
plan
}
}
}abstract class RuleExecutor[TreeType <: TreeNode[_]] {
// Define batches of rules to execute.
protected def batches: Seq[Batch]
// Execute all batches on the plan.
def execute(plan: TreeType): TreeType
}
// Batch execution strategies.
abstract class Strategy
case class Once extends Strategy
case class FixedPoint(maxIterations: Int) extends Strategyobject MyOptimiser extends RuleExecutor[LogicalPlan] {
val batches = Seq(
Batch("Normalisation", Once,
EliminateSubqueryAliases,
RemoveRedundantAliases
),
Batch("Operator Optimisation", FixedPoint(100),
PushDownPredicate,
ConstantFolding,
ColumnPruning
),
Batch("Join Reordering", Once,
CostBasedJoinReorder
)
)
}The Analyzer resolves unresolved logical plans by binding attributes, functions, and tables.
val analyzer = new Analyzer(catalogManager)
val analysedPlan = analyzer.execute(unresolvedPlan)The Optimizer transforms logical plans to improve query performance.
Predicate pushdown:
// Push filters below projections.
PushDownPredicate
// Push filters into join conditions.
PushPredicateThroughJoin
// Push filters to data sources.
PushDownPredicatesProjection pushdown:
// Eliminate unnecessary columns.
ColumnPruning
// Combine adjacent projections.
CollapseProjectConstant folding:
// Evaluate constant expressions.
ConstantFolding
// Simplify expressions.
SimplifyConditionals
SimplifyCastsJoin optimisation:
// Reorder joins for better performance.
CostBasedJoinReorder
// Eliminate redundant joins.
EliminateOuterJoinSubquery optimisation:
// Decorrelate correlated subqueries.
DecorrelateInnerQuery
// Merge scalar subqueries.
MergeScalarSubqueriesCatalyst generates optimised Java bytecode for query execution.
CodegenContext:
Manages code generation state including variable declarations, functions, and class structure.
ExprCode:
Represents generated code for an expression evaluation.
case class ExprCode(
code: Block, // Generated code block
isNull: ExprValue, // Variable for null check
value: ExprValue // Variable for result value
)trait CodegenFallback extends Expression {
// Fallback to interpreted evaluation.
protected def doGenCode(ctx: CodegenContext, ev: ExprCode): ExprCode = {
ctx.references += this
val objectTerm = ctx.addReferenceObj("expression", this)
ExprCode(
code = code"""
boolean ${ev.isNull} = true;
${CodeGenerator.javaType(dataType)} ${ev.value} =
${CodeGenerator.defaultValue(dataType)};
Object result = $objectTerm.eval(${ctx.INPUT_ROW});
if (result != null) {
${ev.isNull} = false;
${ev.value} = (${CodeGenerator.boxedType(dataType)}) result;
}
""",
isNull = ev.isNull,
value = ev.value
)
}
}// Generate unsafe projection from expressions.
val projection = GenerateUnsafeProjection.generate(expressions)
val result = projection(inputRow)
// Generate mutable projection.
val mutableProjection = GenerateMutableProjection.generate(expressions)
val outputRow = mutableProjection(inputRow)
// Generate safe projection.
val safeProjection = GenerateSafeProjection.generate(expressions)// Generate predicate for filter condition.
val predicate = GeneratePredicate.generate(condition)
val passes = predicate.eval(row)// Generate row comparator.
val ordering = GenerateOrdering.generate(sortOrders)
val compareResult = ordering.compare(row1, row2)BooleanType
ByteType
ShortType
IntegerType
LongType
FloatType
DoubleType
StringType
BinaryType
DateType
TimestampType
TimestampNTZType// Array type.
ArrayType(elementType: DataType, containsNull: Boolean)
// Map type.
MapType(keyType: DataType, valueType: DataType, valueContainsNull: Boolean)
// Struct type.
StructType(fields: Seq[StructField])
// Struct field.
StructField(name: String, dataType: DataType, nullable: Boolean, metadata: Metadata)DecimalType(precision: Int, scale: Int)
DecimalType.SYSTEM_DEFAULT // Decimal(38, 18)case class MyCustomFunction(child: Expression) extends UnaryExpression {
// Define output type.
override def dataType: DataType = StringType
// Can the result be null?
override def nullable: Boolean = child.nullable
// Evaluate the expression.
override def eval(input: InternalRow): Any = {
val value = child.eval(input)
if (value == null) {
null
} else {
// Custom logic here.
UTF8String.fromString(value.toString.toUpperCase)
}
}
// Generate code for this expression.
override protected def doGenCode(
ctx: CodegenContext,
ev: ExprCode): ExprCode = {
val childGen = child.genCode(ctx)
ev.copy(code = code"""
${childGen.code}
boolean ${ev.isNull} = ${childGen.isNull};
${CodeGenerator.javaType(dataType)} ${ev.value} =
${CodeGenerator.defaultValue(dataType)};
if (!${ev.isNull}) {
${ev.value} = UTF8String.fromString(
${childGen.value}.toString().toUpperCase());
}
""")
}
// Override for pretty printing.
override def prettyName: String = "my_custom_function"
}object EliminateRedundantCasts extends Rule[LogicalPlan] {
def apply(plan: LogicalPlan): LogicalPlan = plan.transformAllExpressions {
case Cast(child, dataType, _, _) if child.dataType == dataType =>
// Remove cast if types match.
child
}
}plan match {
case Project(projectList, child) =>
// Handle projection.
case Filter(condition, Project(projectList, child)) =>
// Handle filter over projection.
case Join(left, right, joinType, Some(condition), _) =>
// Handle join with condition.
case Aggregate(grouping, aggregates, child) =>
// Handle aggregation.
case _ =>
// Default case.
}// Find all attribute references.
val attributes = expr.collect {
case a: AttributeReference => a
}
// Find all subqueries.
val subqueries = expr.collect {
case s: SubqueryExpression => s
}
// Transform specific expression types.
val transformed = expr.transformUp {
case Add(Literal(0, _), right) => right
case Add(left, Literal(0, _)) => left
}// Resolve attribute by name.
def resolve(attrName: String, input: LogicalPlan): Option[Attribute] = {
input.output.find(_.name == attrName)
}
// Resolve with qualifier.
def resolveQualified(
qualifier: Seq[String],
attrName: String,
input: LogicalPlan): Option[Attribute] = {
input.output.find { attr =>
attr.qualifier.startsWith(qualifier) && attr.name == attrName
}
}// Create schema from attributes.
val schema = StructType(attributes.map { attr =>
StructField(attr.name, attr.dataType, attr.nullable, attr.metadata)
})
// Convert schema to attributes.
val attributes = schema.toAttributes
// Add column to schema.
val newSchema = schema.add("newColumn", StringType, nullable = true)
// Drop column from schema.
val reducedSchema = StructType(schema.filterNot(_.name == "dropColumn"))test("my custom function") {
val input = Literal("hello")
val expr = MyCustomFunction(input)
// Test evaluation.
assert(expr.eval(null) == UTF8String.fromString("HELLO"))
// Test properties.
assert(expr.dataType == StringType)
assert(expr.foldable)
assert(expr.deterministic)
}test("eliminate redundant casts") {
val plan = Project(
Seq(Alias(Cast(AttributeReference("x", IntegerType)(), IntegerType), "y")()),
testRelation
)
val optimised = EliminateRedundantCasts(plan)
val expected = Project(
Seq(Alias(AttributeReference("x", IntegerType)(), "y")()),
testRelation
)
comparePlans(optimised, expected)
}outputSet and references.fastEquals: For quick equality checks without deep comparison.doGenCode for custom expressions.sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/sql/catalyst/src/test/scala/org/apache/spark/sql/catalyst/The Catalyst framework provides:
Key extension points:
Expression)Rule[LogicalPlan])Rule[LogicalPlan])TableProvider)AggregateFunction)When working with Catalyst, focus on immutability, type safety, and efficient tree traversal patterns.
© aehrc, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/spark-catalyst of aehrc/pathling.
Open the folder on GitHubat commit 56a3b4a
Spark Catalyst next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Spark Catalyst this skillaehrc/pathling | 137 | — | ~5.4k | Automated safety check: Pass | Apache-2.0 | |
| Clickhouse IohellangleZ/burn-in-cceverywhere-ralph | 112 | 14 repos | ~2.5k | Automated safety check: Pass | None | |
| Altimate Data Warehouse DelegateAltimateAI/data-engineering-skills | 128 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Expensive Snowflake Query FinderAltimateAI/data-engineering-skills | 128 | — | ~662 | Automated safety check: Pass | MIT | |
| Tinybird Datafile RulesTryGhost/Ghost | 55k | — | ~417 | Automated safety check: Pass | MIT | |
| Optimizing Databricks SQLAltimateAI/data-engineering-skills | 128 | — | ~6.7k | Automated safety check: Pass | MIT |
hellangleZ/burn-in-cceverywhere-ralph
ClickHouse database patterns, query optimization, analytics, and data engineering best practices for high-performance analytical workloads.
AltimateAI/data-engineering-skills
Delegates dbt and warehouse tasks such as lineage, migrations and cost attribution to the altimate-code CLI agent and relays its answer back.
AltimateAI/data-engineering-skills
Ranks the costliest, slowest or heaviest-scanning Snowflake queries from query history and suggests how to optimize them.
TryGhost/Ghost
Rules for writing Tinybird datasources, pipes, endpoints and materialized views, with SQL constraints, optimization habits and deduplication patterns.
AltimateAI/data-engineering-skills
Analyze DBSQL queries, including SQL embedded in notebooks (spark.sql(...), %sql cells), for anti-patterns, lint issues, and performance problems, using Databricks-specific dialect and platform…
ynulihao/AgentSkillOS
Master SQL query optimization, indexing strategies, and EXPLAIN analysis to dramatically improve database performance and eliminate slow queries.
aehrc/pathling
Expert guidance for implementing FHIR RESTful API servers and clients following the HL7 FHIR specification.
aehrc/pathling
Expert guidance for implementing FHIR Bulk Data Access (Flat FHIR) following the HL7 specification.
aehrc/pathling
Expert guidance for using the Databricks CLI to manage Databricks workspaces, clusters, jobs, pipelines, Unity Catalog, SQL warehouses, serving endpoints, secrets, bundles, and all other Databricks…
aehrc/pathling
FHIR RESTful search specification expert with access to the official HL7 search specification text and the formal SearchParameter registry.
aehrc/pathling
Design and generate comprehensive FHIRPath test suites using input domain partitioning and Pathling's DSL test framework.
aehrc/pathling
Expert guidance for implementing FHIR servers using HAPI FHIR Plain Server framework.
Works with
Categories
Expert guidance for working with the Apache Spark Catalyst query optimisation framework. Spark Catalyst is an agent skill from aehrc/pathling. Expert guidance for working with the Apache Spark Catalyst query optimisation framework.
Spark Catalyst fits situations like: working with Spark SQL internals; creating custom expressions; implementing query optimisations; working with logical/physical plans.
Run `npx skills add aehrc/pathling --skill spark-catalyst -a claude-code`. Or copy the skill folder (.claude/skills/spark-catalyst in aehrc/pathling) into .claude/skills/spark-catalyst in your project. Claude Code loads it when a task matches its description.
Run `npx skills add aehrc/pathling --skill spark-catalyst -a codex`. Or copy the skill folder (.claude/skills/spark-catalyst in aehrc/pathling) into .agents/skills/spark-catalyst in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aehrc/pathling --skill spark-catalyst -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spark-catalyst, .gemini/skills/spark-catalyst, .github/skills/spark-catalyst and .opencode/skills/spark-catalyst in your project.
SKILL.md names no scripts, command-line tools or credentials: Spark Catalyst is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Spark Catalyst is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.4k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Spark Catalyst: Clickhouse Io (hellangleZ/burn-in-cceverywhere-ralph, 112 stars), Altimate Data Warehouse Delegate (AltimateAI/data-engineering-skills, 128 stars), Expensive Snowflake Query Finder (AltimateAI/data-engineering-skills, 128 stars) and Tinybird Datafile Rules (TryGhost/Ghost, 55k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
aehrc (a GitHub organization) maintains it in aehrc/pathling, which has 137 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 8, 2026.
Source: aehrc/pathling on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.