AI Data Engineering
ancoleman/ai-design-components
Data pipelines, feature stores, and embedding generation for AI/ML systems.
Migrates Airflow projects from airflow-ai-sdk to apache-airflow-providers-common-ai 0.4.0+.
$ npx skills add astronomer/agents --skill migrating-ai-sdk-to-common-ai -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install astronomer/agents migrating-ai-sdk-to-common-ai --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/astronomer/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/migrating-ai-sdk-to-common-ai .claude/skills/migrating-ai-sdk-to-common-ai && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "migrating-ai-sdk-to-common-ai" agent skill from https://github.com/astronomer/agents/tree/main/skills/migrating-ai-sdk-to-common-ai into .claude/skills/migrating-ai-sdk-to-common-ai/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "migrating-ai-sdk-to-common-ai", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/astronomer/agents/tree/main/skills/migrating-ai-sdk-to-common-aiType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add astronomer/agents --skill migrating-ai-sdk-to-common-ai -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install astronomer/agents migrating-ai-sdk-to-common-ai --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/astronomer/agents.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/migrating-ai-sdk-to-common-ai .agents/skills/migrating-ai-sdk-to-common-ai && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "migrating-ai-sdk-to-common-ai" agent skill from https://github.com/astronomer/agents/tree/main/skills/migrating-ai-sdk-to-common-ai into .agents/skills/migrating-ai-sdk-to-common-ai/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "migrating-ai-sdk-to-common-ai", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add astronomer/agents --skill migrating-ai-sdk-to-common-ai -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install astronomer/agents migrating-ai-sdk-to-common-ai --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/astronomer/agents.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/migrating-ai-sdk-to-common-ai .cursor/skills/migrating-ai-sdk-to-common-ai && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "migrating-ai-sdk-to-common-ai" agent skill from https://github.com/astronomer/agents/tree/main/skills/migrating-ai-sdk-to-common-ai into .cursor/skills/migrating-ai-sdk-to-common-ai/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "migrating-ai-sdk-to-common-ai", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/astronomer/agents.git --path skills/migrating-ai-sdk-to-common-ai--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add astronomer/agents --skill migrating-ai-sdk-to-common-ai -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install astronomer/agents migrating-ai-sdk-to-common-ai --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/astronomer/agents.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/migrating-ai-sdk-to-common-ai .gemini/skills/migrating-ai-sdk-to-common-ai && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "migrating-ai-sdk-to-common-ai" agent skill from https://github.com/astronomer/agents/tree/main/skills/migrating-ai-sdk-to-common-ai into .gemini/skills/migrating-ai-sdk-to-common-ai/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "migrating-ai-sdk-to-common-ai", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install astronomer/agents migrating-ai-sdk-to-common-aiInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add astronomer/agents --skill migrating-ai-sdk-to-common-ai -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/astronomer/agents.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/migrating-ai-sdk-to-common-ai .github/skills/migrating-ai-sdk-to-common-ai && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "migrating-ai-sdk-to-common-ai" agent skill from https://github.com/astronomer/agents/tree/main/skills/migrating-ai-sdk-to-common-ai into .github/skills/migrating-ai-sdk-to-common-ai/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "migrating-ai-sdk-to-common-ai", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add astronomer/agents --skill migrating-ai-sdk-to-common-ai -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install astronomer/agents migrating-ai-sdk-to-common-ai --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/astronomer/agents.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/migrating-ai-sdk-to-common-ai .opencode/skills/migrating-ai-sdk-to-common-ai && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "migrating-ai-sdk-to-common-ai" agent skill from https://github.com/astronomer/agents/tree/main/skills/migrating-ai-sdk-to-common-ai into .opencode/skills/migrating-ai-sdk-to-common-ai/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "migrating-ai-sdk-to-common-ai", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
migrating-ai-sdk-to-common-aiMigrates Airflow projects from airflow-ai-sdk to apache-airflow-providers-common-ai 0.4.0+.
Migrating AI SDK To Common AI is an agent skill from astronomer/agents. Migrates Airflow projects from airflow-ai-sdk to apache-airflow-providers-common-ai 0.4.0+. Use when replacing airflow-ai-sdk with the official Airflow AI provider - migrating LLM decorators (@task.llm, @task.agent, @task.llmbranch, @task.embed), switching from model strings/objects to connection-based LLM configuration, updating imports from airflowaisdk to the new provider, or upgrading an existing common-ai 0.1.x setup to 0.4.x (multimodal prompts, toolsets, embedding operators); also when common-ai provider…
Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Data pipelines and ETL and Embeddings. It works with Apache Airflow, Vercel AI SDK, LlamaIndex and SQL. The repository describes itself as: AI agent tooling for data engineering workflows. The licence is Apache-2.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 486ee63. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python, bash and yaml).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENAI_API_KEYANTHROPIC_API_KEYGOOGLE_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Migrating AI SDK To Common AI loads about 4.7k tokens when it runs. Until then it costs about 157 tokens; SKILL.md has 1,587 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
### Via environment variable (.env)ripts, non-Airflow services sharing the `.env`), leave them in place.] `pydanticai` connection configured in `.env` or connections.yamlAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from astronomer/agents at commit 486ee63, republished under its Apache-2.0 licence (© astronomer). 1,587 words, ~4,737 tokens.
.claude/skills/migrating-ai-sdk-to-common-ai/SKILL.md (or your agent's skills folder).This skill migrates Airflow projects from airflow-ai-sdk to apache-airflow-providers-common-ai (target 0.4.0+), the official Airflow AI provider built on PydanticAI. It also covers upgrading projects already on common-ai 0.1.x, since several capabilities (multimodal prompts, toolsets, embedding operators, structured-output XCom behavior) changed between 0.1.0 and 0.4.0.
CRITICAL: The new provider requires Airflow 3.0+ and (for 0.4.0) pydantic-ai-slim >= 1.71.0. The API surface has changed: LLM configuration moves from code (model strings/objects) to Airflow connections (
pydanticaitype). There is no@task.embedin the new provider; embeddings move to the LlamaIndex integration or a plain@task(see Step 3).
Use the Grep tool with the pattern below to inventory everything that needs to migrate:
airflow_ai_sdk|airflow-ai-sdk|ai_sdk|@task\.llm|@task\.agent|@task\.llm_branch|@task\.embedFrom the results, capture:
airflow-ai-sdk / airflow_ai_sdk@task.llm, @task.agent, @task.llm_branch, @task.embed"gpt-5", or OpenAIModel(...) objects)airflow_ai_sdk.BaseModel subclasses used as output_typeUse this inventory to drive the steps below.
Remove:
airflow-ai-sdk[openai]
# or any variant: airflow-ai-sdk[openai]==0.1.7, airflow-ai-sdk[anthropic], etc.Add:
apache-airflow-providers-common-ai[openai]>=0.4.0Use the latest available 0.x version unless the user has pinned a specific one. Available extras (0.4.0): [openai], [anthropic], [google], [bedrock], [llamaindex], [langchain], [mcp], plus file-format extras ([pdf], [docx], [parquet], [avro]) for DocumentLoaderOperator and [sql]/[common-sql] for the SQL operators. There are no [groq]/[mistral] extras; for those providers install the matching pydantic-ai-slim extra yourself.
Add [llamaindex] if the project migrates @task.embed to the LlamaIndexEmbeddingOperator (recommended, see Step 3). In that case sentence-transformers and torch can usually be removed, which shrinks the image considerably. Keep them only if the project stays on local sentence-transformers embeddings via plain @task.
The new provider uses an Airflow connection instead of model strings or objects in code.
Connection type: pydanticai
Default connection ID: pydanticai_default
AIRFLOW_CONN_PYDANTICAI_DEFAULT='{
"conn_type": "pydanticai",
"password": "<api-key>",
"extra": {
"model": "<provider>:<model-name>"
}
}'The model field uses provider:model format:
| Provider | Example model value |
|---|---|
| OpenAI | openai:gpt-5 |
| Anthropic | anthropic:claude-sonnet-4-20250514 |
google:gemini-2.5-pro | |
| Groq | groq:llama-3.3-70b-versatile |
| Mistral | mistral:mistral-large-latest |
| Bedrock | bedrock:us.anthropic.claude-sonnet-4-20250514-v1:0 |
Set host to the base URL:
AIRFLOW_CONN_PYDANTICAI_CORTEX='{
"conn_type": "pydanticai",
"password": "<api-key>",
"host": "https://my-endpoint.com/v1",
"extra": {
"model": "openai:<model-name>"
}
}'Use the openai: prefix for any OpenAI-compatible API, regardless of the actual provider.
The env var name determines the connection ID:
AIRFLOW_CONN_PYDANTICAI_DEFAULT creates pydanticai_defaultAIRFLOW_CONN_PYDANTICAI_CORTEX creates pydanticai_cortexmodel_id parameter on the decorator/operator (highest)model in connection's extra JSON (fallback)Besides pydanticai, the provider registers vendor-specific connection types: pydanticai-azure (Azure OpenAI: host = endpoint, extra api_version), pydanticai-bedrock (AWS credentials/region in extra), and pydanticai-vertex (GCP project/location in extra). The LlamaIndex and LangChain hooks read API key/host/extra from whatever connection ID they are given, so a single pydanticai_default connection can serve LLM calls and embeddings: one API key entry for the whole project.
# BEFORE (airflow-ai-sdk)
import airflow_ai_sdk as ai_sdk
class MyOutput(ai_sdk.BaseModel):
field: str
@task.llm(
model="gpt-5", # or model=OpenAIModel(...)
system_prompt="You are helpful.",
output_type=MyOutput,
)
def my_task(text: str) -> str:
return text
# AFTER (apache-airflow-providers-common-ai)
from pydantic import BaseModel
class MyOutput(BaseModel):
field: str
@task.llm(
llm_conn_id="pydanticai_default", # Airflow connection ID
system_prompt="You are helpful.",
output_type=MyOutput,
)
def my_task(text: str) -> str:
return textParameter mapping:
| airflow-ai-sdk | common-ai provider | Notes |
|---|---|---|
model="gpt-5" | llm_conn_id="pydanticai_default" | Model specified in connection |
model=OpenAIModel(...) | llm_conn_id="pydanticai_default" | Model + endpoint in connection |
system_prompt="..." | system_prompt="..." | Unchanged |
output_type=MyModel | output_type=MyModel | Unchanged |
result_type=MyModel | output_type=MyModel | result_type was already deprecated |
| (not available) | model_id="openai:gpt-5" | Override connection's model |
| (not available) | require_approval=True | Built-in HITL review |
| (not available) | agent_params={...} | Extra kwargs for pydantic-ai Agent |
| (not available) | serialize_output=True | Force dict shape for BaseModel output |
Multimodal prompts (0.4.0+): the translation function may return a Sequence[UserContent] instead of a string, e.g. for vision:
@task.llm(llm_conn_id="pydanticai_default", system_prompt="...", output_type=ReviewAnalysis)
def analyze(text: str, image_path: str | None = None):
if image_path:
with open(image_path, "rb") as f:
return [text, BinaryContent(data=f.read(), media_type="image/jpeg")]
return textThis matches the old airflow-ai-sdk vision pattern, so vision code migrates unchanged. Note: common-ai 0.1.x only accepted strings — if a project disabled vision to migrate to 0.1.0, re-enable it when bumping to 0.4.0. Non-string prompts are incompatible with require_approval=True / enable_hitl_review=True (both render the prompt as text).
Structured output via XCom (0.4.0 behavior change): with output_type=<BaseModel subclass>, the model instance flows through XCom on Airflow cores whose task SDK has SUPPORTS_OPERATOR_DESERIALIZATION_WALKER (attribute access downstream); on older cores (including Astro Runtime 3.2 task SDK 1.2.x) the provider automatically dumps to a dict (subscript access). Check which shape arrives at runtime before choosing attribute vs dict access downstream, or set serialize_output=True to force the dict shape everywhere. The output_type class must be defined at module scope (nested classes cannot be deserialized from XCom).
# BEFORE
@task.llm_branch(
model="gpt-5",
system_prompt="Choose a team...",
allow_multiple_branches=False,
)
def route(text: str) -> str:
return text
# AFTER
@task.llm_branch(
llm_conn_id="pydanticai_default",
system_prompt="Choose a team...",
allow_multiple_branches=False, # same parameter, unchanged
)
def route(text: str) -> str:
return textOnly change: model= becomes llm_conn_id=.
This has the biggest API change. The Agent is no longer pre-built in user code.
# BEFORE (airflow-ai-sdk) - Agent built at module level
from pydantic_ai import Agent
my_agent = Agent(
"gpt-5",
system_prompt="You are a research assistant.",
tools=[search_tool, lookup_tool],
)
@task.agent(agent=my_agent)
def research(question: str) -> str:
return question
# AFTER (common-ai provider) - No Agent object, config via parameters
from pydantic_ai.toolsets import FunctionToolset
@task.agent(
llm_conn_id="pydanticai_default",
system_prompt="You are a research assistant.",
toolsets=[FunctionToolset(tools=[search_tool, lookup_tool])],
)
def research(question: str) -> str:
return questionParameter mapping:
| airflow-ai-sdk | common-ai provider | Notes |
|---|---|---|
agent=Agent(model, ...) | llm_conn_id="..." | Model from connection |
Agent's system_prompt | system_prompt="..." | Now a decorator param |
Agent's tools=[...] | toolsets=[FunctionToolset(tools=[...])] | Preferred: gets automatic tool-call logging |
Agent's tools=[...] | agent_params={"tools": [...]} | Also works, but no tool-call logging |
Agent's output_type | output_type=MyModel | Now a decorator param |
| (not available) | durable=True | Step-level caching (needs [common.ai] durable_cache_path) |
| (not available) | enable_hitl_review=True | Iterative human review loop (see below) |
Key insight: Everything that was configured on the Agent() constructor now goes into either a top-level decorator parameter or agent_params. The agent_params dict is passed directly to pydantic-ai's Agent constructor. Prefer toolsets over agent_params["tools"]: the operator wraps each toolset in a LoggingToolset, so every tool call appears in the task log with timing.
enable_hitl_review behavior: the task generates a first draft, then blocks until a human acts. The reviewer uses the HITL Review tab/extra link on the task instance (chat UI from the provider's auto-registered hitl_review plugin) to request changes (agent regenerates with the feedback in its message history) or approve. Constraints: requires a string prompt, incompatible with durable=True, and the final (possibly regenerated) output is what flows to XCom. Warn users that the Dag run waits indefinitely at this task unless hitl_timeout is set. For headless testing, the plugin exposes REST endpoints under /hitl-review: GET /sessions/find, POST /sessions/feedback, POST /sessions/approve, POST /sessions/reject (query params dag_id, task_id, run_id, map_index).
The new provider does NOT include an embed decorator. Pick the replacement based on what the project needs:
Option A (recommended): LlamaIndexEmbeddingOperator (0.4.0, [llamaindex] extra). Connection-based, one task embeds the whole document list, and with persist_dir the resulting vector index is persisted for retrieval (pairs with LlamaIndexRetrievalOperator):
from airflow.providers.common.ai.operators.llamaindex_embedding import LlamaIndexEmbeddingOperator
_embeddings = LlamaIndexEmbeddingOperator(
task_id="create_embeddings",
documents=[{"text": "...", "metadata": {"id": 1}}, ...], # templated, accepts XComArg
llm_conn_id="pydanticai_default", # reuses the same connection (API key only)
embed_model="text-embedding-3-small",
persist_dir=f"{AIRFLOW_HOME}/include/my_index", # optional; local path or s3://, gs://, ...
)The operator returns {"chunks": [{"text", "metadata", "vector"}], ...}. Put a stable key into each document's metadata — it round-trips through chunking, so vectors can be mapped back to source records.
Option B: LlamaIndexHook for raw vectors (no operator, no persisted index). Shortest path when vectors go straight to a database:
@task
def create_embeddings(rows):
from airflow.providers.common.ai.hooks.llamaindex import LlamaIndexHook
embed_model = LlamaIndexHook(
llm_conn_id="pydanticai_default",
embed_model="text-embedding-3-small",
).get_embedding_model()
vectors = embed_model.get_text_embedding_batch([r["text"] for r in rows])
return list(zip([r["id"] for r in rows], vectors))Option C: plain @task with sentence-transformers (keeps the old local/offline behavior, no API cost; requires keeping sentence-transformers + torch in requirements):
@task
def embed_texts(texts: list[str]) -> list[list[float]]:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
return model.encode(texts, normalize_embeddings=True).tolist()Note on dimensions: switching from all-MiniLM-L6-v2 (384) to text-embedding-3-small (1536) changes vector size — existing stored embeddings must be regenerated, and fixed-size vector columns (e.g. pgvector vector(384)) need a schema change. Embed all texts in one task/batch call rather than .expand() per text: batching is one API round-trip and avoids per-task model loading.
| Old import | New import |
|---|---|
import airflow_ai_sdk as ai_sdk | Remove entirely |
from airflow_ai_sdk import BaseModel | from pydantic import BaseModel |
from airflow_ai_sdk.models.base import BaseModel | from pydantic import BaseModel |
class Foo(ai_sdk.BaseModel): | class Foo(BaseModel): |
from pydantic_ai import Agent | Remove if Agent was only used for @task.agent |
from pydantic_ai.models.openai import OpenAIModel | Remove (model config in connection now) |
| (new) | from pydantic_ai.toolsets import FunctionToolset for @task.agent toolsets |
The @task.llm, @task.agent, @task.llm_branch decorators are auto-registered by the provider. No explicit import needed beyond from airflow.sdk import task.
pydantic_ai imports for non-decorator usage (e.g., BinaryContent for multimodal) are still valid since the new provider depends on pydantic-ai-slim (>= 1.71.0 for provider 0.4.0).
pydanticai_default:
conn_type: pydanticai
password: <api-key>
extra:
model: "openai:gpt-5"For custom endpoints:
pydanticai_cortex:
conn_type: pydanticai
password: <api-key>
host: https://my-endpoint.com/v1
extra:
model: "openai:llama3.1-8b"The new provider reads model config from the pydanticai connection, so env vars that previously fed the model in code are usually redundant. Before removing any of them, grep the project (and any sibling scripts/services) to confirm nothing else still references them:
OPENAI_API_KEY|OPENAI_BASE_URL|ANTHROPIC_API_KEY|GOOGLE_API_KEYCandidates for removal only if no other code references them:
OPENAI_API_KEY (now in the pydanticai connection's password field)OPENAI_BASE_URL (now in the connection's host field)If anything outside the migrated DAGs still uses them (other DAGs not yet migrated, helper scripts, non-Airflow services sharing the .env), leave them in place.
Keep AIRFLOW_CONN_* env vars for all connections.
After migration, grep the codebase to confirm no stale references remain:
airflow_ai_sdk|airflow-ai-sdk|ai_sdk\.BaseModel|from pydantic_ai import Agent|from pydantic_ai.modelsVerify:
airflow_ai_sdkAgent() objects created for @task.agent (unless used outside decorators)model= parameter on LLM decorators (should be llm_conn_id=)@task.embed replaced (LlamaIndex operator/hook or plain @task); stored embeddings regenerated if the model/dimensions changed[text, BinaryContent(...)] again if they were string-only-restricted under common-ai 0.1.xoutput_type=BaseModel results use the XCom shape that actually arrives (dict on older cores, instance on newer; serialize_output=True pins it)pydanticai connection configured in .env or connections.yamlrequirements.txt has apache-airflow-providers-common-ai[...] instead of airflow-ai-sdk[...]; torch/sentence-transformers removed if no longer usedenable_hitl_review=True or require_approval=True wait for human input, so the test plan must include acting on them (UI tab or /hitl-review REST)These features are available after migration but have no airflow-ai-sdk equivalent:
| Feature | Parameter / API | Since | Description |
|---|---|---|---|
| HITL approval | require_approval=True on @task.llm | 0.1.0 | Pause for human review before returning |
| HITL review loop | enable_hitl_review=True on @task.agent | 0.1.0 | Iterative review with regeneration (chat UI via hitl_review plugin) |
| Durable execution | durable=True on @task.agent | 0.1.0 | Step-level caching for resilience |
| Tool logging | enable_tool_logging=True on @task.agent | 0.1.0 | INFO-level tool call logs (default: on; requires toolsets) |
| Model override | model_id="openai:gpt-5" | 0.1.0 | Override connection's model per-task |
| File analysis | @task.llm_file_analysis | 0.1.0 | Analyze files/images via ObjectStoragePath |
| NL-to-SQL | @task.llm_sql | 0.1.0 | Generate SQL from natural language |
| Multimodal prompts | Translation function returns Sequence[UserContent] | 0.4.0 | Vision and other binary content in @task.llm / @task.agent / @task.llm_branch |
| Pydantic instance via XCom | output_type=BaseModel (with serialize_output opt-out) | 0.4.0 | Instance flows through XCom on capable cores; dict fallback otherwise |
| Embeddings | LlamaIndexEmbeddingOperator (+ persist_dir) | 0.4.0 | Connection-based embeddings + persisted vector index |
| Retrieval | LlamaIndexRetrievalOperator | 0.4.0 | Top-k similarity search over a persisted index |
© astronomer, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/migrating-ai-sdk-to-common-ai of astronomer/agents.
Open the folder on GitHubat commit 486ee63
Migrating AI SDK To Common AI next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Migrating AI SDK To Common AI this skillastronomer/agents | 451 | — | ~4.7k | Automated safety check: Notes | Apache-2.0 | |
| AI Data Engineeringancoleman/ai-design-components | 525 | — | ~3.5k | Automated safety check: Pass | MIT | |
| Senior Data Engineerbenchflow-ai/skillsbench | 1.8k | — | ~5.9k | Automated safety check: Pass | MIT | |
| Senior Data Engineeralirezarezvani/claude-skills | 28k | 3 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Senior Data Engineerdavila7/claude-code-templates | 33k | 1 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Transforming Dataancoleman/ai-design-components | 525 | — | ~3k | Automated safety check: Pass | MIT |
ancoleman/ai-design-components
Data pipelines, feature stores, and embedding generation for AI/ML systems.
benchflow-ai/skillsbench
World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.
alirezarezvani/claude-skills
Data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure.
davila7/claude-code-templates
World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure.
ancoleman/ai-design-components
Transform raw data into analytical assets using ETL/ELT patterns, SQL (dbt), Python (pandas/polars/PySpark), and orchestration (Airflow).
aws/agent-toolkit-for-aws
Selects, investigates, and compares AWS object, file, and block storage services, and answers cost, performance, configuration, security, and troubleshooting questions about storage services.
astronomer/agents
Queries the data warehouse with SQL and answers business questions about data.
astronomer/agents
Queries, manages, and troubleshoots Apache Airflow using the af CLI.
astronomer/agents
Guide for migrating Dagster projects to Apache Airflow 3 on Astro.
astronomer/agents
Workflow and best practices for writing Apache Airflow DAGs.
astronomer/agents
Deploys Airflow DAGs and projects. An agent skill from astronomer/agents.
astronomer/agents
Builds human-in-the-loop (HITL) Airflow workflows - approval gates, form input, and human-driven branching.
Works with
Categories
Migrates Airflow projects from airflow-ai-sdk to apache-airflow-providers-common-ai 0.4.0+. Migrating AI SDK To Common AI is an agent skill from astronomer/agents.0+.
Migrating AI SDK To Common AI fits situations like: replacing airflow-ai-sdk with the official Airflow AI provider - migrating LLM decorators (@task.llm; @task.llmbranch; switching from model strings/objects to connection-based LLM configuration; updating imports from airflowaisdk to the new provider.
Run `npx skills add astronomer/agents --skill migrating-ai-sdk-to-common-ai -a claude-code`. Or copy the skill folder (skills/migrating-ai-sdk-to-common-ai in astronomer/agents) into .claude/skills/migrating-ai-sdk-to-common-ai in your project. Claude Code loads it when a task matches its description.
Run `npx skills add astronomer/agents --skill migrating-ai-sdk-to-common-ai -a codex`. Or copy the skill folder (skills/migrating-ai-sdk-to-common-ai in astronomer/agents) into .agents/skills/migrating-ai-sdk-to-common-ai in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add astronomer/agents --skill migrating-ai-sdk-to-common-ai -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/migrating-ai-sdk-to-common-ai, .gemini/skills/migrating-ai-sdk-to-common-ai, .github/skills/migrating-ai-sdk-to-common-ai and .opencode/skills/migrating-ai-sdk-to-common-ai in your project.
Going by SKILL.md and its folder, Migrating AI SDK To Common AI needs credentials named OPENAI_API_KEY, ANTHROPIC_API_KEY and GOOGLE_API_KEY. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Migrating AI SDK To Common AI is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Migrating AI SDK To Common AI: AI Data Engineering (ancoleman/ai-design-components, 525 stars), Senior Data Engineer (benchflow-ai/skillsbench, 1.8k stars), Senior Data Engineer (alirezarezvani/claude-skills, 28k stars) and Senior Data Engineer (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
astronomer (a GitHub organization) maintains it in astronomer/agents, which has 451 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on October 7, 2026.
Source: astronomer/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.