Agent skill

Synthetic Data Generation

by Red-Hat-AI-Innovation-Team in Red-Hat-AI-Innovation-Team/sdg_hub

Generate synthetic data using sdghub with composable blocks and YAML flows.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Synthetic Data Generation

skills CLI
$ npx skills add Red-Hat-AI-Innovation-Team/sdg_hub --skill synthetic-data-generation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Red-Hat-AI-Innovation-Team/sdg_hub synthetic-data-generation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Red-Hat-AI-Innovation-Team/sdg_hub.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/synthetic-data-generation .claude/skills/synthetic-data-generation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
synthetic-data-generation
GitHub stars
164
Token cost
~3.1k tokens
SKILL.md length
491 words
Files
6 (incl. references)
Skills in repo
4
Repo updated
First seen
Licence
Apache-2.0

At a glance

Generate synthetic data using sdghub with composable blocks and YAML flows.

  • Works in 9 steps: Discover flows → Load and inspect → Configure model → …
  • The user wants to create training datasets
  • SKILL.md covers Choose Your Approach, Approach A: Pre-Built Flows, Approach B: Custom Python… and Approach C: Authoring Custom…, plus 7 more sections
  • Needs OPENAI_API_KEY

What it does

Synthetic Data Generation is an agent skill from Red-Hat-AI-Innovation-Team/sdg_hub. Generate synthetic data using sdghub with composable blocks and YAML flows. Use when the user wants to create training datasets, generate QA pairs, run data generation pipelines, build custom flows, produce synthetic data from documents, use agent frameworks for data generation, or distill MCP tool-use traces. Supports pre-built flows, custom Python scripts, and YAML flow authoring with 20+ blocks, agent connectors (Langflow, LangGraph), MCP tool-use, and 100+ LLM providers via LiteLLM.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/block_reference.md`, `references/flow_patterns.md` and `references/model_configs.md`).

It sits in AI & LLM Engineering, covering Test data and fixtures, Building AI agents and MCP servers. It works with LangGraph and Python. The repository describes itself as: Synthetic Data Generation Toolkit for LLMs. The licence is Apache-2.0.

When your agent uses it

  • The user wants to create training datasets
  • Generate QA pairs
  • Run data generation pipelines
  • Build custom flows

Example prompts

  • “/synthetic-data-generation”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Discover flows
  2. Load and inspect
  3. Configure model
  4. Prepare data and dry run
  5. Generate and save
  6. Define the data contract
  7. Minimal YAML
  8. Create prompt template
  9. Test incrementally

What it can do on your machine

Read from SKILL.md and the folder at commit 31efcbe. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and yaml).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Synthetic Data Generation loads about 3.1k tokens when it runs, and up to ~9.7k if it reads all its reference files. Until then it costs about 130 tokens; SKILL.md has 491 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~130
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Red-Hat-AI-Innovation-Team/sdg_hub at commit 31efcbe, republished under its Apache-2.0 licence (© Red-Hat-AI-Innovation-Team). 491 words, ~3,051 tokens.

Download SKILL.mdSave it as .claude/skills/synthetic-data-generation/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
synthetic-data-generation
description
Generate synthetic data using sdg_hub with composable blocks and YAML flows. Use when the user wants to create training datasets, generate QA pairs, run data generation pipelines, build custom flows, produce synthetic data from documents, use agent frameworks for data generation, or distill MCP tool-use traces. Supports pre-built flows, custom Python scripts, and YAML flow authoring with 20+ blocks, agent connectors (Langflow, LangGraph), MCP tool-use, and 100+ LLM providers via LiteLLM.

Synthetic Data Generation with SDG Hub

Generate synthetic data using composable blocks and flows. Blocks are processing units that transform datasets; flows chain blocks into pipelines defined in YAML.

Core concept: dataset -> Block_1 -> Block_2 -> Block_3 -> enriched_dataset

Choose Your Approach

ApproachWhen to Use
Pre-built flowStandard pipeline exists for your task (QA generation, text analysis, red-teaming, RAG eval, MCP distillation)
Custom PythonQuick experiments, ad-hoc generation, custom logic
Custom YAML flowReusable pipeline, team sharing, complex multi-block workflows
Agent-basedNeed external agent frameworks (Langflow, LangGraph) or MCP tool-use in your pipeline

Approach A: Pre-Built Flows

Step 1: Discover flows
python
# play.py
from sdg_hub import FlowRegistry

# List all flows
for f in FlowRegistry.list_flows():
    print(f"- {f['name']} (tags: {f.get('tags', [])})")

# Search by tag
FlowRegistry.search_flows(tag="qa-generation")

Consult references/pre_built_flows.md for the full catalog with descriptions and required inputs.

Step 2: Load and inspect
python
from sdg_hub import Flow, FlowRegistry

path = FlowRegistry.get_flow_path("Flow Name or ID")
flow = Flow.from_yaml(path)
flow.print_info()

# Check what dataset columns are needed
reqs = flow.get_dataset_requirements()
if reqs:
    print(f"Required columns: {reqs.required_columns}")
Step 3: Configure model
python
import os

flow.set_model_config(
    model="openai/gpt-4o-mini",
    api_key=os.environ.get("OPENAI_API_KEY")
)

# For local models (vLLM, Ollama)
flow.set_model_config(
    model="meta-llama/Llama-3.3-70B-Instruct",
    api_base="http://localhost:8000/v1",
    api_key="EMPTY"
)

See references/model_configs.md for all supported providers (OpenAI, Anthropic, Azure, vLLM, Ollama, Together, Groq, Bedrock, etc.).

Step 4: Prepare data and dry run
python
import pandas as pd

df = pd.DataFrame({"document": ["Your text here..."]})

# Validate dataset against flow requirements
errors = flow.validate_dataset(df)
if errors:
    print(f"Fix these: {errors}")

# Dry run with 2 samples -- do this before every full run
dry = flow.dry_run(df, sample_size=2)
print(f"Success: {dry['execution_successful']}")
for block in dry['blocks_executed']:
    print(f"  {block['block_name']}: {block['execution_time_seconds']:.2f}s")
Step 5: Generate and save
python
# Full run with checkpointing for large datasets
result = flow.generate(
    df,
    checkpoint_dir="./checkpoints",
    save_freq=100,
    max_concurrency=5
)

result.to_parquet("output.parquet")

Approach B: Custom Python Scripts

Use blocks directly for ad-hoc experiments.

Basic: Single block
python
# play.py
from sdg_hub.core.blocks import LLMChatBlock
import pandas as pd

block = LLMChatBlock(
    block_name="gen",
    input_cols="messages",
    output_cols="response",
    model="openai/gpt-4o-mini",
    api_key="sk-...",
    temperature=0.7
)

df = pd.DataFrame({
    "messages": [[
        {"role": "system", "content": "You generate QA pairs."},
        {"role": "user", "content": "Generate a fun fact about Python."}
    ]]
})

result = block(df)
print(result["response"].iloc[0])
Chain: Multiple blocks
python
from sdg_hub.core.blocks import LLMChatBlock, TagParserBlock
import pandas as pd

# Step 1: Generate
llm = LLMChatBlock(
    block_name="gen",
    input_cols="messages",
    output_cols="response",
    model="openai/gpt-4o-mini",
    api_key="sk-..."
)

# Step 2: Parse with tags
parser = TagParserBlock(
    block_name="parse",
    input_cols="response",
    output_cols=["question", "answer"],
    start_tags=["<question>", "<answer>"],
    end_tags=["</question>", "</answer>"]
)

df = pd.DataFrame({
    "messages": [[
        {"role": "user", "content": "Generate a QA pair. Use <question>...</question> and <answer>...</answer> tags."}
    ]]
})

result = parser(llm(df))
print(result[["question", "answer"]])
Batch processing for large datasets
python
from tqdm import tqdm

def process_in_batches(df, block, batch_size=50):
    results = []
    for i in tqdm(range(0, len(df), batch_size)):
        batch = df.iloc[i:i+batch_size].copy()
        results.append(block(batch))
    return pd.concat(results, ignore_index=True)

See references/block_reference.md for all 20+ available blocks and their configurations.

Approach C: Authoring Custom Flow YAMLs

Build incrementally -- start with one block, test, add the next.

Step 1: Define the data contract
python
# play.py - Clarify inputs and outputs first
import pandas as pd

input_df = pd.DataFrame({
    "document": ["Climate change is accelerating..."],
    "domain": ["environment"]
})
print("Input columns:", list(input_df.columns))

expected_outputs = ["document", "domain", "question", "response"]
print("Expected output:", expected_outputs)
Step 2: Minimal YAML
yaml
# flow.yaml
metadata:
  name: "My QA Flow"
  version: "0.1.0"
  author: "Your Name"
  description: "Generate QA pairs from documents"
  dataset_requirements:
    required_columns: ["document"]

blocks:
  - block_type: "PromptBuilderBlock"
    block_config:
      block_name: "build_prompt"
      input_cols: ["document"]
      output_cols: "messages"
      prompt_config_path: "prompts/qa.yaml"

  - block_type: "LLMChatBlock"
    block_config:
      block_name: "generate"
      input_cols: "messages"
      output_cols: "raw_response"
      temperature: 0.7
      async_mode: true

  - block_type: "TagParserBlock"
    block_config:
      block_name: "parse"
      input_cols: "raw_response"
      output_cols: ["question", "response"]
      start_tags: ["<question>", "<answer>"]
      end_tags: ["</question>", "</answer>"]
Step 3: Create prompt template
yaml
# prompts/qa.yaml (relative to flow.yaml)
- role: system
  content: |
    You generate question-answer pairs from documents.

- role: user
  content: |
    Generate one question and answer from this document.
    Use <question>...</question> and <answer>...</answer> tags.

    {document}
Step 4: Test incrementally
python
# play.py
from sdg_hub import Flow
import pandas as pd

flow = Flow.from_yaml("flow.yaml")
flow.set_model_config(model="openai/gpt-4o-mini", api_key="sk-...")

df = pd.DataFrame({"document": ["Python was created by Guido van Rossum in 1991."]})

# Dry run first
dry = flow.dry_run(df, sample_size=1)
print(f"Success: {dry['execution_successful']}")

# Full run
if dry['execution_successful']:
    result = flow.generate(df)
    print(result[["document", "question", "response"]])

See references/yaml_schema.md for the complete YAML structure and references/flow_patterns.md for common patterns (quality filtering, parallel paths, multi-step extraction).

Approach D: Agent and MCP Pipelines

Agent frameworks (Langflow, LangGraph)

Use AgentBlock to call external agent frameworks as pipeline steps:

python
from sdg_hub.core.blocks.agent import AgentBlock

block = AgentBlock(
    block_name="my_agent",
    agent_framework="langflow",       # or "langgraph"
    agent_url="http://localhost:7860/api/v1/run/my-flow",
    agent_api_key="your-key",
    input_cols=["question"],
    output_cols=["agent_response"],
    extract_response=True
)

result = block.generate(dataset)

In YAML flows, configure agent blocks with set_agent_config():

python
flow = Flow.from_yaml("flow.yaml")
if flow.is_agent_config_required():
    flow.set_agent_config(
        agent_framework="langgraph",
        agent_url="http://localhost:8123",
        agent_api_key="your-key"
    )
MCP tool-use distillation

MCPAgentBlock connects an LLM to a remote MCP server for agentic tool-use. The LLM calls tools in a loop, producing full traces for training data:

yaml
- block_type: "MCPAgentBlock"
  block_config:
    block_name: "mcp_agent"
    input_cols: "messages"
    output_cols: "agent_trace"
    mcp_server_url: "http://localhost:3000/mcp"
    max_iterations: 10

See the pre-built MCP Server Distillation flow in references/pre_built_flows.md for a complete pipeline.

Show full SKILL.md (191 more words)Show less

Flow Methods Quick Reference

python
flow = Flow.from_yaml("flow.yaml")

# Model configuration
flow.set_model_config(model="...", api_key="...", blocks=["specific_block"])
flow.is_model_config_required()
flow.get_default_model()
flow.get_model_recommendations()

# Agent configuration
flow.set_agent_config(agent_framework="...", agent_url="...", agent_api_key="...")
flow.is_agent_config_required()

# Dataset validation
flow.validate_dataset(df)
flow.get_dataset_requirements()

# Execution
flow.dry_run(df, sample_size=2)
flow.generate(df, checkpoint_dir="./ckpt", save_freq=100, max_concurrency=5)

# Inspection
flow.print_info()
flow.to_yaml("output_flow.yaml")

Block Discovery

python
from sdg_hub.core.blocks import BlockRegistry

BlockRegistry.discover_blocks()                    # Rich table of all blocks
BlockRegistry.list_blocks(category="llm")          # By category
BlockRegistry.list_blocks(grouped=True)            # Grouped by category
BlockRegistry.categories()                         # All categories

Data I/O

python
import pandas as pd

# Load
df = pd.read_csv("input.csv")
df = pd.read_parquet("input.parquet")
df = pd.read_json("input.jsonl", lines=True)

# From HuggingFace
from datasets import load_dataset
df = load_dataset("your_dataset", split="train").to_pandas()

# Save
result.to_parquet("output.parquet")
result.to_csv("output.csv", index=False)
result.to_json("output.jsonl", orient="records", lines=True)

# Push to HuggingFace Hub
from datasets import Dataset
Dataset.from_pandas(result).push_to_hub("username/dataset")

Quality Checklist

Before using generated data:

  • Dry run succeeded with sample_size=2?
  • Output columns are correct?
  • Sample outputs look reasonable (spot-check 5-10)?
  • No excessive nulls or empty values?
  • Data saved to durable storage?

Common Issues

"Column X not found" -- Input data is missing a required column. Run flow.get_dataset_requirements() to see what the flow expects, then check your DataFrame columns.

Empty or null outputs -- The LLM response didn't match the parser pattern. Check the raw LLM output before parsing, and adjust your prompt template or parser config.

Rate limit errors -- Reduce max_concurrency in flow.generate() or add timeout and num_retries to set_model_config().

Slow generation -- Use async_mode: true on LLMChatBlock, increase max_concurrency, or use checkpointing to resume interrupted runs.

Model not responding -- Verify your model config works with a single-sample test:

python
from sdg_hub.core.blocks import LLMChatBlock
block = LLMChatBlock(block_name="test", input_cols="messages", output_cols="r", model="...", api_key="...")
block(pd.DataFrame({"messages": [[{"role": "user", "content": "hello"}]]}))

Reference Files

Detailed documentation for specific topics:

  • references/block_reference.md -- All 20+ blocks with YAML configs and usage examples
  • references/pre_built_flows.md -- Catalog of pre-built flows with inputs, outputs, and usage
  • references/model_configs.md -- LLM provider configurations (OpenAI, Anthropic, vLLM, Ollama, etc.)
  • references/yaml_schema.md -- Complete flow YAML structure and validation rules
  • references/flow_patterns.md -- Common composition patterns (LLM chain, quality filtering, parallel paths, agent integration)

© Red-Hat-AI-Innovation-Team, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in .claude/skills/synthetic-data-generation of Red-Hat-AI-Innovation-Team/sdg_hub.

  • SKILL.md
  • references/block_reference.md
  • references/flow_patterns.md
  • references/model_configs.md
  • references/pre_built_flows.md
  • references/yaml_schema.md

Open the folder on GitHubat commit 31efcbe

Compare with similar skills

Synthetic Data Generation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Synthetic Data Generation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Synthetic Data Generation this skillRed-Hat-AI-Innovation-Team/sdg_hub164—~3.1kAutomated safety check: PassApache-2.0
Cloudbase Agent PythonTencentCloudBase/CloudBase-AI-Toolkit1.1k2 repos~2.9kAutomated safety check: NotesMIT
Tool Designagentailor/fullstack-langgraph-nextjs-agent132—~3.2kAutomated safety check: PassMIT
Agent Squad Python Guide2FastLabs/agent-squad7.8k—~4.7kAutomated safety check: PassApache-2.0
Add Example AgentGetBindu/Bindu10k—~1.1kAutomated safety check: NotesCustom licence
Failproof AI SDK IntegrationFailproofAI/failproofai5.3k—~6kAutomated safety check: PassCustom licence

Similar skills

  • Cloudbase Agent Python

    TencentCloudBase/CloudBase-AI-Toolkit

    Build production-ready AI agent backends using the CloudBase Agent Python SDK — create agents with LangGraph/CrewAI/LlamaIndex, serve them via FastAPI with AG-UI protocol streaming +…

    1.1k GitHub starsUsed in 2 repos~2.9k tokens
    AI & LLM EngineeringAuto-check: notes
  • Tool Design

    agentailor/fullstack-langgraph-nextjs-agent

    Design and verify tools that AI agents can actually use — for any framework or language (MCP servers, LangChain/LangGraph, function-calling, raw JSON schema; TypeScript, Python, or otherwise).

    132 GitHub stars~3.2k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Agent Squad Python Guide

    2FastLabs/agent-squad

    Map of the agent-squad Python framework for async multi-agent orchestration: which agent, classifier, storage and tool provider to pick, and the pitfalls to avoid.

    7.8k GitHub stars~4.7k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Add Example Agent

    GetBindu/Bindu

    Add a new self-contained example agent under examples/. An agent skill from GetBindu/Bindu.

    10k GitHub stars~1.1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Failproof AI SDK Integration

    FailproofAI/failproofai

    Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

    5.3k GitHub stars~6k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Omnigent Framework Detection

    omnigent-ai/omnigent

    Scans Python agent code for framework imports and recommends the matching Omnigent executor type, or says when the framework is not natively supported yet.

    11k GitHub stars~610 tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from Red-Hat-AI-Innovation-Team/sdg_hub

  • Setup Guide

    Red-Hat-AI-Innovation-Team/sdg_hub

    A skill your agent uses when the user wants to set up synthetic data generation for the first time, or when sdghub is not yet installed/configured in the current environment.

    164 GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • Data Generation

    Red-Hat-AI-Innovation-Team/sdg_hub

    A skill your agent uses when the user wants to run synthetic data generation via scripts — detect environment, execute a flow, and present results.

    164 GitHub stars~381 tokensUpdated 2 days ago
    Auto-check passed
  • Flow Browser

    Red-Hat-AI-Innovation-Team/sdg_hub

    A skill your agent uses when the user wants to list, search, or inspect available SDG flows and data generation pipelines.

    164 GitHub stars~343 tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about Synthetic Data Generation

What does Synthetic Data Generation do?

Generate synthetic data using sdghub with composable blocks and YAML flows. Synthetic Data Generation is an agent skill from Red-Hat-AI-Innovation-Team/sdg_hub. Generate synthetic data using sdghub with composable blocks and YAML flows.

When should I use Synthetic Data Generation?

Synthetic Data Generation fits situations like: the user wants to create training datasets; generate QA pairs; run data generation pipelines; build custom flows.

How do I install Synthetic Data Generation in Claude Code?

Run `npx skills add Red-Hat-AI-Innovation-Team/sdg_hub --skill synthetic-data-generation -a claude-code`. Or copy the skill folder (.claude/skills/synthetic-data-generation in Red-Hat-AI-Innovation-Team/sdg_hub) into .claude/skills/synthetic-data-generation in your project. Claude Code loads it when a task matches its description.

How do I install Synthetic Data Generation in Codex?

Run `npx skills add Red-Hat-AI-Innovation-Team/sdg_hub --skill synthetic-data-generation -a codex`. Or copy the skill folder (.claude/skills/synthetic-data-generation in Red-Hat-AI-Innovation-Team/sdg_hub) into .agents/skills/synthetic-data-generation in your project. Codex loads it when a task matches its description.

Can I use Synthetic Data Generation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Red-Hat-AI-Innovation-Team/sdg_hub --skill synthetic-data-generation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/synthetic-data-generation, .gemini/skills/synthetic-data-generation, .github/skills/synthetic-data-generation and .opencode/skills/synthetic-data-generation in your project.

What does Synthetic Data Generation need to run?

Going by SKILL.md and its folder, Synthetic Data Generation needs credentials named OPENAI_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY.

Does Synthetic Data Generation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Synthetic Data Generation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Synthetic Data Generation use?

Synthetic Data Generation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Synthetic Data Generation use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.6k tokens, read only when the agent opens those files.

What are the alternatives to Synthetic Data Generation?

Skills that share tags, products or a category with Synthetic Data Generation: Cloudbase Agent Python (TencentCloudBase/CloudBase-AI-Toolkit, 1.1k stars), Tool Design (agentailor/fullstack-langgraph-nextjs-agent, 132 stars), Agent Squad Python Guide (2FastLabs/agent-squad, 7.8k stars) and Add Example Agent (GetBindu/Bindu, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Synthetic Data Generation?

Red-Hat-AI-Innovation-Team (a GitHub organization) maintains it in Red-Hat-AI-Innovation-Team/sdg_hub, which has 164 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 9, 2026.

Source: Red-Hat-AI-Innovation-Team/sdg_hub on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.