A skill your agent uses when querying auto-collected CI data from test runs in BigQuery (cidataautodl dataset) including risk analysis, disruption, CPU metrics, audit logs, operator state, and retry…

Apache-2.0Auto-check passedDatabases

Install Autodl

skills CLI
$ npx skills add openshift-eng/ai-helpers --skill autodl -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install openshift-eng/ai-helpers autodl --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/openshift-eng/ai-helpers.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/bigquery-ci-data/skills/autodl .claude/skills/autodl && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
autodl
GitHub stars
120
Token cost
~2.9k tokens
SKILL.md length
1,097 words
Files
1
Skills in repo
118
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when querying auto-collected CI data from test runs in BigQuery (cidataautodl dataset) including risk analysis, disruption, CPU metrics, audit logs, operator state, and retry…

  • Works in 4 steps: Monitor tests in openshift/origin… → Data is written as *autodl.json files in… → The ci-data-loader picks up these files… → …
  • Querying auto-collected CI data from test runs in BigQuery (cidataautodl dataset) including risk analysis
  • SKILL.md covers When to Use This Skill, Common Columns, Tables and Query Examples, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Autodl is an agent skill from openshift-eng/ai-helpers. Use when querying auto-collected CI data from test runs in BigQuery (cidataautodl dataset) including risk analysis, disruption, CPU metrics, audit logs, operator state, and retry statistics

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Databases, covering Data warehousing and Statistics. It works with Google BigQuery. The repository describes itself as: Developer productivity tools for Claude Code & other AI assistants. The licence is Apache-2.0.

When your agent uses it

  • Querying auto-collected CI data from test runs in BigQuery (cidataautodl dataset) including risk analysis
  • Retry statistics

Example prompts

  • “/autodl”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Monitor tests in openshift/origin collect data during test execution
  2. Data is written as *autodl.json files in job artifacts
  3. The ci-data-loader picks up these files and uploads to BigQuery
  4. New columns in schemas are auto-added to existing tables; removed columns are never deleted (cross-release integrity)

What it can do on your machine

Read from SKILL.md and the folder at commit a627176. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are sql).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Autodl loads about 2.9k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 1,097 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from openshift-eng/ai-helpers at commit a627176, republished under its Apache-2.0 licence (© openshift-eng). 1,097 words, ~2,935 tokens.

Download SKILL.mdSave it as .claude/skills/autodl/SKILL.md (or your agent's skills folder).
name
autodl
description
Use when querying auto-collected CI data from test runs in BigQuery (ci_data_autodl dataset) including risk analysis, disruption, CPU metrics, audit logs, operator state, and retry statistics

Autodl Tables

Query data automatically collected during OpenShift CI test runs, stored in openshift-ci-data-analysis.ci_data_autodl. This data is generated by monitor tests and analysis code in openshift/origin and uploaded by the ci-data-loader after each job run.

Follow the foundations skill for cost safety, caching, and execution workflow.

When to Use This Skill

  • Investigating risk analysis verdicts for job runs
  • Analyzing test retry behavior and flake patterns
  • Examining kube-apiserver audit log patterns (latency, request counts, watch storms)
  • Tracking operator state transitions during test runs
  • Correlating CPU usage with test failures
  • Understanding DNS disruption during tests
  • Checking cluster instance types and configuration

Common Columns

All autodl tables include three columns added by the ci-data-loader:

ColumnTypeNotes
JobRunNameSTRINGProw job run identifier — join key to other datasets
PartitionTimeTIMESTAMPPartition column — always filter on this
SourceSTRINGData source identifier

The JobRunName can be used to correlate autodl data with job runs in openshift-gce-devel.ci_analysis_us.jobs (match against prowjob_build_id or extract from prowjob_url).

Tables

Risk Analysis
risk_analysis_overall_results

Overall risk analysis verdict for a job run — the aggregate risk level across all tests.

ColumnTypeNotes
RiskLevelINTEGERNumeric risk level
RiskNameSTRINGHuman-readable risk name
JobRunTestCountINTEGERTotal tests in the run
JobRunTestFailuresINTEGERTotal test failures
NeverStableJobSTRINGWhether this job has ever been stable
HistoricalRunTestCountINTEGERHistorical test count for comparison

Use case: Find job runs with high risk levels, correlate risk verdicts with actual job outcomes.

risk_analysis_test_results

Per-test risk analysis — the risk level assigned to each individual test based on historical pass rates.

ColumnTypeNotes
TestNameSTRINGFull test name
TestIDINTEGERStable test identifier
RiskLevelINTEGERNumeric risk level
RiskNameSTRINGHuman-readable risk name
CurrentRunsINTEGERRecent run count for this test
CurrentPassesINTEGERRecent pass count
CurrentPassPercentageFLOATRecent pass rate

Use case: Identify which tests contributed most to a run's risk assessment, find tests with declining pass rates.

risk_analysis_api_requests

Metadata about HTTP requests to the Sippy risk analysis API during test runs.

ColumnTypeNotes
RequestCountINTEGERNumber of API requests made
StartTimeTIMESTAMPWhen the request started
DurationSecondsFLOATRequest duration
ErrorSTRINGError message if request failed
BytesReadINTEGERResponse size

Use case: Debug risk analysis API performance issues or failures.

Test Execution
retry_statistics

Per-test retry statistics when tests are retried during a run.

ColumnTypeNotes
TestNameSTRINGFull test name
RetryStrategySTRINGStrategy used for retries
TotalAttemptsINTEGERTotal attempts made
SuccessfulAttemptsINTEGERPassing attempts
FailedAttemptsINTEGERFailing attempts
FinalOutcomeSTRINGFinal test result
TotalDurationMillisecondsINTEGERTotal time across all attempts
MaxRetriesAllowedINTEGERRetry limit
FirstAttemptDurationMillisecondsINTEGERDuration of first attempt
AverageAttemptDurationMillisecondsINTEGERAverage attempt duration
JobNameSTRINGProw job name
JobTypeSTRINGperiodic, presubmit, postsubmit
PullNumberSTRINGPR number (presubmits)
RepoNameSTRINGGitHub repo
RepoOwnerSTRINGGitHub org
PullShaSTRINGCommit SHA
ReleaseImageLatestSTRINGTarget release image
ReleaseImageInitialSTRINGInitial release image (upgrades)

Use case: Analyze retry effectiveness, find tests that always fail on first attempt but pass on retry, measure retry overhead.

run_suite_options

Configuration used to run the test suite.

ColumnTypeNotes
ClusterStabilitySTRINGCluster stability mode
RandomSeedINTEGERRandom seed for test ordering
WorkerNodesINTEGERNumber of worker nodes
TotalNodesINTEGERTotal cluster nodes
ParallelismINTEGERTest parallelism level

Use case: Correlate test failures with cluster size or parallelism settings.

high_cpu_e2e_tests

Tests that overlapped with high CPU usage intervals.

ColumnTypeNotes
TestNameSTRINGFull test name
SuccessINTEGER1 = pass, 0 = fail

Use case: Find tests that fail due to high CPU or cause high CPU on the cluster.

duration-metrics

Named duration metrics (install time, upgrade time, etc.).

ColumnTypeNotes
nameSTRINGMetric name (e.g. "install", "upgrade")
durationINTEGERDuration in milliseconds

Use case: Track install/upgrade duration trends across releases and platforms.

Kube-APIServer Audit Analysis
audit_latency_counts

Histogram-bucketed latency counts for kube-apiserver requests, from audit logs.

ColumnTypeNotes
LatencyTypeSTRINGType of latency measurement
ResourceSTRINGAPI resource
VerbSTRINGHTTP verb
BucketFLOATLatency bucket threshold
CountINTEGERRequests exceeding this bucket

Use case: Identify API resources with high latency, find slow verbs, detect apiserver performance regressions.

Show full SKILL.md (426 more words)Show less
audit_resource_requests_per_user

API request counts per user/service-account, from audit logs.

ColumnTypeNotes
UserSTRINGUser or service account (cleaned)
ResourceSTRINGAPI resource
VerbSTRINGHTTP verb
HttpStatusINTEGERResponse status code
RequestCountINTEGERNumber of requests

Use case: Find noisy controllers, identify unexpected API callers, detect request storms.

operator_watch_requests

Watch request counts per operator against the kube-apiserver, from audit logs.

ColumnTypeNotes
ControlPlaneTopologySTRINGe.g. "HighlyAvailable", "SingleReplica"
PlatformTypeSTRINGCloud platform
OperatorSTRINGOperator name
WatchRequestCountINTEGERNumber of watch requests

Use case: Detect watch storms, find operators with excessive watch counts, compare across topologies.

Cluster Health
operator_state_metrics

ClusterOperator state transitions (Available, Progressing, Degraded) during the test run.

ColumnTypeNotes
OperatorSTRINGClusterOperator name
StateSTRING"Available", "Progressing", "Degraded"
CountINTEGERNumber of transitions
TotalSecondsFLOATTotal time in this state
MaxIndividualDurationSecondsFLOATLongest single period in this state

Use case: Find operators that flap between states, track degraded duration across releases.

dns_disruption_stats

DNS disruption summary during the test run.

ColumnTypeNotes
IntervalCountINTEGERNumber of disruption intervals
TotalDurationSecondsINTEGERTotal disruption time

Use case: Track DNS disruption trends, correlate with network configuration variants.

node_cpu_usage_timeline

Time-series per-node CPU usage sampled during the test run.

ColumnTypeNotes
TimestampTIMESTAMPSample time
NodeNameSTRINGNode name
NodeRoleSTRINGmaster, worker, etc.
CPUUsageFLOATCPU usage percentage

Use case: Correlate CPU spikes with test failures, identify resource-starved nodes.

interval_duration_sum

Total duration of monitor intervals by source type.

ColumnTypeNotes
IntervalSourceSTRING"MetricsEndpointDown", "CPUMonitor"
TotalDurationSecondsINTEGERTotal duration

Use case: Track metrics endpoint availability, CPU monitoring coverage.

cluster_instance_types

Cloud instance types used by nodes in the cluster (AWS, Azure, GCP).

ColumnTypeNotes
PlatformSTRINGCloud provider
RegionSTRINGCloud region
ZoneSTRINGAvailability zone
RoleSTRINGNode role
InstanceTypeSTRINGInstance type (e.g. m5.xlarge)
SuiteSTRINGTest suite

Use case: Correlate failures with instance types, track what hardware CI uses.

Query Examples

Find high-risk job runs in the last week
sql
SELECT
  JobRunName,
  RiskLevel,
  RiskName,
  JobRunTestCount,
  JobRunTestFailures
FROM `openshift-ci-data-analysis.ci_data_autodl.risk_analysis_overall_results`
WHERE PartitionTime >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 7 DAY)
  AND RiskLevel >= 3
ORDER BY RiskLevel DESC, JobRunTestFailures DESC
Find tests with worst retry rates
sql
SELECT
  TestName,
  COUNT(*) AS total_runs,
  COUNTIF(TotalAttempts > 1) AS runs_with_retries,
  ROUND(COUNTIF(TotalAttempts > 1) * 100.0 / COUNT(*), 1) AS retry_pct,
  ROUND(AVG(IF(TotalAttempts > 1, TotalAttempts, NULL)), 1) AS avg_attempts_when_retried,
  COUNTIF(FinalOutcome = 'passed') / COUNT(*) AS final_pass_rate
FROM `openshift-ci-data-analysis.ci_data_autodl.retry_statistics`
WHERE PartitionTime >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 7 DAY)
GROUP BY TestName
HAVING total_runs >= 5
ORDER BY retry_pct DESC
Find operators with most degraded time
sql
SELECT
  Operator,
  COUNT(*) AS run_count,
  AVG(TotalSeconds) AS avg_degraded_seconds,
  MAX(MaxIndividualDurationSeconds) AS worst_degraded_seconds
FROM `openshift-ci-data-analysis.ci_data_autodl.operator_state_metrics`
WHERE PartitionTime >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 7 DAY)
  AND State = 'Degraded'
  AND Count > 0
GROUP BY Operator
ORDER BY avg_degraded_seconds DESC
Find noisiest API callers
sql
SELECT
  User,
  Resource,
  Verb,
  SUM(RequestCount) AS total_requests
FROM `openshift-ci-data-analysis.ci_data_autodl.audit_resource_requests_per_user`
WHERE PartitionTime >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 7 DAY)
GROUP BY User, Resource, Verb
ORDER BY total_requests DESC
LIMIT 20

Data Pipeline

The autodl data pipeline works as follows:

  1. Monitor tests in openshift/origin collect data during test execution
  2. Data is written as *autodl.json files in job artifacts
  3. The ci-data-loader picks up these files and uploads to BigQuery
  4. New columns in schemas are auto-added to existing tables; removed columns are never deleted (cross-release integrity)

Source code for all table definitions is in openshift/origin:

  • Data loader framework: pkg/dataloader/types.go
  • Individual monitor tests: pkg/monitortests/ subdirectories
  • Risk analysis: pkg/riskanalysis/cmd.go
  • Retry statistics: pkg/test/ginkgo/retries.go
  • Suite options: pkg/test/ginkgo/cmd_runsuite.go
  • Duration metrics: pkg/e2eanalysis/e2e_analysis.go

© openshift-eng, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/bigquery-ci-data/skills/autodl of openshift-eng/ai-helpers.

Open the folder on GitHubat commit a627176

Compare with similar skills

Autodl next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Autodl compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Autodl this skillopenshift-eng/ai-helpers120—~2.9kAutomated safety check: PassApache-2.0
Databrain Intelligenceinfometa/workbuddyskills344—~8kAutomated safety check: PassNone
Data Warehouse Experimentationrampstackco/claude-skills940—~7.3kAutomated safety check: PassMIT
Altimate Data Warehouse DelegateAltimateAI/data-engineering-skills128—~1.4kAutomated safety check: PassMIT
dbt Snowflake to BigQuery Translatorgoogle/skills21k—~2.7kAutomated safety check: PassApache-2.0
Bigquery Bigframesgoogle/skills21k1 repos~1.3kAutomated safety check: PassApache-2.0

Similar skills

  • Databrain Intelligence

    infometa/workbuddyskills

    DataBrain intelligence data query assistant. An agent skill from infometa/workbuddyskills.

    344 GitHub stars~8k tokensUpdated yesterday
    DatabasesAuto-check passed
  • Data Warehouse Experimentation

    rampstackco/claude-skills

    Running experiments out of the data warehouse instead of via dedicated experiment platforms.

    940 GitHub stars~7.3k tokensUpdated yesterday
    DatabasesAuto-check passed
  • Altimate Data Warehouse Delegate

    AltimateAI/data-engineering-skills

    Delegates dbt and warehouse tasks such as lineage, migrations and cost attribution to the altimate-code CLI agent and relays its answer back.

    128 GitHub stars~1.4k tokensUpdated 7 days ago
    DatabasesAuto-check passed
  • Translates Snowflake dbt SQL models into standardized BigQuery SQL, keeping Jinja constructs and tracking progress in a migration tasks file.

    21k GitHub stars~2.7k tokensUpdated yesterday
    DatabasesAuto-check passed
  • Bigquery Bigframes

    google/skills

    Official

    Generates Python code using BigQuery DataFrames (BigFrames).

    21k GitHub starsUsed in 1 repo~1.3k tokens
    DatabasesAuto-check passed
  • Dbt Model Index

    warpdotdev/oz-skills

    Provide a lookup index of dbt models (BigQuery tables) to guide query writing against a data warehouse.

    825 GitHub stars~915 tokensUpdated 1 mo ago
    DatabasesAuto-check passed

More from openshift-eng/ai-helpers

All 118 skills in this repo
  • Investigate CI Reliability

    openshift-eng/ai-helpers

    Find and independently validate actionable reliability defects across OpenShift release jobs and presubmits, then export portable issue handoffs.

    120 GitHub stars~1.9k tokensUpdated 2 days ago
    Auto-check passed
  • Address Review PR

    openshift-eng/ai-helpers

    Fetch and address all PR review comments — categorize by priority, make code changes, post replies, and push.

    120 GitHub stars~2.9k tokensUpdated 2 days ago
    Auto-check passed
  • Categorize Activity Types

    openshift-eng/ai-helpers

    Categorize Jira issues into Red Hat Sankey Activity Type categories using MCP Jira tools.

    120 GitHub stars~2.4k tokensUpdated 2 days ago
    Auto-check passed
  • Has Review Work

    openshift-eng/ai-helpers

    Decide whether a GitHub PR has unanswered authorized review comments or new required CI failures worth a follow-up agent.

    120 GitHub stars~1.9k tokensUpdated 2 days ago
    Auto-check passed
  • Must Gather Analyzer

    openshift-eng/ai-helpers

    Analyze OpenShift must-gather diagnostic data including cluster operators, pods, nodes, and network components.

    120 GitHub stars~2.3k tokensUpdated 2 days ago
    Auto-check passed
  • Payload Autodl JSON

    openshift-eng/ai-helpers

    Schema for the autodl JSON data file produced by payload-analysis for database ingestion — you must use this skill whenever generating the autodl JSON file

    120 GitHub stars~2.6k tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about Autodl

What does Autodl do?

A skill your agent uses when querying auto-collected CI data from test runs in BigQuery (cidataautodl dataset) including risk analysis, disruption, CPU metrics, audit logs, operator state, and retry…. Autodl is an agent skill from openshift-eng/ai-helpers.

When should I use Autodl?

Autodl fits situations like: querying auto-collected CI data from test runs in BigQuery (cidataautodl dataset) including risk analysis; retry statistics.

How do I install Autodl in Claude Code?

Run `npx skills add openshift-eng/ai-helpers --skill autodl -a claude-code`. Or copy the skill folder (plugins/bigquery-ci-data/skills/autodl in openshift-eng/ai-helpers) into .claude/skills/autodl in your project. Claude Code loads it when a task matches its description.

How do I install Autodl in Codex?

Run `npx skills add openshift-eng/ai-helpers --skill autodl -a codex`. Or copy the skill folder (plugins/bigquery-ci-data/skills/autodl in openshift-eng/ai-helpers) into .agents/skills/autodl in your project. Codex loads it when a task matches its description.

Can I use Autodl in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add openshift-eng/ai-helpers --skill autodl -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/autodl, .gemini/skills/autodl, .github/skills/autodl and .opencode/skills/autodl in your project.

What does Autodl need to run?

SKILL.md names no scripts, command-line tools or credentials: Autodl is instructions for the agent only.

Does Autodl access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Autodl safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Autodl use?

Autodl is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Autodl use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Autodl?

Skills that share tags, products or a category with Autodl: Databrain Intelligence (infometa/workbuddyskills, 344 stars), Data Warehouse Experimentation (rampstackco/claude-skills, 940 stars), Altimate Data Warehouse Delegate (AltimateAI/data-engineering-skills, 128 stars) and dbt Snowflake to BigQuery Translator (google/skills, 21k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Autodl?

openshift-eng (a GitHub organization) maintains it in openshift-eng/ai-helpers, which has 120 GitHub stars. The repository holds 118 skills in this directory. The repository was last updated on October 6, 2026.

Source: openshift-eng/ai-helpers on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.