Agent skill

Logs Analysis

by openshift-eng in openshift-eng/ai-helpers

Analyze system and application log data from sosreport archives, extracting error patterns, kernel panics, OOM events, service failures, and application crashes from journald logs and traditional…

Apache-2.0Auto-check passedDevelopment

Install Logs Analysis

skills CLI
$ npx skills add openshift-eng/ai-helpers --skill logs-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install openshift-eng/ai-helpers logs-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/openshift-eng/ai-helpers.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/sosreport/skills/logs-analysis .claude/skills/logs-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
logs-analysis
GitHub stars
120
Token cost
~2.6k tokens
SKILL.md length
710 words
Files
1
Skills in repo
118
Repo updated
First seen
Licence
Apache-2.0

At a glance

Analyze system and application log data from sosreport archives, extracting error patterns, kernel panics, OOM events, service failures, and application crashes from journald logs and traditional…

  • Works in 6 steps: Identify Available Log Sources → Analyze Journald Logs → Analyze System Logs (var/log) → …
  • Tasks that involve Root cause analysis
  • SKILL.md covers When to Use This Skill, Prerequisites, Key Log Locations in Sosreport and Implementation Steps, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Logs Analysis is an agent skill from openshift-eng/ai-helpers. Analyze system and application log data from sosreport archives, extracting error patterns, kernel panics, OOM events, service failures, and application crashes from journald logs and traditional log files within the sosreport directory structure to identify root causes of system failures and issues

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Root cause analysis. The repository describes itself as: Developer productivity tools for Claude Code & other AI assistants. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Root cause analysis

Example prompts

  • “/logs-analysis”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Identify Available Log Sources
  2. Analyze Journald Logs
  3. Analyze System Logs (var/log)
  4. Count and Categorize Errors
  5. Analyze Application-Specific Logs
  6. Generate Log Analysis Summary

What it can do on your machine

Read from SKILL.md and the folder at commit a627176. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Logs Analysis loads about 2.6k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 710 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from openshift-eng/ai-helpers at commit a627176, republished under its Apache-2.0 licence (© openshift-eng). 710 words, ~2,636 tokens.

Download SKILL.mdSave it as .claude/skills/logs-analysis/SKILL.md (or your agent's skills folder).
name
logs-analysis
description
Analyze system and application log data from sosreport archives, extracting error patterns, kernel panics, OOM events, service failures, and application crashes from journald logs and traditional log files within the sosreport directory structure to identify root causes of system failures and issues

Logs Analysis Skill

This skill provides detailed guidance for analyzing logs from sosreport archives, including journald logs, system logs, kernel messages, and application logs.

When to Use This Skill

Use this skill when:

  • Analyzing the /sosreport:analyze command's log analysis phase
  • Investigating specific log-related errors or warnings in a sosreport
  • Performing deep-dive analysis of system failures from logs
  • Identifying patterns and root causes in system logs

Prerequisites

  • Sosreport archive must be extracted to a working directory
  • Path to the sosreport root directory must be known
  • Basic understanding of Linux log structure and journald

Key Log Locations in Sosreport

Sosreports contain logs in several locations:

  1. Journald logs: sos_commands/logs/journalctl_*

    • journalctl_--no-pager_--boot - Current boot logs
    • journalctl_--no-pager - All available logs
    • journalctl_--no-pager_--priority_err - Error priority logs
  2. Traditional system logs: var/log/

    • messages - System-level messages
    • dmesg - Kernel ring buffer
    • secure - Authentication and security logs
    • cron - Cron job logs
  3. Application logs: var/log/ (varies by application)

    • httpd/ - Apache logs
    • nginx/ - Nginx logs
    • audit/audit.log - SELinux audit logs

Implementation Steps

Step 1: Identify Available Log Sources
  1. Check for journald logs:

    bash
    ls -la sos_commands/logs/journalctl_* 2>/dev/null || echo "No journald logs found"
  2. Check for traditional system logs:

    bash
    ls -la var/log/{messages,dmesg,secure} 2>/dev/null || echo "No traditional logs found"
  3. Identify application-specific logs:

    bash
    find var/log/ -type f -name "*.log" 2>/dev/null | head -20
Step 2: Analyze Journald Logs
  1. Parse journalctl output for error patterns:

    bash
    # Look for common error indicators
    grep -iE "(error|failed|failure|critical|panic|segfault|oom)" sos_commands/logs/journalctl_--no-pager 2>/dev/null | head -100
  2. Identify OOM (Out of Memory) killer events:

    bash
    grep -i "out of memory\|oom.*kill" sos_commands/logs/journalctl_--no-pager 2>/dev/null
  3. Find kernel panics:

    bash
    grep -i "kernel panic\|bug:\|oops:" sos_commands/logs/journalctl_--no-pager 2>/dev/null
  4. Check for segmentation faults:

    bash
    grep -i "segfault\|sigsegv\|core dump" sos_commands/logs/journalctl_--no-pager 2>/dev/null
  5. Extract service failures:

    bash
    grep -i "failed to start\|failed with result" sos_commands/logs/journalctl_--no-pager 2>/dev/null
Step 3: Analyze System Logs (var/log)
  1. Check messages for errors:

    bash
    # If file exists and is readable
    if [ -f var/log/messages ]; then
      grep -iE "(error|failed|failure|critical)" var/log/messages | tail -100
    fi
  2. Check dmesg for hardware issues:

    bash
    if [ -f var/log/dmesg ]; then
      grep -iE "(error|fail|warning|i/o error|bad sector)" var/log/dmesg
    fi
  3. Analyze authentication logs:

    bash
    if [ -f var/log/secure ]; then
      grep -iE "(failed|failure|invalid|denied)" var/log/secure | tail -50
    fi
Step 4: Count and Categorize Errors
  1. Count errors by severity:

    bash
    # Critical errors
    grep -ic "critical\|panic\|fatal" sos_commands/logs/journalctl_--no-pager 2>/dev/null || echo "0"
    
    # Errors
    grep -ic "error" sos_commands/logs/journalctl_--no-pager 2>/dev/null || echo "0"
    
    # Warnings
    grep -ic "warning\|warn" sos_commands/logs/journalctl_--no-pager 2>/dev/null || echo "0"
  2. Find most frequent error messages:

    bash
    grep -iE "(error|failed)" sos_commands/logs/journalctl_--no-pager 2>/dev/null | \
      sed 's/^.*\]: //' | \
      sort | uniq -c | sort -rn | head -10
  3. Extract timestamps for error timeline:

    bash
    # Get first and last error timestamps
    grep -i "error" sos_commands/logs/journalctl_--no-pager 2>/dev/null | \
      head -1 | awk '{print $1, $2, $3}'
    grep -i "error" sos_commands/logs/journalctl_--no-pager 2>/dev/null | \
      tail -1 | awk '{print $1, $2, $3}'
Step 5: Analyze Application-Specific Logs
  1. Identify application logs:

    bash
    find var/log/ -type f \( -name "*.log" -o -name "*_log" \) 2>/dev/null
  2. Check for stack traces and exceptions:

    bash
    # Python tracebacks
    grep -A 10 "Traceback (most recent call last)" var/log/*.log 2>/dev/null | head -50
    
    # Java exceptions
    grep -B 2 -A 10 "Exception\|Error:" var/log/*.log 2>/dev/null | head -50
  3. Look for common application errors:

    bash
    # Database connection errors
    grep -i "connection.*refused\|connection.*timeout\|database.*error" var/log/*.log 2>/dev/null
    
    # HTTP/API errors
    grep -E "HTTP [45][0-9]{2}|status.*[45][0-9]{2}" var/log/*.log 2>/dev/null | head -20
Show full SKILL.md (431 more words)Show less
Step 6: Generate Log Analysis Summary

Create a structured summary with the following information:

  1. Error Statistics:

    • Total critical errors
    • Total errors
    • Total warnings
    • Time range of errors (first to last)
  2. Critical Findings:

    • Kernel panics (with timestamps)
    • OOM killer events (with victim processes)
    • Segmentation faults (with process names)
    • Service failures (with service names)
  3. Top Error Messages (sorted by frequency):

    • Error message
    • Count
    • First occurrence timestamp
    • Affected component/service
  4. Application-Specific Issues:

    • Stack traces found
    • Database errors
    • Network/connectivity errors
    • Authentication failures
  5. Log File Locations:

    • Provide paths to specific log files for manual investigation
    • Indicate which logs contain the most relevant information

Error Handling

  1. Missing log files:

    • If journalctl logs are missing, fall back to var/log/* files
    • If traditional logs are missing, document this in the summary
    • Some sosreports may have limited logs due to collection parameters
  2. Large log files:

    • For files larger than 100MB, sample the beginning and end
    • Use head -n 10000 and tail -n 10000 to avoid memory issues
    • Inform user that analysis is based on sampling
  3. Compressed logs:

    • Check for .gz files in var/log/
    • Use zgrep instead of grep for compressed files
    • Example: zgrep -i "error" var/log/messages*.gz
  4. Binary log formats:

    • Some logs may be in binary format (e.g., journald binary logs)
    • Rely on sos_commands/logs/journalctl_* text outputs
    • Do not attempt to parse binary files directly

Output Format

The log analysis should produce:

bash
LOG ANALYSIS SUMMARY
====================

Time Range: {first_log_entry} to {last_log_entry}

ERROR STATISTICS
----------------
Critical: {count}
Errors: {count}
Warnings: {count}

CRITICAL FINDINGS
-----------------
Kernel Panics: {count}
  - {timestamp}: {panic_message}

OOM Killer Events: {count}
  - {timestamp}: Killed {process_name} (PID: {pid})

Segmentation Faults: {count}
  - {timestamp}: {process_name} segfaulted

Service Failures: {count}
  - {service_name}: {failure_reason}

TOP ERROR MESSAGES
------------------
1. [{count}x] {error_message}
   First seen: {timestamp}
   Component: {component}

2. [{count}x] {error_message}
   First seen: {timestamp}
   Component: {component}

APPLICATION ERRORS
------------------
Stack Traces: {count} found in {log_files}
Database Errors: {count}
Network Errors: {count}
Auth Failures: {count}

LOG FILES FOR INVESTIGATION
---------------------------
- Primary: {sosreport_path}/sos_commands/logs/journalctl_--no-pager
- System: {sosreport_path}/var/log/messages
- Kernel: {sosreport_path}/var/log/dmesg
- Security: {sosreport_path}/var/log/secure
- Application: {sosreport_path}/var/log/{app_specific}

RECOMMENDATIONS
---------------
1. {actionable_recommendation_based_on_findings}
2. {actionable_recommendation_based_on_findings}

Examples

Example 1: OOM Killer Analysis
bash
# Detect OOM events
grep -B 5 -A 15 "Out of memory" sos_commands/logs/journalctl_--no-pager

# Output interpretation:
# - Which process was killed
# - Memory state at the time
# - What triggered the OOM
Example 2: Service Failure Pattern
bash
# Find failed services
grep "failed to start\|Failed with result" sos_commands/logs/journalctl_--no-pager | \
  awk -F'[][]' '{print $2}' | sort | uniq -c | sort -rn

# This shows which services failed most frequently
Example 3: Timeline of Errors
bash
# Create error timeline
grep -i "error\|fail" sos_commands/logs/journalctl_--no-pager | \
  awk '{print $1, $2, $3}' | sort | uniq -c

# Shows error frequency over time

Tips for Effective Analysis

  1. Start with critical errors: Focus on panics, OOMs, and segfaults first
  2. Look for patterns: Repeated errors often indicate systemic issues
  3. Check timestamps: Correlate errors with the reported incident time
  4. Consider context: Read surrounding log lines for context
  5. Cross-reference: Correlate log findings with resource analysis
  6. Check all log sources: Check both journald and traditional logs, as some events may only appear in one
  7. Document findings: Note file paths and line numbers for reference

Common Log Patterns to Look For

  1. OOM Killer: "Out of memory: Kill process" → Memory pressure issue
  2. Segfault: "segfault at" → Application crash, possible bug
  3. I/O Error: "I/O error" in dmesg → Hardware or filesystem issue
  4. Connection Refused: "Connection refused" → Service not running or firewall
  5. Permission Denied: "Permission denied" → SELinux, file permissions, or ACL issue
  6. Timeout: "timeout" → Network or resource contention
  7. Failed to start: "Failed to start" → Service configuration or dependency issue

See Also

  • Resource Analysis Skill: For correlating log errors with resource constraints
  • System Configuration Analysis Skill: For investigating service failures
  • Network Analysis Skill: For investigating connectivity errors

© openshift-eng, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/sosreport/skills/logs-analysis of openshift-eng/ai-helpers.

Open the folder on GitHubat commit a627176

Compare with similar skills

Logs Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Logs Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Logs Analysis this skillopenshift-eng/ai-helpers120—~2.6kAutomated safety check: PassApache-2.0
Code Design Rationale Investigatorcursor/plugins10k9 repos~2.6kAutomated safety check: PassNone
OpenLogi macOS Permissions TriageAprilNEA/OpenLogi23k—~2.5kAutomated safety check: NotesApache-2.0
Bug Finder for daisyUIsaadeghi/daisyui43k—~2.3kAutomated safety check: PassMIT
Root Cause Debugginggarrytan/gstack136k—~1.4kAutomated safety check: PassMIT
Graph-Based Bug Tracingtirth8205/code-review-graph32k1 repos~287Automated safety check: PassMIT

Similar skills

  • Official

    Digs into why code is shaped the way it is by checking git history, pull requests and connected tools in parallel, then reporting a cited read on the tradeoffs.

    10k GitHub starsUsed in 9 repos~2.6k tokens
    DevelopmentAuto-check passed
  • Decides whether an OpenLogi device problem on macOS is a privacy-permission (TCC) problem, using agent log lines, and says which identity needs which grant.

    23k GitHub stars~2.5k tokensUpdated 4 days ago
    DevelopmentAuto-check: notes
  • Bug Finder for daisyUI

    saadeghi/daisyui

    Investigates suspected bugs in the daisyUI monorepo through read-only analysis, then writes a decision-ready fix plan in tmp/bugs without changing any product code.

    43k GitHub stars~2.3k tokensUpdated 8 days ago
    DevelopmentAuto-check passed
  • Root Cause Debugging

    garrytan/gstack

    Investigates bugs, errors and stack traces in phases and requires a root-cause hypothesis to be confirmed before any fix is written.

    136k GitHub stars~1.4k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Graph-Based Bug Tracing

    tirth8205/code-review-graph

    Traces a bug through a code knowledge graph, following callers, callees and execution flow before opening source files, within a small token budget.

    32k GitHub starsUsed in 1 repo~287 tokens
    DevelopmentAuto-check passed
  • Om Auto Fix Issue

    go-musicfox/go-musicfox

    Fix or implement a tracker issue end to end from a single command — takes an issue id or a plain problem description (filed first via om-prepare-issue), classifies, then drives the bug autofix chain…

    2.6k GitHub starsUsed in 1 repo~5k tokens
    DevelopmentAuto-check: notes

More from openshift-eng/ai-helpers

All 118 skills in this repo
  • Investigate CI Reliability

    openshift-eng/ai-helpers

    Find and independently validate actionable reliability defects across OpenShift release jobs and presubmits, then export portable issue handoffs.

    120 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Address Review PR

    openshift-eng/ai-helpers

    Fetch and address all PR review comments — categorize by priority, make code changes, post replies, and push.

    120 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Categorize Activity Types

    openshift-eng/ai-helpers

    Categorize Jira issues into Red Hat Sankey Activity Type categories using MCP Jira tools.

    120 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Has Review Work

    openshift-eng/ai-helpers

    Decide whether a GitHub PR has unanswered authorized review comments or new required CI failures worth a follow-up agent.

    120 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Must Gather Analyzer

    openshift-eng/ai-helpers

    Analyze OpenShift must-gather diagnostic data including cluster operators, pods, nodes, and network components.

    120 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Payload Autodl JSON

    openshift-eng/ai-helpers

    Schema for the autodl JSON data file produced by payload-analysis for database ingestion — you must use this skill whenever generating the autodl JSON file

    120 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Logs Analysis

What does Logs Analysis do?

Analyze system and application log data from sosreport archives, extracting error patterns, kernel panics, OOM events, service failures, and application crashes from journald logs and traditional…. Logs Analysis is an agent skill from openshift-eng/ai-helpers.

When should I use Logs Analysis?

Logs Analysis fits situations like: tasks that involve Root cause analysis.

How do I install Logs Analysis in Claude Code?

Run `npx skills add openshift-eng/ai-helpers --skill logs-analysis -a claude-code`. Or copy the skill folder (plugins/sosreport/skills/logs-analysis in openshift-eng/ai-helpers) into .claude/skills/logs-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Logs Analysis in Codex?

Run `npx skills add openshift-eng/ai-helpers --skill logs-analysis -a codex`. Or copy the skill folder (plugins/sosreport/skills/logs-analysis in openshift-eng/ai-helpers) into .agents/skills/logs-analysis in your project. Codex loads it when a task matches its description.

Can I use Logs Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add openshift-eng/ai-helpers --skill logs-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/logs-analysis, .gemini/skills/logs-analysis, .github/skills/logs-analysis and .opencode/skills/logs-analysis in your project.

What does Logs Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: Logs Analysis is instructions for the agent only. Our summary lists: Python 3.

Does Logs Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Logs Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Logs Analysis use?

Logs Analysis is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Logs Analysis use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Logs Analysis?

Skills that share tags, products or a category with Logs Analysis: Code Design Rationale Investigator (cursor/plugins, 10k stars), OpenLogi macOS Permissions Triage (AprilNEA/OpenLogi, 23k stars), Bug Finder for daisyUI (saadeghi/daisyui, 43k stars) and Root Cause Debugging (garrytan/gstack, 136k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Logs Analysis?

openshift-eng (a GitHub organization) maintains it in openshift-eng/ai-helpers, which has 120 GitHub stars. The repository holds 118 skills in this directory. The repository was last updated on October 6, 2026.

Source: openshift-eng/ai-helpers on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.