Agent skill

Content Core

by lfnovo in lfnovo/content-core

Extract text content from external sources — URLs, PDFs, documents, YouTube videos, Reddit posts, and audio/video files.

MITAuto-check: notesDocuments & Office

Install Content Core

skills CLI
$ npx skills add lfnovo/content-core --skill content-core -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install lfnovo/content-core content-core --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/lfnovo/content-core.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/content-core .claude/skills/content-core && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
content-core
GitHub stars
174
Token cost
~1.5k tokens
SKILL.md length
512 words
Files
1
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Extract text content from external sources — URLs, PDFs, documents, YouTube videos, Reddit posts, and audio/video files.

  • You need to read
  • SKILL.md covers Purpose, Prerequisites, Capabilities and CLI Usage, plus 3 more sections
  • Calls uvx, uv and curl; reaches astral.sh and youtube.com; needs OPENAI_API_KEY
  • Summarize content from a URL

What it does

Content Core is an agent skill from lfnovo/content-core. Extract text content from external sources — URLs, PDFs, documents, YouTube videos, Reddit posts, and audio/video files. Use when you need to read, analyze, or summarize content from a URL, file, or media source.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office. It works with Reddit, YouTube and Model Context Protocol. The repository describes itself as: Extract what matters from any media source. The licence is MIT.

When your agent uses it

  • You need to read
  • Summarize content from a URL

Example prompts

  • “/content-core”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit c4725c0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uvx
    • uv
    • curl
    • sh
    • brew
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • astral.sh
    • youtube.com
    • reddit.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Content Core loads about 1.5k tokens when it runs. Until then it costs about 56 tokens; SKILL.md has 512 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~56
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePipes a well-known installer script into a shellSKILL.md:27
    - **macOS/Linux**: `curl -LsSf https://astral.sh/uv/install.sh | sh`
  • NotePipes a well-known installer script into a shellSKILL.md:28
    `powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"`

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from lfnovo/content-core at commit c4725c0, republished under its MIT licence (© lfnovo). 512 words, ~1,532 tokens.

Download SKILL.mdSave it as .claude/skills/content-core/SKILL.md (or your agent's skills folder).
name
content-core
description
Extract text content from external sources — URLs, PDFs, documents, YouTube videos, Reddit posts, and audio/video files. Use when you need to read, analyze, or summarize content from a URL, file, or media source.

Purpose

Content Core extracts text from external sources so you can read, analyze, or summarize them. Use it whenever you need content from a URL, PDF, document, YouTube video, Reddit post, or audio/video file.

Most extraction works without API keys. Only audio/video transcription and summarization require an LLM API key (e.g., OPENAI_API_KEY).

Prerequisites

Content Core runs via uvx (zero-install) which requires uv to be available.

Check if uv is installed
bash
uv --version

If uv is not found, help the user install it:

  • macOS/Linux: curl -LsSf https://astral.sh/uv/install.sh | sh
  • Windows: powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
  • Homebrew: brew install uv
  • pip: pip install uv

After installation, the user may need to restart their shell or run source ~/.bashrc / source ~/.zshrc for uv to be available on PATH.

Capabilities

SourceExamplesAPI Key Needed
Web pagesAny URLNo
YouTubeVideo transcript (watch, live, shorts URLs)No
RedditPost + comments via public JSONNo
DocumentsPDF, DOCX, PPTX, XLSX, EPUB, MarkdownNo
AudioMP3, WAV, M4A, FLAC, OGGYes (STT)
VideoMP4, AVI, MOV, MKVYes (STT)
Plain text / HTMLRaw text, auto-detects HTMLNo

CLI Usage

All commands use uvx content-core which runs without installation.

bash
# Check the installed version
uvx content-core --version
Extract content
bash
# From a URL
uvx content-core extract "https://example.com"

# From a file
uvx content-core extract document.pdf

# From a YouTube video (watch, live, and shorts URLs all work)
uvx content-core extract "https://www.youtube.com/watch?v=VIDEO_ID"

# From a Reddit post
uvx content-core extract "https://www.reddit.com/r/sub/comments/POST_ID/title/"

# JSON output (includes title, content, metadata)
uvx content-core extract --format json "https://example.com"

# With a specific extraction engine
uvx content-core extract --engine firecrawl "https://example.com"
uvx content-core extract --engine docling document.pdf
Docling enrichment flags (for advanced document processing)
bash
# Enable formula extraction (LaTeX)
uvx content-core extract --engine docling --formulas paper.pdf

# Enable image descriptions and chart data extraction
uvx content-core extract --engine docling --pictures paper.pdf

# Disable OCR (faster, for PDFs with embedded text)
uvx content-core extract --engine docling --no-ocr paper.pdf
Summarize content

Requires an LLM API key (OPENAI_API_KEY or another provider).

bash
# Summarize text
uvx content-core summarize "Long text here..."

# With context to guide the summary
uvx content-core summarize --context "bullet points" "Long text..."

# Pipe extraction into summarization
uvx content-core extract "https://example.com" | uvx content-core summarize --context "key takeaways"
Configuration
bash
# View current config
uvx content-core config list

# Set persistent defaults
uvx content-core config set llm_provider anthropic
uvx content-core config set llm_model claude-sonnet-5
uvx content-core config set url_engine firecrawl

# Delete a config value
uvx content-core config delete llm_provider

# See all available config keys
uvx content-core config --help

MCP Usage

Content Core can also run as an MCP server. It may or may not be available in your current environment.

Check availability

Look for content-core in the list of available MCP servers. If available, you will have access to these tools:

extract_content

Extracts text from a URL or file. No API key needed for most sources.

extract_content(url="https://example.com")
extract_content(file_path="/path/to/document.pdf")
extract_content(url="https://youtube.com/watch?v=ID")

# With engine override (firecrawl, jina, crawl4ai, simple, docling)
extract_content(file_path="paper.pdf", engine="docling")

# With Docling enrichment
extract_content(file_path="paper.pdf", engine="docling", formulas=true, pictures=true)
summarize_content

Summarizes text using an LLM. Requires an API key.

summarize_content(content="Long text...", context="bullet points")

If summarization fails with an API key error, fall back to extract_content and return the raw content instead.

Show full SKILL.md (213 more words)Show less

Guidelines

  • For small/medium content (articles, short pages): prefer MCP tools if available — they are async and more efficient
  • For large content (long documents, full books, lengthy transcripts): prefer the CLI via a shell, redirecting output to a file (uvx content-core extract "URL" > output.md). This avoids flooding the agent's context window with large payloads. Read only the relevant sections from the file as needed.
  • If MCP is not available, always use the CLI via a shell with uvx content-core
  • For URLs: extraction works without any API key
  • For audio/video: requires OPENAI_API_KEY (or another STT provider key)
  • For summarization: requires an LLM API key
  • When summarization is unavailable, extract the raw content and summarize it yourself
  • Use --format json when you need structured metadata (title, source type, identified type)
  • For large documents with formulas or charts, use --engine docling with --formulas or --pictures

Error Handling

  • If uvx is not found: help the user install uv (see Prerequisites above)
  • If extraction returns empty content: the source may be behind a paywall or require authentication
  • If MCP tools are not available: fall back to CLI via uvx content-core
  • If summarization fails with API key error: use extract_content instead and summarize the content yourself
  • If a specific engine fails: try without --engine to use the auto-detection fallback chain

© lfnovo, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/content-core of lfnovo/content-core.

Open the folder on GitHubat commit c4725c0

Compare with similar skills

Content Core next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Content Core compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Content Core this skilllfnovo/content-core174—~1.5kAutomated safety check: NotesMIT
PullmdAeternaLabsHQ/pullmd486—~2.6kAutomated safety check: PassAGPL-3.0
Influencer Discoverytigerless-labs/influencer-discovery211—~2.5kAutomated safety check: NotesNone
Videodevourdatawhalechina/video-devour156—~3kAutomated safety check: PassApache-2.0
Social PostHao0321/claude-skill-social-post727—~2.9kAutomated safety check: PassMIT
Content Trend Researcheralirezarezvani/claude-code-skill-factory8802 repos~2.1kAutomated safety check: PassMIT

Similar skills

  • Pullmd

    AeternaLabsHQ/pullmd

    Read any web page, document, or YouTube video as clean Markdown using PullMD.

    486 GitHub stars~2.6k tokensUpdated 17 days ago
    Documents & OfficeAuto-check passed
  • Influencer Discovery

    tigerless-labs/influencer-discovery

    Find the bloggers/creators who can help promote your work, capture their contact info, and append them to the target sheet in Google Sheets.

    211 GitHub stars~2.5k tokensUpdated 16 days ago
    Marketing & SEOAuto-check: notes
  • Videodevour

    datawhalechina/video-devour

    使用 VideoDevour 把视频(B站/YouTube/抖音/X 链接、微信视频号分享链接或本地文件)处理成中文图文报告。当用户要求"处理这个视频"、"视频转笔记/报告/图文大纲"、"下载并总结B站/YouTube/抖音/X/视频号视频"时使用。支持搜索视频、查询链接信息、一键生成带关键帧的图文报告(精简/详细),改写成量子速读/公众号文章/小红书笔记、导出…

    156 GitHub stars~3k tokensUpdated 17 days ago
    Documents & OfficeAuto-check passed
  • Social Post

    Hao0321/claude-skill-social-post

    依使用者真實貼文與成效寫 Facebook/Instagram/YouTube/Threads/X 文案,包含 ChatGPT Chat 的「寫文」「Mode C」「用我的格式/口氣」「黑底白字」;本機工作台、規劃、確認後發布、留言回覆及成效學習。使用者說「發文」「文案」「Social Post 介面」「回覆留言」「查流量」「把數據訓練進去」「比較貼文」「優化 pattern」時使用。

    727 GitHub stars~2.9k tokensUpdated 5 days ago
    Writing & ContentAuto-check passed
  • Content Trend Researcher

    alirezarezvani/claude-code-skill-factory

    Advanced content and topic research skill that analyzes trends across Google Analytics, Google Trends, Substack, Medium, Reddit, LinkedIn, X, blogs, podcasts, and YouTube to generate data-driven…

    880 GitHub starsUsed in 2 repos~2.1k tokens
    Writing & ContentAuto-check passed
  • Socialpost

    Hao0321/claude-skill-social-post

    在 ChatGPT 聊天中規劃、撰寫與修改 Facebook、Instagram、Threads、YouTube、X 貼文;處理 Mode C、黑底白字、作者語氣、版型與成效證據。使用者說「發文」「寫文」「文案」「用我的口氣/格式」「Mode C」「黑底白字」「分析貼文」時啟用。

    727 GitHub stars~332 tokensUpdated 5 days ago
    Writing & ContentAuto-check passed

Questions about Content Core

What does Content Core do?

Extract text content from external sources — URLs, PDFs, documents, YouTube videos, Reddit posts, and audio/video files. Content Core is an agent skill from lfnovo/content-core. Extract text content from external sources — URLs, PDFs, documents, YouTube videos, Reddit posts, and audio/video files.

When should I use Content Core?

Content Core fits situations like: you need to read; summarize content from a URL.

How do I install Content Core in Claude Code?

Run `npx skills add lfnovo/content-core --skill content-core -a claude-code`. Or copy the skill folder (skills/content-core in lfnovo/content-core) into .claude/skills/content-core in your project. Claude Code loads it when a task matches its description.

How do I install Content Core in Codex?

Run `npx skills add lfnovo/content-core --skill content-core -a codex`. Or copy the skill folder (skills/content-core in lfnovo/content-core) into .agents/skills/content-core in your project. Codex loads it when a task matches its description.

Can I use Content Core in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lfnovo/content-core --skill content-core -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/content-core, .gemini/skills/content-core, .github/skills/content-core and .opencode/skills/content-core in your project.

What does Content Core need to run?

Going by SKILL.md and its folder, Content Core needs the command-line tools its instructions call (uvx, uv, curl, sh, brew and pip) and credentials named OPENAI_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY.

Does Content Core access the network?

SKILL.md names 3 domains. In commands or code: astral.sh, youtube.com and reddit.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Content Core safe to install?

Our automated static check of SKILL.md found notes only (pipes a well-known installer script into a shell), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Content Core use?

Content Core is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Content Core use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Content Core?

Skills that share tags, products or a category with Content Core: Pullmd (AeternaLabsHQ/pullmd, 486 stars), Influencer Discovery (tigerless-labs/influencer-discovery, 211 stars), Videodevour (datawhalechina/video-devour, 156 stars) and Social Post (Hao0321/claude-skill-social-post, 727 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Content Core?

lfnovo (a GitHub user) maintains it in lfnovo/content-core, which has 174 GitHub stars. The repository was last updated on October 3, 2026.

Source: lfnovo/content-core on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.