Agent skill

Efficient Web Research

by sickn33 in sickn33/agentic-awesome-skills

Protocol for token-efficient web research. An agent skill from sickn33/agentic-awesome-skills.

MITAuto-check passed

Install Efficient Web Research

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill efficient-web-research -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills efficient-web-research --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/efficient-web-research .claude/skills/efficient-web-research && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
efficient-web-research
GitHub stars
47k
Used in
1 other repo
Token cost
~2.7k tokens
SKILL.md length
898 words
Files
1
Skills in repo
1,394
Repo updated
First seen
Licence
MIT

At a glance

Protocol for token-efficient web research. An agent skill from sickn33/agentic-awesome-skills.

  • Running search queries
  • SKILL.md covers When to Use, Core Principle, Step 1 — Classify the Input and GitHub Protocol, plus 9 more sections
  • Reaches api.github.com

What it does

Efficient Web Research is an agent skill from sickn33/agentic-awesome-skills. Protocol for token-efficient web research. Use when accessing URLs, GitHub repos, or running search queries. Prevents full-page fetching waste.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It works with GitHub. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Running search queries

Example prompts

  • “/efficient-web-research”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 1e53ce2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Efficient Web Research loads about 2.7k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 898 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit 1e53ce2, republished under its MIT licence (© sickn33). 898 words, ~2,745 tokens.

Download SKILL.mdSave it as .claude/skills/efficient-web-research/SKILL.md (or your agent's skills folder).
name
efficient-web-research
description
Protocol for token-efficient web research. Use when accessing URLs, GitHub repos, or running search queries. Prevents full-page fetching waste.
source
community
date_added
2026-09-04
risk
safe

Efficient Web Research Skill

A protocol for accessing web content in the most token-efficient, accurate, and structured way — using the right tool at the right depth, and stopping as soon as the question is answerable.


When to Use

  • Use this skill when the task matches this description: Protocol for token-efficient web research. Use when accessing URLs, GitHub repos, or running search queries. Prevents full-page fetching waste.

Core Principle

Fetch the minimum needed to answer. Skim before you dive. Stop when you can answer.

Every unnecessary fetch wastes tokens and adds noise. This skill enforces a layered approach where you escalate fetch depth only when shallower layers fail.


Step 1 — Classify the Input

Before fetching anything, identify what kind of input you received:

Input TypeExampleGo To
GitHub repo URLgithub.com/user/repoGitHub Protocol
Specific page URLdocs.python.org/3/library/osURL Protocol
Topic / query (no URL)"how does RAFT consensus work"Search Protocol
Multiple URLsList of linksMulti-URL Protocol
PDF / file link.pdf, .txt, .md URLFile Protocol

GitHub Protocol

Use when input is a GitHub URL (repo, file, PR, issue, etc.)

Step 1 — Parse the URL
github.com/{owner}/{repo}                → Repo root
github.com/{owner}/{repo}/tree/{branch}  → Directory
github.com/{owner}/{repo}/blob/{branch}/{path} → Single file
github.com/{owner}/{repo}/issues/{n}     → Issue
github.com/{owner}/{repo}/pull/{n}       → Pull request
Step 2 — Use GitHub API (preferred over scraping)

Always prefer the GitHub API. It returns clean JSON — no HTML parsing needed.

# Repo metadata (name, description, language, stars, topics)
GET https://api.github.com/repos/{owner}/{repo}

# File tree (see what files exist — very cheap)
GET https://api.github.com/repos/{owner}/{repo}/git/trees/{ref}?recursive=1

# Single file content (base64 encoded)
GET https://api.github.com/repos/{owner}/{repo}/contents/{path}?ref={ref}

# README only (usually enough to understand the repo)
GET https://api.github.com/repos/{owner}/{repo}/readme
Step 3 — Layered Fetch for Repos
Layer 1 (always do first):
  → Fetch repo metadata + README only
  → Can you answer the user's question now? YES → STOP. NO → continue.

Layer 2 (only if needed):
  → Fetch file tree to understand structure
  → Identify the 1-3 most relevant files based on the question
  → Can you answer now? YES → STOP. NO → continue.

Layer 3 (last resort):
  → Fetch specific relevant files only (never fetch all files)
  → Prioritize: main entry point, config files, key modules
Token Rules for GitHub
  • README alone answers ~70% of "what does this repo do" questions — always try it first
  • Never fetch more than 3 files in a single research turn
  • If a file exceeds ~300 lines, read only the top (imports + class/function signatures)
  • Decode base64 content from API before passing to context

URL Protocol

Use when the user gives a specific non-GitHub URL (docs, articles, blogs, etc.)

Step 1 — Assess the URL type
Site typeLikely works withNotes
Static docs / MDN / ReadTheDocsread_url_contentFast, clean, cheap
News articles / blogsread_url_contentUsually fine
SPAs / React/Next.js appsbrowser_subagentJS-rendered
Auth-gated pagesbrowser_subagentNeeds login
Raw GitHub files (raw.githubusercontent)read_url_contentDirect text
Step 2 — Layered Fetch
Layer 1 — Skim
  → Fetch the URL with read_url_content
  → Read only headings (H1, H2, H3) and first paragraph
  → Does this page contain what the user needs? NO → try a different URL or search. YES → continue.

Layer 2 — Targeted Extract
  → If the page has anchor links (e.g. /docs/page#section), fetch with the anchor
  → Extract only the relevant section (200–500 tokens max)
  → Can you answer? YES → STOP.

Layer 3 — Full Fetch
  → Fetch full page, strip boilerplate (nav, footer, ads, cookie banners, sidebars)
  → Cap at 2000 tokens. Summarize before passing to answer.

Layer 4 — Browser Subagent (last resort only)
  → Use ONLY if read_url_content returns empty, garbled, or JS-placeholder content
  → Instruct subagent: "Navigate to [URL], wait for content to load, extract [specific section]"
  → Do NOT use browser_subagent for static pages — it's expensive
What to Strip from Fetched Pages

Always remove before using fetched content:

  • Navigation menus and breadcrumbs
  • Cookie banners and GDPR notices
  • "Related articles" / "You might also like" blocks
  • Footer content (copyright, links)
  • Social share buttons
  • Ads and sponsored content

Extract and keep:

  • Main article / documentation body
  • Code blocks
  • Tables with data
  • Numbered steps or procedures

Search Protocol

Use when the user gives a topic, question, or query — not a specific URL.

Step 1 — Sharpen the Query Before Searching

Do NOT search the raw user query. Transform it first:

Raw: "how to deploy fastapi on aws"
Sharpened: "fastapi AWS deployment tutorial 2024"

Raw: "python async vs threads"
Sharpened: "Python asyncio vs threading performance comparison"

Raw: "best way to structure react project"
Sharpened: "React project folder structure best practices"

Query sharpening rules:

  • Add specificity: version numbers, technology names, "tutorial" / "guide" / "comparison"
  • Add recency if relevant: current year
  • Remove filler words: "how do I", "what is the", "can you explain"
  • For code questions: add the language + framework name explicitly
Step 2 — Search and Select
1. Run search_web with the sharpened query
2. Get results (titles + snippets)
3. Scan titles + snippets ONLY — do not fetch yet
4. Pick the TOP 1-2 most relevant results (max 3 in complex cases)
5. Skip results from: forums (if docs exist), aggregator blogs, paywalled sites
6. Prefer: official docs, GitHub repos, well-known tech blogs, academic sources
Step 3 — Fetch Selected Results

Apply the URL Protocol (above) to each selected URL. Process results one at a time — only fetch the second URL if the first didn't answer the question.

  • Never read more than 3 URLs per search query
  • If the snippet already contains the answer → do NOT fetch the full page, use the snippet
  • For factual questions (dates, names, simple facts) → snippet is usually enough
  • For procedural questions (how to do X) → fetch 1 relevant page, targeted section only

Show full SKILL.md (354 more words)Show less

Multi-URL Protocol

Use when the user provides a list of URLs to compare or summarize.

1. Skim all URLs first (Layer 1 fetch for each)
2. Group by relevance to the user's question
3. Deep-fetch only the most relevant 1-3 URLs
4. Summarize each in 3-5 sentences before combining
5. Never dump raw content from multiple pages — always summarize per-source first

File Protocol

Use when URL points directly to a file (PDF, .txt, .md, .csv, etc.)

  • .md / .txt / .csv → read_url_content works directly, read full content
  • .pdf → Use browser_subagent or a PDF extraction tool; extract text only
  • .json / .yaml → read_url_content, parse structure, summarize schema + key values
  • Large files (>500 lines) → Read first 100 lines + last 20 lines + search for relevant sections

Anti-Patterns (Never Do These)

Anti-patternWhy it's badDo this instead
Fetching full page for a simple factWastes 1000s of tokensUse snippet or targeted anchor
Using browser_subagent for static sitesVery expensiveUse read_url_content first
Searching with the raw user queryVague resultsSharpen query first
Fetching 5+ search resultsToken explosionMax 3, stop when answered
Dumping raw HTML into contextNoisy, wastefulAlways strip to Markdown
Fetching "just in case"Unnecessary tokensOnly fetch what's needed to answer
Re-fetching the same URLRedundantCache result in context, reuse
Fetching entire GitHub repoExtremely wastefulREADME + targeted files only

Decision Flowchart (Quick Reference)

Input received
│
├─ GitHub URL?
│   ├─ Fetch README + metadata via API
│   ├─ Answered? → STOP
│   ├─ Need more? → Fetch file tree, pick 1-3 files
│   └─ Still need more? → Fetch specific files only
│
├─ Specific URL?
│   ├─ Try read_url_content → skim headings
│   ├─ Answered? → STOP
│   ├─ Need more? → Targeted section fetch
│   ├─ Still need more? → Full fetch, stripped
│   └─ JS-rendered / broken? → browser_subagent (last resort)
│
├─ Topic/query?
│   ├─ Sharpen query
│   ├─ search_web → scan snippets
│   ├─ Snippet enough? → Answer from snippet, STOP
│   ├─ Need more? → Fetch top 1 result (targeted)
│   └─ Still need more? → Fetch top 2nd result (targeted)
│
└─ List of URLs?
    ├─ Skim all (Layer 1 each)
    ├─ Deep fetch top 1-3 relevant ones
    └─ Summarize per-source, then combine

Output Format Rules

After fetching, structure your response as:

Source: [URL or "Web search for: query"]
Summary: [2-5 sentences of what was found]
Answer: [Direct answer to user's question]
Confidence: [High / Medium / Low — based on source quality]

For multiple sources:

Source 1: ...
Source 2: ...
Combined Answer: ...

Never output:

  • Raw HTML fragments
  • Full page dumps
  • Unattributed information
  • More than needed to answer the question

Token Budget Guide

OperationApproximate token costWhen to use
GitHub README fetch~300–800 tokensAlways first for repos
GitHub API metadata~200 tokensAlways for repos
Skim (headings only)~100–200 tokensAlways first for URLs
Targeted section fetch~300–600 tokensWhen skim isn't enough
Full page fetch (stripped)~1000–2000 tokensOnly when targeted fails
browser_subagent~2000–5000 tokensLast resort only
Search snippet scan~300–500 tokensAlways before fetching

Rule of thumb: If you're about to spend >2000 tokens on a fetch, ask yourself if there's a cheaper path first.


Limitations

  • JavaScript Reliance: Standard fetching may not fully render Single Page Applications (SPAs). You must fallback to browser_subagent for these, which is slower and more expensive.
  • Paywalls & Protections: This skill cannot bypass CAPTCHAs, bot protections (e.g., strict Cloudflare rules), or hard paywalls.
  • GitHub API Limits: Frequent GitHub API requests without authentication may hit rate limits.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/efficient-web-research of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit 1e53ce2

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Efficient Web Research next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Efficient Web Research compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Efficient Web Research this skillsickn33/agentic-awesome-skills47k1 repos~2.7kAutomated safety check: PassMIT
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Diagnosing Superpowers Sessionsobra/superpowers296k3 repos~1.7kAutomated safety check: PassMIT
GitHub Deep Researchbytedance/deer-flow83k5 repos~1.3kAutomated safety check: PassMIT
Greplooponyx-dot-app/onyx32k4 repos~3.3kAutomated safety check: PassMIT
Update V8 Versionopeninterpreter/openinterpreter69k2 repos~845Automated safety check: PassApache-2.0

Similar skills

  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.

    296k GitHub starsUsed in 3 repos~1.7k tokens
    Agent WorkflowsAuto-check passed
  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    83k GitHub starsUsed in 5 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Greploop

    onyx-dot-app/onyx

    Iteratively improves a PR (GitHub), MR (GitLab), or shelved changelist (Perforce) until Greptile gives it a 5/5 confidence score with zero unresolved comments.

    32k GitHub starsUsed in 4 repos~3.3k tokens
    DevelopmentAuto-check passed
  • Update V8 Version

    openinterpreter/openinterpreter

    Bumps the pinned v8 and rusty_v8 versions in Codex, validates the release-candidate path with the v8-canary check, and traces failures to upstream build changes.

    69k GitHub starsUsed in 2 repos~845 tokens
    DevOps & CloudAuto-check passed
  • Check PR

    onyx-dot-app/onyx

    Checks a GitHub, GitLab, or Perforce (p4) pull request (or merge request, or shelved changelist) for unresolved review comments, failing status checks, and incomplete PR descriptions.

    32k GitHub starsUsed in 2 repos~2.3k tokens
    DevelopmentAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,394 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Works with

Questions about Efficient Web Research

What does Efficient Web Research do?

Protocol for token-efficient web research. An agent skill from sickn33/agentic-awesome-skills. Efficient Web Research is an agent skill from sickn33/agentic-awesome-skills. Protocol for token-efficient web research.

When should I use Efficient Web Research?

Efficient Web Research fits situations like: running search queries.

How do I install Efficient Web Research in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill efficient-web-research -a claude-code`. Or copy the skill folder (skills/efficient-web-research in sickn33/agentic-awesome-skills) into .claude/skills/efficient-web-research in your project. Claude Code loads it when a task matches its description.

How do I install Efficient Web Research in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill efficient-web-research -a codex`. Or copy the skill folder (skills/efficient-web-research in sickn33/agentic-awesome-skills) into .agents/skills/efficient-web-research in your project. Codex loads it when a task matches its description.

Can I use Efficient Web Research in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill efficient-web-research -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/efficient-web-research, .gemini/skills/efficient-web-research, .github/skills/efficient-web-research and .opencode/skills/efficient-web-research in your project.

What does Efficient Web Research need to run?

SKILL.md names no scripts, command-line tools or credentials: Efficient Web Research is instructions for the agent only. Our summary lists: Python 3.

Does Efficient Web Research access the network?

SKILL.md names 1 domain. In commands or code: api.github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Efficient Web Research safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Efficient Web Research use?

Efficient Web Research is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Efficient Web Research use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Efficient Web Research?

Skills that share tags, products or a category with Efficient Web Research: PR Babysitter (openinterpreter/openinterpreter, 69k stars), Diagnosing Superpowers Sessions (obra/superpowers, 296k stars), GitHub Deep Research (bytedance/deer-flow, 83k stars) and Greploop (onyx-dot-app/onyx, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Efficient Web Research?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,304 GitHub stars. The repository holds 1,394 skills in this directory. The repository was last updated on October 6, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.