Agent skill

Research Material Scout

by huangruiteng in huangruiteng/CS-Notes

A skill your agent uses when the user asks Codex to research, find learning materials, process "素材:" links, "请你读" / "精读" a material, build a material radar, or use SenSight-like broad information…

MITAuto-check passedResearch & Science

Install Research Material Scout

skills CLI
$ npx skills add huangruiteng/CS-Notes --skill research-material-scout -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install huangruiteng/CS-Notes research-material-scout --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/huangruiteng/CS-Notes.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/research-material-scout .claude/skills/research-material-scout && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
research-material-scout
GitHub stars
4k
Token cost
~8.3k tokens
SKILL.md length
4,330 words
Files
4 (incl. scripts, references)
Skills in repo
39
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user asks Codex to research, find learning materials, process "素材:" links, "请你读" / "精读" a material, build a material radar, or use SenSight-like broad information…

  • Works in 5 steps: Pick the relevant ETCLOVG layer first:… → Inspect the candidate repo's README,… → Prefer implementation entries that can… → …
  • The user asks Codex to research
  • SKILL.md covers Mission, Boundary With LoopX Material, Source Strategy and SenSight Backend, plus 7 more sections
  • Runs Python scripts from its folder; calls python3; reaches arxiv.org and github.com

What it does

Research Material Scout is an agent skill from huangruiteng/CS-Notes. Use when the user asks Codex to research, find learning materials, process "素材:" links, "请你读" / "精读" a material, build a material radar, or use SenSight-like broad information retrieval for career learning and Agent infra tracking. Do not use the career/Agent-infra routing bias for user-directed 整理笔记 into a named note.

Its SKILL.md is about 8.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/decision-driven-scouting.md`, `references/paper-reading-protocol.md` and `scripts/multi_source_paper_explore.py`).

It sits in Research & Science. The licence is MIT.

When your agent uses it

  • The user asks Codex to research
  • Find learning materials
  • Process 素材: links
  • 请你读 / 精读 a material

Example prompts

  • “links,”
  • “/research-material-scout”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Pick the relevant ETCLOVG layer first: execution, tooling, context, lifecycle, observability, verification, or governance.
  2. Inspect the candidate repo's README, docs, releases, issues, or paper before ranking it.
  3. Prefer implementation entries that can change the user's artifacts: sandbox boundary, tool registry/schema, state/context contract…
  4. Do not dump many frameworks into Top30. Promote only high-signal projects through the current managed candidate authority, with read…
  5. Treat stars and catalog summaries as recall/ranking hints, not truth.

What it can do on your machine

Read from SKILL.md and the folder at commit f7b4e92. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • arxiv.org
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Research Material Scout loads about 8.3k tokens when it runs, and up to ~14k if it reads all its reference files. Until then it costs about 87 tokens; SKILL.md has 4,330 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~87
When it runs · the whole SKILL.md, loaded when a task matches
~8.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~14k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from huangruiteng/CS-Notes at commit f7b4e92, republished under its MIT licence (© huangruiteng). 4,330 words, ~8,283 tokens.

Download SKILL.mdSave it as .claude/skills/research-material-scout/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
research-material-scout
description
Use when the user asks Codex to research, find learning materials, process "素材:" links, "请你读" / "精读" a material, build a material radar, or use SenSight-like broad information retrieval for career learning and Agent infra tracking. Do not use the career/Agent-infra routing bias for user-directed `整理笔记` into a named note.

Research Material Scout

Use this skill when the user asks for research, material discovery, learning-material triage, or sends links with the 素材: prefix. Treat 调研: as the explicit active-research directive.

Mission

Build a high-signal learning and career material pipeline for the user.

Use the user's current, explicitly confirmed Decision Context to select the next learning action. Read its current version and distinguish user facts, source evidence, and assistant proposals. Career themes are defaults, not permanent weights: new responsibilities, a decision deadline, or a completed reading can change the order. Keep private context out of public queries and skill text.

Before recommending a first read, check existing Notes, managed material records, and the user's latest read-completion statements. Material importance and remaining reading effort are separate: a central paper already understood may need only one missing experiment, a design delta, or a later revisit. Do not invent a completed/archive transition; use the active lifecycle contract and an explicit disposition.

For user-directed note integration, the named target and source's primary domain take precedence over career mapping. Common discovery lanes include agent runtime/evaluation/memory, model learning, serving and RL systems, product workflows, and career signals; select among them from the current task rather than a fixed percentage split.

For decision-driven X discovery or benchmark reading recommendations, read the focused scout protocol.

Boundary With LoopX Material

This skill is the CS-Notes source-discovery and exact-reading adapter. It owns source recall, primary-source verification, reader maps, domain scoring, and the private CS-Notes landing decision.

When the connected project or goal explicitly activates Material Lifecycle, use the project-local managed loopx-material skill for generic store inventory, candidate/archive transitions, lossless migration, ranked-entry rebuild, bounded rerank, owner-gated apply, rollback, and audit. Pass exact-read evidence and the project-specific score into that workflow; do not duplicate or weaken its authority and losslessness gates here.

For CS-Notes, a managed authority pointer under .local/material-lifecycle/authority/current.json means the old candidate and archive Markdown files are immutable legacy content backing. Never append a candidate, archive an item, or edit the old Top30 in those files. Use the project-local loopx-material workflow and the private managed intake adapter. An explicit 素材: request authorizes candidate intake plus a ranking settlement in the same user workflow. Keep intake and ranking as separate receipts and, when ranking changes, separate authority revisions so CAS, preview, rollback, and audit remain independent.

Every intake must end with one explicit disposition: top_window, ranked_backlog, or no_change. Determine high value from exact-read evidence, S/A/B tier, current Decision Context, overlap, and artifact convertibility rather than tier alone. A high-value material must gain verified ranked membership in the Top30 or ranked backlog; it may not remain only an unranked candidate. no_change requires a reason and is reserved for lower-value or substantially duplicative material. If the Top30 is full, move displaced entries into the ranked backlog without loss. Do not move protected anchors without a changed Decision Context or explicit owner instruction.

The installed project skill does not itself activate Material Lifecycle. Without an active project- or goal-scoped profile, declared source authority, and current native owner authorization, this skill may research or read material but must not rewrite the managed store. An already authorized ordinary project source does not require a new Goal or a borrowed agent identity.

Source Strategy

Prefer primary sources for technical conclusions:

  • Papers: arXiv, conference pages, official PDFs.
  • Code: GitHub repositories, README, issues, releases.
  • Products: official docs, release notes, engineering blogs.
  • Internal/user links: Lark/ByteTech only when accessible through user-provided context or approved local tools.

Use social media, 微信公众号, 小红书, X/Twitter, 微博, and aggregators as discovery signals, not final truth.

When extracting mechanisms from a repo, quickstart, prompt, or code file, include file-level links in the user-facing answer and persisted notes/archive. For GitHub sources, prefer commit-pinned permalinks and record the commit read; avoid writing only a repo name or bare filename when a concrete source file drove the claim.

Use SenSight as the primary broad-recall backend when it is available:

  • recent AI/Agent/RL/serving dynamics,
  • social-platform opinions,
  • author/researcher recent posts,
  • link reading across 微信公众号 / 小红书 / X / 微博,
  • keyword monitoring.

Codex remains responsible for verification, ranking, summarization, and local persistence.

Use ordinary web search, platform-specific readers, and Agent-Reach-like local tools as fallback or source-level readers, not as the primary discovery layer.

Implementation-First Catalogs

For Agent Harness / agent infra exploration, use implementation-first catalogs as a source-discovery layer before broad web search when available.

Current primary catalog:

  • Agent Harness Engineering implementation-first catalog: https://github.com/Picrew/awesome-agent-harness

Use it as an index, not as evidence by itself:

  1. Pick the relevant ETCLOVG layer first: execution, tooling, context, lifecycle, observability, verification, or governance.
  2. Inspect the candidate repo's README, docs, releases, issues, or paper before ranking it.
  3. Prefer implementation entries that can change the user's artifacts: sandbox boundary, tool registry/schema, state/context contract, handoff/workflow loop, trace/eval schema, policy/audit layer.
  4. Do not dump many frameworks into Top30. Promote only high-signal projects through the current managed candidate authority, with read status and concrete artifact deltas.
  5. Treat stars and catalog summaries as recall/ranking hints, not truth.
arXiv / Paper Reading Route

For arXiv papers, do not jump straight to PDF extraction unless HTML is unavailable.

Resolution order:

  1. Normalize the paper id and version from any of these forms:
    • https://arxiv.org/abs/<id>
    • https://arxiv.org/pdf/<id>
    • https://arxiv.org/html/<id>vN
    • bare title + discovered arXiv id.
  2. Try arXiv HTML first:
    • if a version is known, open https://arxiv.org/html/<id>vN;
    • if only bare id is known, inspect the abs page for the current version, then try https://arxiv.org/html/<id>v<version>;
    • also try https://arxiv.org/html/<id> if versioned HTML is not obvious.
  3. Use the abs page for metadata, title, authors, abstract, version history, and links.
  4. For 精读, method-heavy papers, or cases where HTML lacks needed appendix, math, algorithm, prompt, or caption detail, try the arXiv TeX source before PDF fallback:
    • download https://arxiv.org/src/<id> into .local/paper-cache/<id>-src.tar.gz;
    • unpack into .local/paper-cache/<id>-src/;
    • locate the entry .tex file, usually main.tex or the file containing \documentclass;
    • recursively inspect included .tex, .bib, figure captions, tables, algorithm blocks, appendices, and prompt/templates.
  5. Fall back to PDF only when HTML/source is missing, blocked, malformed, or lacks the needed figures/tables.
  6. If using PDF fallback, extract text into .local/paper-cache/ and explicitly say that the read path was PDF fallback.

Why this matters:

  • arXiv HTML preserves section anchors, table/figure order, equation context, and is easier for user-side parallel reading.
  • arXiv TeX source often preserves appendices, captions, algorithms, prompts, and bibliography context better than PDF text extraction.
  • PDF extraction can lose figures, captions, math, and table structure; it is acceptable for quick scanning but weaker for 请你读 / 精读.
  • For 请你读 / 精读, always provide the user-facing HTML link when it exists, even if Codex also used the PDF for extraction.
Paper / Research Reading Protocol

For papers, research reports, benchmark papers, method repos, arXiv / OpenReview links, and paper collections, 请你读 / 精读 must use the protocol in references/paper-reading-protocol.md.

Use it as a progressive-disclosure reference rather than copying it into every answer. It synthesizes Keshav's three-pass method, CMU 11-785's paper-reading recitation, and academic / PhD / AI research lenses into a concrete output contract, plus selected ideas from a local scan of research-related skills.

Operational defaults:

  • 请你读 = read first, then answer with Keshav pass 1 plus targeted pass 2 on decision-relevant sections; escalate selected parts to pass 3 only when the material is high-value.
  • 精读 = same output schema, but default to pass 2 plus selective pass 3: virtually reimplement the method, challenge assumptions, and produce artifact deltas.
  • For paper radars or collections, first triage all visible papers with pass 1, then deep-read only the highest-leverage subset.
  • Always state what was actually read: HTML, TeX source, PDF, repo paths, figures/tables, appendix, code, or only metadata.
Paper Radar Pattern

For recurring or paper-heavy exploration, use daily-paper-reader as a design reference, not as a dependency to install. Its useful increment is the pipeline shape: intent profiles -> multi-lane recall -> fusion/rerank -> evidence scoring -> deep/quick/carryover selection.

Adopt these patterns:

  1. Convert the user's current goal into 1-3 intent profiles. Each profile should have:
    • a short tag,
    • 3-6 atomic English keywords for exact/BM25-style recall,
    • 1-4 semantic intent queries for embedding/web search,
    • an explicit source lane such as arXiv, OpenReview, conference proceedings, domain preprint servers, GitHub, official docs, social, or implementation catalog.
  2. Run more than one lane when possible. Combine exact keywords, semantic queries, and source-specific searches instead of relying on one giant query.
    • For AI / agent infra paper search, treat arXiv, OpenReview, and major venue lanes (NeurIPS, ICLR, ICML, ACL, EMNLP, AAAI) as high-value paper sources.
    • For cross-domain or bio/chem/medical-adjacent topics, optionally add bioRxiv, medRxiv, and ChemRxiv as separate lanes.
    • Do not sweep every source by default. Pick lanes that match the topic, then record which lane found each candidate.
  3. Fuse and rank qualitatively:
    • keep at least one strong result from each active lane before global ranking,
    • prefer candidates hit by multiple lanes,
    • require an evidence sentence for every S/A candidate.
  4. Split results into deep, quick, background, and carryover:
    • deep: user should personally read or Codex should deep-read next;
    • quick: Codex summary is enough now;
    • background: useful but not active;
    • carryover: high-signal but not yet processed, keep it visible for the next exploration pass instead of letting daily freshness bury it.
  5. For papers where figures/tables are central, inspect HTML, TeX, PDF figures, tables, or repo assets before ranking when feasible.

Do not copy these parts by default: GitHub Actions / Pages deployment, Supabase schema, remote public embedding/rerank services, front-end panels, API keys, or the exact prompt text. Keep the local skill lightweight and source-agnostic.

Multi-Source Paper Exploration Mesh

For paper-heavy exploration, absorb daily-paper-reader's paper data-source coverage rather than its hosted search/deployment stack. The key improvement is to search several paper-source lanes in parallel and then merge them with Codex's existing source stack.

Default paper lanes:

  • arXiv: fast preprint and recent-paper recall.
  • OpenReview: ICLR / NeurIPS / ICML / AAAI submissions, public reviews, decisions, and withdrawn-public papers when visible.
  • Venue lanes: NeurIPS, ICLR, ICML, ACL, EMNLP, AAAI; use official venue pages, OpenReview, ACL Anthology, AAAI/OJS, proceedings pages, or targeted web search as appropriate.
  • Domain preprint lanes: bioRxiv, medRxiv, ChemRxiv; use only when the topic is bio / medical / chemistry / scientific-agent adjacent.

Combine these with non-paper lanes:

  • GitHub repos, release notes, project pages, Papers with Code-style project pages when available;
  • official product or framework docs;
  • implementation-first catalogs for agent infra;
  • SenSight / social / aggregators as broad recall, never as final fact sources.

Operational rule:

  1. For a focused paper search, choose at least 2 relevant paper lanes plus 1 implementation or docs lane when possible.
  2. For a broader paper radar, choose 3-6 lanes and keep per-lane coverage visible.
  3. Do not wait for a local database or hosted search service to exist. Use official source pages, platform APIs, source-specific search, ordinary web search, and available local readers in parallel.
  4. Rank by cross-lane agreement, primary-source quality, code/data availability, and direct artifact impact.
  5. Record which lane found each S/A candidate and which high-value lane was checked but produced no useful hit.

If this skill has scripts/multi_source_paper_explore.py, use it as the first-pass paper-source recall helper for paper-heavy tasks:

bash
python3 .codex/skills/research-material-scout/scripts/multi_source_paper_explore.py \
  --query "<topic>" \
  --query "<alternate wording or known title>" \
  --sources arxiv,openreview,openalex,biorxiv,medrxiv,chemrxiv,venue-hints

Repeat --query for intent-query expansion. Narrow --sources to topic-fit lanes when domain preprint servers are likely irrelevant.

Use a dual-track exploration flow for real paper-heavy work:

  1. Run the script for structured paper-lane recall and per-lane evidence.
  2. In parallel, run Codex's existing web / SenSight / GitHub / official-doc search for project pages, repos, social/aggregator leads, implementation evidence, and missed terminology.
  3. Merge the two result sets before ranking. Treat script-only hits as candidates to verify, and web-only hits as leads to trace back to primary papers / code / docs.

The script is only a recall layer; it complements rather than replaces the older exploration stack.

SenSight Backend

Discover callable SenSight tools or the current project's private .local/sensight-skill-source/sensight directory. Do not assume a previous machine's absolute path or installed version. If the source exists, read its SKILL.md and actual command help before invoking it. Do not execute placeholder installers.

When SenSight is absent or unavailable, record that limitation and continue with web search plus an authenticated platform reader and primary-source verification. Optional discovery tooling should not block already authorized research. If its authentication is essential, follow its current authorization flow; never treat an auth response as source content.

Useful Actions

Use these actions as retrieval, not as final authority:

NeedSenSight action
AI industry deep dive / high-quality articlesretrieve_summarize
Latest AI papersdaily_paper
Latest AI/company technical blogsdaily_blog
Weekly model releasesweekly_model
Model reputation / user sentimentmodel_sentiment
Hot events, general news, trend searchsearch_events
Platform hot listsget_event_board
Social semantic search across X/小红书/微博/公众号social_search
Recent posts from a specific author/accountsearch_author_posts

Examples:

bash
python3 scripts/sensight.py retrieve_summarize \
  --query "Agent infra 最新进展" \
  --enhance_query "最近一周 Agent infra、OpenClaw、Claude Code、agent memory、long-running coding agent 的高质量技术动态" \
  --size 20 \
  --result_form article_summary

python3 scripts/sensight.py social_search \
  --query "GPT 5.4 评价" \
  --platforms 1 2 3 4 \
  --size 20

python3 scripts/sensight.py search_author_posts \
  --platform 1 \
  --author_name "Anthropic"

Use public metadata when sufficient; for incomplete posts, replies, or signed-in search, read the installed ego-browser skill and follow its current TaskSpace and API contract. Do not copy stale browser methods into this skill. Reuse the user's sign-in; inspect visible posts and original links, and close only task-owned surfaces.

Capture relevant author, date, post URL, primary-source links, read scope, and the claim to verify. Avoid unrelated recommendations, DMs, account statistics, or a bulk feed dump. Inspect needed images through supported browser/media tools; do not download remote media to work around display restrictions. Store private evidence under .local/. A social post can be a first-party statement of its author's claim, but reported experimental results still require the paper, code, or data protocol.

Agent-Reach Complement

Agent-Reach (https://github.com/Panniantong/Agent-Reach) is useful as a complementary local scaffolding layer, especially when SenSight is unavailable, too aggregated, or lacks a channel.

Use it as a design reference or optional install, not the default primary source.

What it adds:

  • source-level tools rather than platform-side aggregation,
  • web reading through Jina Reader,
  • YouTube / Bilibili transcript extraction through yt-dlp,
  • GitHub through gh,
  • RSS through feedparser,
  • Reddit / Twitter / 小红书 / 抖音 / LinkedIn via separate upstream CLIs or MCP tools,
  • agent-reach doctor style capability diagnostics.

When it helps more than SenSight:

  • Need to inspect a specific URL/video/repo/thread, not just discover candidates.
  • Need an open-source, auditable local route.
  • Need video subtitles, RSS feeds, GitHub issues/PRs, or Reddit threads.
  • SenSight auth is unavailable or results are too summarized.

When SenSight should stay primary:

  • Broad topic discovery.
  • Recent social sentiment.
  • Cross-platform opinion summaries.
  • AI papers/blogs/model-release radar.
  • Low-maintenance material scouting.

Do not install Agent-Reach automatically unless the user asks. It may install many dependencies and configure cookies/proxies. If installed, keep secrets/cookies local and never commit them.

Intake Workflow

Directive convention:

  • 素材:<link/text> means intake. Read, classify, summarize, preserve the original link, and append one managed candidate through exact-read evidence, immutable content backing, authority CAS, readback, and receipt.
  • 调研:<question/topic> means active research. Use broad recall plus source verification, then write high-signal candidates and recommendations.
  • 整理笔记:<link/text> or "整理笔记 + named note/theme" means direct note integration. The output surface is Notes/, not the candidate library by default. Named target and source-domain taxonomy win over career priority.
  • 请你读:<material id/link/title> means Codex reads first, then returns an illustrated mechanism-first summary and a reader map in the conversation. Do not organize or edit notes during this command. For papers / research artifacts, load references/paper-reading-protocol.md and follow its output contract. 精读 is a compatibility alias with the same output boundaries, but defaults to deeper pass-3 reconstruction when warranted.
  • 读完:<material id/link/title + user notes> means close the reading loop. Update the best Notes/ landing before archiving unless the content is private-only. Treat user-highlighted points as retention requirements: concrete prompts, env flags, schema fields, tool/API names, figures/tables, failure cases, doubts, and comparison phrases should be preserved in the public note when safe and conceptually useful; .local may keep raw/private/full detail, but must not be the only landing for points the user explicitly asked to remember.
  • 继续调研 means continue the latest active-research theme, but only if adding new sources or a new decision-relevant synthesis.

For each user-provided material:

  1. Classify the request intent before choosing tools or files:
    • 整理笔记: use the repository note-integration workflow. If the user named a file/section, inspect that target first; if not, classify the source's primary contribution such as model algorithm, training method, inference system, agent runtime, product strategy, or career signal, then find the matching note. Do not default to Agent infra just because the material is AI-related.
    • 素材: intake to the current managed material authority. Do not write the legacy candidate Markdown.
    • 调研: prioritize by the user's current confirmed Decision Context.
    • 请你读 / 精读: read and explain in the conversation, with relevant images displayed inline. Save only source caches, figure assets and reading evidence under .local/ or the task artifact directory; do not edit Notes/, its indexes, or the material lifecycle/ranking. Note integration requires an explicit 整理笔记 or 读完 request; an artifact delta here is a proposal, not permission to implement it.
Show full SKILL.md (1,592 more words)Show less
Illustrated Reading Output

For 请你读 / 精读, show relevant images in the final answer, not only image links or a claim that figures were inspected. Prefer a small selection of source figures, table screenshots or source-rendered diagrams that explain the core mechanism and evidence. Inspect each image and explain its labels, axes, main comparison and limitations next to it. Cite the original figure/page. If the source has no suitable readable figure, create a faithful explanatory diagram and explicitly label it as a Codex illustration, not an original figure or measured result. Never fabricate unseen figures; if source access blocks an image, state the gap. Keep reading assets outside Notes/ until note integration is explicitly requested.

User-Highlighted Detail Preservation

When the user provides a numbered/bulleted readout, quoted phrase, prompt snippet, schema, env var, failure case, or says "这个值得作为专题section / 概念级别 / 这个点要记", treat it as first-class source material rather than optional color.

The note-writing rules below apply to 整理笔记 / 读完. During 请你读 / 精读, cover the user's details in the illustrated conversation answer; do not turn detail preservation into unsolicited note edits.

  • Account for every explicit user point before closing the turn: either in Notes/, in .local with a privacy/version reason, or intentionally skipped with a reason reported to the user.
  • Prefer distilled-but-concrete preservation in Notes/: short prompt excerpts or paraphrases, tool names, env flags, schema fields, benchmark names, mode names, key tables/figures, and caveats. Do not over-compress them into only an abstract framework.
  • If a detail is version-sensitive or platform-specific, label it as such and preserve the source/version boundary; do not hide it solely in .local.
  • If the user asks for a "专题section" or "概念级别", promote or create an independent section instead of burying the point under a neighboring topic.
  • .local is for raw cache, private URLs, full prompt copies, and sensitive/internal detail. It supplements public notes; it is not a substitute for durable Notes/ synthesis.
  1. Preserve the original link in the final note title.
    • If the URL contains disposable login tokens, access tokens, auth codes, session IDs, or other sensitive query parameters, use them only for reading and persist only the stable URL with those parameters stripped. Note that the original link was sanitized.
  2. Try the best reader first:
    • 微信公众号 -> wechat-article-reader
    • 小红书 -> xiaohongshu-reader
    • 飞书 / Lark -> lark-doc / lark-wiki
    • arXiv / paper -> arXiv HTML route first, abs metadata second, PDF fallback last
    • GitHub -> GitHub tools or gh
    • X/Twitter concrete link -> public oEmbed first, then ego-browser authenticated snapshot / DOM extraction if needed
    • Web pages -> official web search / browser / Playwright as needed
  3. If unreadable, mark as Unread and ask for pasted text, screenshot, export, or accessible copy.
  4. Classify into S/A/B/Unread:
    • S: user should personally read and convert into an artifact.
    • A: Codex summary is enough unless the theme becomes active.
    • B: useful background, tool lead, or product observation.
    • Unread: not read; never pretend.
  5. If Material Lifecycle is active, write exact-read evidence and staged content, then use the active project source adapter documented in .local/material-lifecycle/README.md to intake one candidate or revise its existing stable ref. Verify an intake adds exactly one record, or a revision preserves identity, lifecycle and membership, then create the separate Decision Context-backed ranking revision in the same workflow. Verify the disposition, ranked membership when required, authority readback, readable projection, audit receipt, and rollback path before reporting completion.

Active Research Workflow

When proactively finding materials:

  1. Start from the user's current goals, not generic trends.
  2. For paper-heavy or recurring themes, draft a small intent-profile plan first: tags, exact keywords, semantic queries, and source lanes. Include arXiv / OpenReview / venue lanes (NeurIPS, ICLR, ICML, ACL, EMNLP, AAAI) when relevant, and add domain preprint lanes (bioRxiv, medRxiv, ChemRxiv) only when the topic warrants them. Then query across at least two source types when possible: paper/code/docs/social.
  3. For paper-heavy tasks, run multiple selected paper-source lanes in parallel where possible, then merge them with code/docs/social/implementation lanes. Do not make a local database or hosted search service a prerequisite for exploration.
  4. Prefer fewer, higher-quality materials over broad dumps. Keep per-lane coverage visible so one popular source does not crowd out a strategically important niche source.
  5. For each candidate, capture:
    • title and original URL,
    • source type and read status,
    • query/profile lane that found it,
    • one-paragraph summary,
    • evidence sentence for the ranking,
    • why it matters to the user's career goal,
    • recommended action.
  6. If a social/aggregator/source-search item points to a paper or repo, follow the paper/repo before ranking.
  7. End with a selection split: S/A/B/Unread plus deep / quick / background / carryover; carryover items must have a reason and a next trigger.

Self-Verification

Before reporting that a research task is done:

  1. Verify the retrieval backend state:
    • paper-source lanes checked, skipped as not relevant, or unavailable;
    • SenSight result received, or
    • SenSight auth-blocked and fallback source path used, or
    • SenSight not relevant for this specific URL/material.
  2. Check at least two source types for active research whenever possible, such as paper + repo, official docs + social discussion, or product page + engineering blog.
  3. For every S/A candidate, include:
    • original URL,
    • read status,
    • source type,
    • query/profile lane,
    • evidence sentence,
    • why it matters to the user's current confirmed Decision Context,
    • next action.
  4. Confirm the candidate stable ref exists exactly once in the current managed catalog, the immutable content backing digest verifies, and an intake receipt records CAS/readback success.
  5. For 请你读 / 精读, check the final answer follows references/paper-reading-protocol.md when applicable, visibly includes relevant images with source attribution and explanation, and contains a concrete "用户本人还需要读什么" reader map. Verify that no note/index or lifecycle/ranking edits were made as part of reading. If the answer is "不用读原文", still name the inspected sections and provide a substitute-quality digest.
  6. For tool / standard / API / framework bundles, check that the answer starts with background and workflow introduction: why this thing exists, what pain it solves, what breaks without it, and how each component is positioned. Then explain each component as a standalone material before mapping to Agent Harness / OpenViking. Do not start directly from jargon, fields, or claim maps, and do not let the project mapping crowd out the source-content explanation.
  7. For 读完 / note integration, run a user-highlighted point audit: every explicit bullet, numbered item, prompt snippet, schema field, env flag, comparison phrase, or doubt from the user is either present in Notes/, present only in .local with a privacy/version reason, or intentionally skipped with a reason reported.
  8. If any source could not be read, say so explicitly and ask for paste/screenshot/export only when necessary.

Quality Bar

Reject or demote materials that are:

  • pure hype without implementation detail,
  • duplicate commentary on already captured material,
  • unrelated to the user's current decision and learning priorities,
  • not traceable to a primary source when factual claims matter.

Output Style

Be concise. Tell the user what was added, where it was added, and the key judgment.

When writing into Notes/:

  • Use Typora-friendly block math for real equations: $$...$$.
  • Do not put formulas in ```text code fences; reserve code fences for schemas, field lists, commands, and pseudocode.
  • If a figure from the source or user-provided material is essential to understanding the mechanism, save it into the target note's relative asset folder (for example Notes/AI-Applied-Algorithms/) and link it with a relative Markdown path. Prefer primary-source figures when available.
  • For high-value cross-layer case studies, do not create a monolithic material section by default. First extract the general mechanism into the highest-level framework section, then add only small local deltas to existing eval / memory / runtime / tooling sections. Keep the source case as evidence, not as the organizing axis.
  • Preserve user-highlighted concrete details in Notes/ when safe: prompt snippets, command/env flags, schema fields, API/tool names, mode names, version boundaries, and failure cases. Prefer short excerpts or paraphrases over long raw prompt dumps, and mark version-sensitive/platform-specific details instead of silently dropping them.

For 请你读 / 精读, Codex should read first and then provide a mechanism-first guide rather than a broad reading plan. Because personal original-reading recommendations are conservative, Codex-summary-enough answers must be more detailed, not thinner: include enough background, source-content explanation, core design, fields/schemas, evidence, artifact mapping, and caveats to substitute for the user's first-pass read. Use a two-focus structure: first explain the material itself, then map it to the user's current artifact. For papers / research artifacts, load references/paper-reading-protocol.md; the short form below is the minimum answer shape:

  1. 一句话判断.
  2. 背景和工作流介绍: why this material exists, what pain it solves, and where it sits in the broader ecosystem.
  3. 我实际读了什么: HTML/PDF/repo/code/figures/tables/appendix, plus unread parts.
  4. 原材料内容卡: for each paper/tool/standard/repo in the bundle, explain its standalone purpose, core abstraction, important fields/APIs, common usage, and limits.
  5. Claim map: main claims, evidence, confidence, and what would make each claim false.
  6. 精要内容和核心设计: problem, boundary, input/output, data/interface format, workflow, metrics, baselines, main results, limitations, and key figure/table/code path.
  7. 核心机制: 3-6 numbered mechanisms, each with "what the author does/proves" and "how the user should interpret it".
  8. 对用户 artifact 的改造建议: schema, feedback signal, benchmark variant, TODO, steering, or interview/deep-dive line; discuss proposals without implementing changes during reading.
  9. 用户本人还需要读什么: mandatory reader map with concrete original sections, figures, tables, code paths, and a decision for each: must-read, optional, or skippable. Do not only say "读摘要即可"; name the exact parts that justify that decision.
  10. 边读边核验的问题: 3-6 sharp checks, especially leakage, counterfactual reliability, metric validity, transferability to Agent Harness / TAU2.
  11. Display the relevant images inline with captions and reading guidance. Do not organize notes or archive during 请你读 / 精读; integrate notes after explicit 整理笔记 / 读完, and archive only after 读完.

For high-value materials, include the next concrete action, such as:

  • "精读并产出一页 design delta",
  • "由 Codex 先读论文 PDF 并摘要",
  • "只保留为产品观察",
  • "转成 agent-harness TODO / benchmark idea".

© huangruiteng, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in .codex/skills/research-material-scout of huangruiteng/CS-Notes.

  • SKILL.md
  • references/decision-driven-scouting.md
  • references/paper-reading-protocol.md
  • scripts/multi_source_paper_explore.py

Open the folder on GitHubat commit f7b4e92

Compare with similar skills

Research Material Scout next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Research Material Scout compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Research Material Scout this skillhuangruiteng/CS-Notes4k—~8.3kAutomated safety check: PassMIT
Hypothesis Generationspacering-net/codeg3.9k14 repos~3.6kAutomated safety check: NotesMIT
GitHub Deep Researchbytedance/deer-flow84k4 repos~1.3kAutomated safety check: PassMIT
Nature Paper CardYuan1z0825/nature-skills47k2 repos~2.1kAutomated safety check: PassApache-2.0
Content Research Writerweapp-tailwindcss/weapp-tailwindcss1.9k25 repos~3.5kAutomated safety check: PassMIT
Last30daysmvanhorn/last30days-skill64k—~7.9kAutomated safety check: NotesMIT

Similar skills

  • Hypothesis Generation

    spacering-net/codeg

    Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.

    3.9k GitHub starsUsed in 14 repos~3.6k tokens
    Research & ScienceAuto-check: notes
  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    84k GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Nature Paper Card

    Yuan1z0825/nature-skills

    Builds a structured deep-reading card for one scientific paper, covering methods, how experiments support claims, limitations and research ideas, with a script to prepare the source.

    47k GitHub starsUsed in 2 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Content Research Writer

    weapp-tailwindcss/weapp-tailwindcss

    Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section.

    1.9k GitHub starsUsed in 25 repos~3.5k tokens
    Research & ScienceAuto-check passed
  • Last30days

    mvanhorn/last30days-skill

    Research what people actually say about any topic in the last 30 days.

    64k GitHub stars~7.9k tokensUpdated yesterday
    Research & ScienceAuto-check: notes
  • Peer Review

    spacering-net/codeg

    Structured manuscript/grant review with checklist-based evaluation.

    3.9k GitHub starsUsed in 17 repos~5.9k tokens
    Research & ScienceAuto-check: notes

More from huangruiteng/CS-Notes

All 39 skills in this repo
  • CLI Creator

    huangruiteng/CS-Notes

    Build a composable CLI for Codex from API docs, an OpenAPI spec, existing curl examples, an SDK, a web app, an admin tool, or a local script.

    4k GitHub starsUsed in 2 repos~2.7k tokens
    Auto-check passed
  • Codex Thread Heartbeat

    huangruiteng/CS-Notes

    Inspect and manage guarded Codex App-native or launchd heartbeats for Codex main control threads.

    4k GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check passed
  • Slack

    huangruiteng/CS-Notes

    A skill your agent uses when you need to control Slack from Clawdbot via the slack tool, including reacting to messages or pinning/unpinning items in Slack channels or DMs.

    4k GitHub starsUsed in 10 repos~578 tokens
    Auto-check passed
  • Codex Thread Reader

    huangruiteng/CS-Notes

    Locate and read a Codex thread by a codex thread link, thread id, or rollout path across all local CODEXHOME directories (~/.codex, ~/.codex-gpt, ...).

    4k GitHub stars~1.4k tokensUpdated 2 days ago
    Auto-check passed
  • GitHub

    huangruiteng/CS-Notes

    Interact with GitHub using the gh CLI. An agent skill from huangruiteng/CS-Notes.

    4k GitHub starsUsed in 27 repos~279 tokens
    Auto-check passed
  • AI Hotspots

    huangruiteng/CS-Notes

    Track and synthesize current AI hotspots into a bilingual HTML daily report.

    4k GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check passed

Questions about Research Material Scout

What does Research Material Scout do?

A skill your agent uses when the user asks Codex to research, find learning materials, process "素材:" links, "请你读" / "精读" a material, build a material radar, or use SenSight-like broad information…. Research Material Scout is an agent skill from huangruiteng/CS-Notes. Use when the user asks Codex to research, find learning materials, process "素材:" links, "请你读" / "精读" a material, build a material radar, or use SenSight-like broad information retrieval for career learning and Agent infra tracking.

When should I use Research Material Scout?

Research Material Scout fits situations like: the user asks Codex to research; find learning materials; process 素材: links; 请你读 / 精读 a material.

How do I install Research Material Scout in Claude Code?

Run `npx skills add huangruiteng/CS-Notes --skill research-material-scout -a claude-code`. Or copy the skill folder (.codex/skills/research-material-scout in huangruiteng/CS-Notes) into .claude/skills/research-material-scout in your project. Claude Code loads it when a task matches its description.

How do I install Research Material Scout in Codex?

Run `npx skills add huangruiteng/CS-Notes --skill research-material-scout -a codex`. Or copy the skill folder (.codex/skills/research-material-scout in huangruiteng/CS-Notes) into .agents/skills/research-material-scout in your project. Codex loads it when a task matches its description.

Can I use Research Material Scout in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add huangruiteng/CS-Notes --skill research-material-scout -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/research-material-scout, .gemini/skills/research-material-scout, .github/skills/research-material-scout and .opencode/skills/research-material-scout in your project.

What does Research Material Scout need to run?

Going by SKILL.md and its folder, Research Material Scout needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Research Material Scout access the network?

SKILL.md names 2 domains. In commands or code: arxiv.org and github.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Research Material Scout safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Research Material Scout use?

Research Material Scout is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Research Material Scout use?

About 8.3k tokens (SKILL.md is roughly 33k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.4k tokens, read only when the agent opens those files.

What are the alternatives to Research Material Scout?

Skills that share tags, products or a category with Research Material Scout: Hypothesis Generation (spacering-net/codeg, 3.9k stars), GitHub Deep Research (bytedance/deer-flow, 84k stars), Nature Paper Card (Yuan1z0825/nature-skills, 47k stars) and Content Research Writer (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Research Material Scout?

huangruiteng (a GitHub user) maintains it in huangruiteng/CS-Notes, which has 4,001 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on October 8, 2026.

Source: huangruiteng/CS-Notes on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.