Chroma Vector Database
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
Ancient text restoration, attribution, dating, contextualization, and embedding via Aeneas (Latin) / Ithaca (Ancient Greek).
$ npx skills add google-deepmind/science-skills --skill predictingthepast -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install google-deepmind/science-skills predictingthepast --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/predictingthepast .claude/skills/predictingthepast && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "predictingthepast" agent skill from https://github.com/google-deepmind/science-skills/tree/main/skills/predictingthepast into .claude/skills/predictingthepast/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "predictingthepast", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/google-deepmind/science-skills/tree/main/skills/predictingthepastType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add google-deepmind/science-skills --skill predictingthepast -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install google-deepmind/science-skills predictingthepast --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/predictingthepast .agents/skills/predictingthepast && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "predictingthepast" agent skill from https://github.com/google-deepmind/science-skills/tree/main/skills/predictingthepast into .agents/skills/predictingthepast/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "predictingthepast", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google-deepmind/science-skills --skill predictingthepast -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install google-deepmind/science-skills predictingthepast --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/predictingthepast .cursor/skills/predictingthepast && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "predictingthepast" agent skill from https://github.com/google-deepmind/science-skills/tree/main/skills/predictingthepast into .cursor/skills/predictingthepast/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "predictingthepast", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/google-deepmind/science-skills.git --path skills/predictingthepast--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add google-deepmind/science-skills --skill predictingthepast -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install google-deepmind/science-skills predictingthepast --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/predictingthepast .gemini/skills/predictingthepast && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "predictingthepast" agent skill from https://github.com/google-deepmind/science-skills/tree/main/skills/predictingthepast into .gemini/skills/predictingthepast/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "predictingthepast", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install google-deepmind/science-skills predictingthepastInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add google-deepmind/science-skills --skill predictingthepast -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/predictingthepast .github/skills/predictingthepast && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "predictingthepast" agent skill from https://github.com/google-deepmind/science-skills/tree/main/skills/predictingthepast into .github/skills/predictingthepast/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "predictingthepast", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google-deepmind/science-skills --skill predictingthepast -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install google-deepmind/science-skills predictingthepast --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/predictingthepast .opencode/skills/predictingthepast && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "predictingthepast" agent skill from https://github.com/google-deepmind/science-skills/tree/main/skills/predictingthepast into .opencode/skills/predictingthepast/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "predictingthepast", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
predictingthepastAncient text restoration, attribution, dating, contextualization, and embedding via Aeneas (Latin) / Ithaca (Ancient Greek).
Predictingthepast is an agent skill from google-deepmind/science-skills. Ancient text restoration, attribution, dating, contextualization, and embedding via Aeneas (Latin) / Ithaca (Ancient Greek). Use when asked to "restore", "attribute", "date", "contextualize", "find parallels", "where was it written", "when was it written", "embed", or "analyze" an ancient text, inscription, or epigraphic document, or when the user mentions "Aeneas", or "Ithaca".
Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including scripts and reference files (for example `references/iphi-region-sub-loc.json`, `references/led-region-sub-loc.json` and `references/output_format.md`).
It sits in AI & LLM Engineering, covering Embeddings. The repository describes itself as: GDM Science Skills to speed up agentic scientific workflows with better grounding and higher token efficiency. Integrate insights from AlphaGenome, AFDB, UniProt and 30+ other… The licence is Apache-2.0.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 6883275. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
zenodo.orggithub.compredictingthepast.cominscriptions.packhum.orgedr-edr.itedh.ub.uni-heidelberg.demanfredclauss.deFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Predictingthepast loads about 3.7k tokens when it runs, and up to ~93k if it reads all its reference files. Until then it costs about 100 tokens; SKILL.md has 1,480 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from google-deepmind/science-skills at commit 6883275, republished under its Apache-2.0 licence (© google-deepmind). 1,480 words, ~3,653 tokens.
.claude/skills/predictingthepast/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.Aeneas (Latin) and Ithaca (Ancient Greek) perform four tasks on ancient texts:
uv: Read the uv skill and follow its Setup instructions to ensure
uv is installed and on PATH.
User Notification: If .licenses/predictingthepast_LICENSE.txt does not already exist in the workspace root directory then (1) prominently notify the user to check the terms at https://github.com/google-deepmind/predictingthepast/blob/main/README.md#license--disclaimer, and include the appropriate citation and the full dataset acknowledgement, and that use of these datasets should acknowledge and cite the original data sources. Then (2) create the file recording the notification text and timestamp.
preprocess.py, run_inference.py,
visualize_results.py). Present model output as-is — never supplement or
override it with external lookups.Present the restoration markup characters, then ask the user for their text:
?:donat in ??????????rtis#:donat in #rtis-:prolixin---s fecit_:prolixin_s fecitAfter presenting this list, ask the user to provide the text they want to submit for analysis.
Clean input text before inference:
uv run <SKILL_DIR>/scripts/preprocess.py \
--language=latin \
--input="raw text here..."Or from a file:
uv run <SKILL_DIR>/scripts/preprocess.py \
--language=greek \
--input_file=/tmp/input.txt \
--output_file=/tmp/cleaned.txt0,
strips editorial brackets [] and (), removes punctuation, filters to
valid chars (abcdefghiklmnopqrstuvxyz plus 0 . - _ ? # <space>)0,
applies PHI cleaning (bracket normalization, sigma conversion), filters to
Greek alphabet (αβγδεζηθικλμνξοπρςστυφχψωϛ plus 0 . - _ ? # <space>)- if shorter).##. No adjacent ?# or #?.? sequences count toward total.#, ask how many characters to restore and set
--restore_max_len accordingly.Confirm with the user before proceeding if either applies:
?
characters, or uses # with --restore_max_len > 10, warn: "This
restoration involves N characters which will take approximately M minutes
(restoration time scales roughly linearly ~10 s per additional ? on a
high-end CPU machine: 5 → ~1 min, 10 → ~2.5 min, 20 → ~5 min, 30 → ~8 min).
Do you want to proceed, or simplify the query first (e.g. fewer ? marks,
shorter --restore_max_len, or restoring section by section)?"These factors compound: a complex restoration across multiple windows will be substantially slower than either factor alone.
Each task is controlled by its own flag. At least one must be provided:
--attribute — geographical + chronological attribution--restore — text restoration (requires ? or # in input)--contextualize — parallel inscription retrievalAny combination is valid. All three can be used together.
When --embedding is provided, a text embedding vector is also generated
alongside the other tasks.
# Attribution + Restoration (text with gaps)
uv run <SKILL_DIR>/scripts/run_inference.py \
--language=latin \
--input="cleaned text with ???" \
--attribute --restore \
--output_json=/tmp/results.json
# Attribution + Contextualization (no gaps)
uv run <SKILL_DIR>/scripts/run_inference.py \
--language=latin \
--input="cleaned text" \
--attribute --contextualize \
--output_json=/tmp/results.json
# All tasks
uv run <SKILL_DIR>/scripts/run_inference.py \
--language=latin \
--input="cleaned text with ???" \
--attribute --restore --contextualize \
--output_json=/tmp/results.json
# Generate dashboard
uv run \
<SKILL_DIR>/scripts/visualize_results.py \
--input=/tmp/results.json \
--output=/tmp/dashboard.htmlOutputs are always returned to the console (JSON/Summary) and must always be written to an HTML dashboard for rich visualization.
Serving HTML artifacts: After generating any HTML file (the dashboard or
any other HTML artifact you create), start a background HTTP server so the user
can view it in their browser. Run the server in the foreground of the background
task process (without trailing &) to prevent the process from being
terminated. Use a random high port to avoid conflicts.
In the example below, replace <HTML_FILE> with the absolute path of the HTML
file you produced (e.g. /tmp/dashboard.html or C:\tmp\dashboard.html).
For Linux/macOS:
# Pick an unused port and start the server as a persistent process
HTML_FILE=/tmp/dashboard.html # ← set to the actual HTML file path
PORT=$(export PATH="$HOME/.local/bin:$PATH" && uv run python -c "import socket; s=socket.socket(); s.bind(('',0)); print(s.getsockname()[1]); s.close()")
FILENAME=$(basename "$HTML_FILE")
echo "View URL: http://localhost:${PORT}/${FILENAME}"
export PATH="$HOME/.local/bin:$PATH" && uv run python -m http.server "$PORT" --bind 0.0.0.0 --directory "$(dirname "$HTML_FILE")"For Windows (PowerShell):
# Pick an unused port and start the server as a persistent process
$HTML_FILE = "C:\tmp\dashboard.html" # ← set to the actual HTML file path
$PORT = & "$HOME\.local\bin\uv" run python -c "import socket; s=socket.socket(); s.bind(('',0)); print(s.getsockname()[1]); s.close()"
$FILENAME = Split-Path $HTML_FILE -Leaf
$DIRNAME = Split-Path $HTML_FILE -Parent
Write-Host "View URL: http://localhost:${PORT}/${FILENAME}"
& "$HOME\.local\bin\uv" run python -m http.server $PORT --bind 0.0.0.0 --directory $DIRNAMEThen provide http://localhost:<port>/<filename> (or
http://127.0.0.1:<port>/<filename>) as the primary clickable link in your
response. Additionally, provide http://<hostname>:<port>/<filename> as a
fallback link for remote setups. Do not give the user a raw file:// path.
--input: Text string (mutually exclusive with --input_file)--input_file: Path to UTF-8 text file--language: latin or greek--models_dir: Model files directory (default:
$HOME/.predictingthepast/models)--attribute: Run geographical and chronological attribution--restore: Run text restoration (requires ? or # in input)--contextualize: Run parallel inscription retrieval--contextualize_top_k: Number of top parallels to return (default: 10)--contextualize_exclude_test_valid: Exclude test/validation texts from
contextualization results. Filters on the internal numeric id (where id % 10 is 3 or 4).--embedding: Generate a text embedding vector (default: off). When
set, the output includes an embedding field with the full-precision float
vector.--output_json: Save combined JSON to file (for visualize_results.py)--restore_beam_width: Number of candidate restorations to explore in
beam search. Higher = more thorough but slower (default: 100)--restore_max_len: Maximum total characters to restore across all ?
and # gaps (default: 15)--restore_temperature: Sampling temperature. Lower = conservative,
higher = creative (default: 1.0)--window_overlap: Overlap fraction for long-text windowing (default:
0.33)Each window must be between 25 and 750 characters. If the input text exceeds 750 characters, it is automatically split into overlapping windows (default 33% overlap). Each window is run through the model independently, and the geographical and chronological attribution results are averaged across all windows. Restoration and contextualization are run per-window and concatenated.
Use --output_json to save the combined JSON to a file. For the full
--output_json format including all fields for attribution, restoration, and
contextualization, see output_format.md.
Present results for every task requested by the user. Always generate the
HTML dashboard, serve it via the background HTTP server described above, and
provide the http:// URL as a clickable link. Format years as BCE/CE.
Bold the restored characters in the top prediction. Show top-10 beam-search candidates in a numbered table (rank, restored text, score). Summarise the most salient words (not characters) — do not dump the raw saliency array.
Show the top 10 regions in a ranked table (rank, region, score). Highlight the top-1 prediction in the narrative. Summarise the most salient words (not characters) — do not dump the raw saliency array.
Report top year and weighted average year. Describe the distribution shape (peak decade, probable range) — do not list all 160 bins. Summarise the most salient words (not characters) — do not dump the raw saliency array.
Use jp or write your own code to extract the contextualization data from
--output_json. Show the top parallels in a table (rank, ID, Trismegistos ID,
region, date range, score, partner link). Quote the full text of the top
parallels from the JSON.
If the user requests a specific number of parallels, use --contextualize_top_k
to control how many are returned.
If the user asks to exclude test/validation texts, add
--contextualize_exclude_test_valid. This filters on the internal numeric
id (where id % 10 is 3 for test or 4 for validation).
--embedding is set)State the vector dimensionality (384 floats).
For more information, see: predictingthepast.com
IMPORTANT: The first time you return inference results to the user, remind them to cite the appropriate paper below (Ithaca for Greek, Aeneas for Latin), and the dataset acknowledgements below.
Ithaca (Ancient Greek): > Assael, Y., Sommerschield, T., Shillingford, B. et al. Restoring and attributing ancient texts using deep neural networks. Nature 603, 280–283 (2022).
Aeneas (Latin): > Assael, Y., Sommerschield, T., Cooley, A. et al. Contextualizing ancient texts with generative neural networks. Nature 645, 141–147 (2025).
IMPORTANT: The first time you present inference results to the user, include the appropriate citation and the full dataset acknowledgement for the language used (see below).
The Ithaca and Aeneas projects rely on the availability of high-quality datasets of ancient Greek and Latin inscriptions, respectively. These datasets were built through centuries of scholarly collection and decades of digital editorial work.
Ithaca (Ancient Greek): In particular, Ithaca draws on a 2019 version of the Searchable Greek Inscriptions database made available under "Fair Use" license by the Packard Humanities Institute, generously supported by David Packard: inscriptions.packhum.org.
Aeneas (Latin): Aeneas was trained on data from:
These resources bring together a large proportion of published inscriptions in a searchable digital format.
© google-deepmind, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 9 other files (scripts, references) in skills/predictingthepast of google-deepmind/science-skills.
Open the folder on GitHubat commit 6883275
Predictingthepast next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Predictingthepast this skillgoogle-deepmind/science-skills | 3.2k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | |
| Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~2.3k | Automated safety check: Pass | MIT | |
| CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~1.7k | Automated safety check: Pass | MIT | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Codebase Managementgiancarloerra/SocratiCode | 3.3k | 1 repos | ~1.8k | Automated safety check: Pass | AGPL-3.0 | |
| Sentence-Transformers Training Routerhuggingface/skills | 11k | 1 repos | ~2.6k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
giancarloerra/SocratiCode
Set up, index, and manage SocratiCode codebase indexing. An agent skill from giancarloerra/SocratiCode.
huggingface/skills
Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.
JimLiu/science-skills
Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.
google-deepmind/science-skills
Retrieve and analyze AlphaFold predicted structures for a protein.
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
google-deepmind/science-skills
Query the ChEMBL database for bioactive molecules, drug targets, bioactivity data, approved drugs, and chemical structures.
google-deepmind/science-skills
Query ClinicalTrials.gov via APIv2. An agent skill from google-deepmind/science-skills.
google-deepmind/science-skills
A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
Categories
Ancient text restoration, attribution, dating, contextualization, and embedding via Aeneas (Latin) / Ithaca (Ancient Greek). Predictingthepast is an agent skill from google-deepmind/science-skills. Ancient text restoration, attribution, dating, contextualization, and embedding via Aeneas (Latin) / Ithaca (Ancient Greek).
Predictingthepast fits situations like: asked to restore; where was it written; analyze an ancient text; epigraphic document.
Run `npx skills add google-deepmind/science-skills --skill predictingthepast -a claude-code`. Or copy the skill folder (skills/predictingthepast in google-deepmind/science-skills) into .claude/skills/predictingthepast in your project. Claude Code loads it when a task matches its description.
Run `npx skills add google-deepmind/science-skills --skill predictingthepast -a codex`. Or copy the skill folder (skills/predictingthepast in google-deepmind/science-skills) into .agents/skills/predictingthepast in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google-deepmind/science-skills --skill predictingthepast -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/predictingthepast, .gemini/skills/predictingthepast, .github/skills/predictingthepast and .opencode/skills/predictingthepast in your project.
Going by SKILL.md and its folder, Predictingthepast needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3.
SKILL.md names 7 domains. As links in the text: zenodo.org, github.com, predictingthepast.com, inscriptions.packhum.org, edr-edr.it, edh.ub.uni-heidelberg.de and manfredclauss.de. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Predictingthepast is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 89k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Predictingthepast: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars) and Codebase Management (giancarloerra/SocratiCode, 3.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
google-deepmind (a GitHub organization) maintains it in google-deepmind/science-skills, which has 3,216 GitHub stars. The repository holds 40 skills in this directory. The repository was last updated on September 15, 2026.
Source: google-deepmind/science-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.