Agent skill

Google Image API Skill

by browser-act in browser-act/skills

This skill helps users automatically extract structured image data from Google Images via BrowserAct API.

MITAuto-check passedMarketing & SEO

Install Google Image API Skill

skills CLI
$ npx skills add browser-act/skills --skill google-image-api-skill -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install browser-act/skills google-image-api-skill --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/browser-act/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/solutions/search-research/google-image-api-skill .claude/skills/google-image-api-skill && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
google-image-api-skill
GitHub stars
6.1k
Used in
1 other repo
Token cost
~1.7k tokens
SKILL.md length
776 words
Files
2 (incl. scripts)
Skills in repo
87
Repo updated
First seen
Licence
MIT

At a glance

This skill helps users automatically extract structured image data from Google Images via BrowserAct API.

  • Works in 5 steps: No hallucinations, ensuring stable and… → No CAPTCHA issues: No need to deal with… → No IP restrictions or geo-blocking: No… → …
  • Tasks that involve Market research
  • SKILL.md covers 📖 Introduction, ✨ Features, 🔑 API Key Guide and 🛠️ Input Parameters, plus 4 more sections
  • Runs Python scripts from its folder; calls python; needs BROWSERACT_API_KEY

What it does

Google Image API Skill is an agent skill from browser-act/skills. This skill helps users automatically extract structured image data from Google Images via BrowserAct API. Agent should proactively apply this skill when users express needs like finding images for specific keywords, gathering product style images for competitors, building visual datasets at scale, scanning visual search results for market research, tracking localized image trends by country, compiling related image thumbnails and links, extracting image titles and source logos, fetching click through URLs from…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/google_image_api.py`).

It sits in Marketing & SEO, covering Market research and Logo and visual identity. The repository describes itself as: Browser automation CLI built for AI agents. Break through anti-bot walls, hand off to humans across platforms when stuck. Parallel multi-task execution, independent multi-session… The licence is MIT.

When your agent uses it

  • Tasks that involve Market research
  • Tasks that involve Logo and visual identity

Example prompts

  • “/google-image-api-skill”

Requirements

  • Python 3
  • A credential in BROWSERACT_API_KEY

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. No hallucinations, ensuring stable and accurate data extraction: Pre-set workflows avoid generative AI hallucinations.
  2. No CAPTCHA issues: No need to deal with reCAPTCHA or other verification challenges.
  3. No IP restrictions or geo-blocking: No need to handle regional IP limitations.
  4. Agile execution speed: Faster task execution compared to pure AI-driven browser automation solutions.
  5. High cost-effectiveness: Significantly reduces data acquisition costs compared to AI solutions that consume a large number of tokens.

What it can do on your machine

Read from SKILL.md and the folder at commit 11c057b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • browseract.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • BROWSERACT_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Google Image API Skill loads about 1.7k tokens when it runs. Until then it costs about 189 tokens; SKILL.md has 776 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~189
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from browser-act/skills at commit 11c057b, republished under its MIT licence (© browser-act). 776 words, ~1,665 tokens.

Download SKILL.mdSave it as .claude/skills/google-image-api-skill/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
google-image-api-skill
description
This skill helps users automatically extract structured image data from Google Images via BrowserAct API. Agent should proactively apply this skill when users express needs like finding images for specific keywords, gathering product style images for competitors, building visual datasets at scale, scanning visual search results for market research, tracking localized image trends by country, compiling related image thumbnails and links, extracting image titles and source logos, fetching click through URLs from image results, monitoring competitor visual assets, sourcing creative content for specific topics, looking up product pictures in different regions, collecting structured image metadata without opening detail pages.

Google Image API Automation Skill

📖 Introduction

This skill provides users with one-click image data extraction directly from Google Images using the BrowserAct Google Image API template. It allows you to search with keywords, set country and language, control scroll depth and result limits, returning clean, structured image metadata directly via API.

✨ Features

  1. No hallucinations, ensuring stable and accurate data extraction: Pre-set workflows avoid generative AI hallucinations.
  2. No CAPTCHA issues: No need to deal with reCAPTCHA or other verification challenges.
  3. No IP restrictions or geo-blocking: No need to handle regional IP limitations.
  4. Agile execution speed: Faster task execution compared to pure AI-driven browser automation solutions.
  5. High cost-effectiveness: Significantly reduces data acquisition costs compared to AI solutions that consume a large number of tokens.

🔑 API Key Guide

Before running, you must check the BROWSERACT_API_KEY environment variable. If it is not set, do not take any further action; you should request and wait for the user to provide it collaboratively. The Agent must inform the user at this point:

"Since you haven't configured the BrowserAct API Key yet, please go to the BrowserAct Console to get your Key first."

🛠️ Input Parameters

The Agent should flexibly configure the following parameters according to user needs when calling the script:

  1. KeyWords (Search keywords)

    • Type: string
    • Description: Search keywords used on Google Images.
    • Example: flower, ai agent, tesla
  2. Country (Country or region bias)

    • Type: string
    • Description: Country or region bias for results.
    • Supported values: us, gb, ca, au, de, fr, es, jp, kr
    • Default: us
  3. Language (UI language)

    • Type: string
    • Description: UI language for the Google Images session and returned text.
    • Supported values: en, zh-CN, zh-TW, ja, ko, fr, de, es
    • Default: en
  4. Scroll_count (Number of scroll actions)

    • Type: number
    • Description: Number of scroll actions to load more image results.
    • Default: 5
  5. Datelimit (Maximum items)

    • Type: number
    • Description: Maximum number of items to extract from the results list.
    • Default: 50

The Agent should execute the following independent script to achieve "results with one command":

bash
# Example invocation
python -u ./scripts/google_image_api.py "KeyWords" "Country" "Language" Scroll_count Datelimit
⏳ Execution Status Monitoring

Since this task involves automated browser operations, it may take a considerable amount of time (several minutes). The script will continuously output status logs with timestamps while running (e.g., [14:30:05] Task Status: running). Agent Notice:

  • While waiting for the script to return results, please keep an eye on the terminal output.
  • As long as the terminal is outputting new status logs, it means the task is running normally; do not mistake it for a deadlock or unresponsiveness.
  • If the status remains unchanged for a long time or the script stops outputting without returning a result, then consider triggering the retry mechanism.
Show full SKILL.md (333 more words)Show less

📊 Data Output

After successful execution, the script will parse and print the results directly from the API response. The results include:

  • is_product: Whether the result is detected as a product-style listing
  • link: Click-through URL associated with the result
  • title: Image result title or caption text
  • source_logo: Source site logo URL
  • source: Source site name shown in results
  • related_content_id: Google Images related content identifier
  • thumbnail: Thumbnail image URL
  • index: Result index in the list

⚠️ Error Handling & Retry

During the execution of the script, if an error occurs (such as network fluctuation or task failure), the Agent should follow this logic:

  1. Check the output:

    • If the output contains "Invalid authorization", it means the API Key is invalid or expired. In this case, do not retry; guide the user to check and provide the correct API Key.
    • If the output does not contain "Invalid authorization" but the task execution fails (for example, the output starts with Error: or the result is empty), the Agent should automatically try executing the script one more time.
  2. Retry limit:

    • Automatic retry is limited to once. If the second attempt still fails, stop retrying and report the specific error message to the user.

🌟 Typical Use Cases

  1. Visual Content Sourcing: Finding specific imagery for creative research and design content.
  2. Competitor Asset Monitoring: Scanning Google Images for competitor product styles and logos.
  3. Market Visual Research: Building datasets of product listings across various countries.
  4. Localized Image Trends: Tracking what images appear for specific terms in Japan (jp) or France (fr).
  5. E-commerce Discovery: Extracting click-through links to track down where products are sold.
  6. Data Enrichment: Fetching thumbnails and high-level titles associated with keywords.
  7. Brand Tracking: Finding instances of specific brands appearing as image results.
  8. SEO Keyword Visualization: Checking the visual results that rank for chosen SEO keywords.
  9. Automated Content Aggregation: Delivering daily list-level visual metadata for specific topics.
  10. Global Image Search: Finding images related to global events or personalities in their native languages.

© browser-act, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in solutions/search-research/google-image-api-skill of browser-act/skills.

  • SKILL.md
  • scripts/google_image_api.py

Open the folder on GitHubat commit 11c057b

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in browser-act/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Google Image API Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Google Image API Skill compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Google Image API Skill this skillbrowser-act/skills6.1k1 repos~1.7kAutomated safety check: PassMIT
Creative Directorsmixs/creative-director-skill246—~5.1kAutomated safety check: PassCC-BY-4.0
SEO Setupalisamadiii/Portfolio180—~1.7kAutomated safety check: PassNone
Scenario Brand Kitscenario-labs/skills898—~3.4kAutomated safety check: PassMIT
Od Brand Guidelinescriptogus/agent-evolve-network288—~831Automated safety check: PassApache-2.0
Magic Hour Brand Kithashgraph-online/awesome-codex-plugins1.2k—~1.1kAutomated safety check: PassMIT

Similar skills

  • Creative Director

    smixs/creative-director-skill

    AI creative director with recursive self-assessment. An agent skill from smixs/creative-director-skill.

    246 GitHub stars~5.1k tokensUpdated 2 mo ago
    Marketing & SEOAuto-check passed
  • SEO Setup

    alisamadiii/Portfolio

    Full SEO/metadata setup and audit for client websites (Astro, Next.js, any static site).

    180 GitHub stars~1.7k tokensUpdated yesterday
    Marketing & SEOAuto-check passed
  • Scenario Brand Kit

    scenario-labs/skills

    A skill your agent uses when building a visual identity or brand kit with Scenario: a logo, wordmark, or icon as editable SVG, a color palette and typography spec, logo variants (mono, reversed…

    898 GitHub stars~3.4k tokensUpdated yesterday
    Marketing & SEOAuto-check passed
  • Od Brand Guidelines

    criptogus/agent-evolve-network

    Apply Anthropic's official brand colors and typography to artifacts for consistent visual identity and professional design standards Use when the user asks for brand guidelines work, or mentions od…

    288 GitHub stars~831 tokensUpdated 28 days ago
    Marketing & SEOAuto-check passed
  • Magic Hour Brand Kit

    hashgraph-online/awesome-codex-plugins

    Create, extend or apply a reusable brand kit with Magic Hour imagery, exact logos and colors, editable layouts and a practical brand guide.

    1.2k GitHub stars~1.1k tokensUpdated yesterday
    Marketing & SEOAuto-check passed
  • Customer Research

    Nexus-JPF/note-companion

    When the user wants to conduct, analyze, or synthesize customer research.

    869 GitHub starsUsed in 5 repos~3.2k tokens
    Marketing & SEOAuto-check passed

More from browser-act/skills

All 87 skills in this repo
  • Amazon ASIN Lookup

    browser-act/skills

    Fetches structured Amazon product details such as title, price, ratings and availability for a given ASIN through BrowserAct's lookup API template.

    6.1k GitHub starsUsed in 2 repos~1.5k tokens
    Auto-check passed
  • Amazon Best Sellers Finder

    browser-act/skills

    Extracts structured Amazon product data for a keyword and marketplace through the BrowserAct API, including titles, prices, ratings, reviews, sales volume and promotions.

    6.1k GitHub starsUsed in 2 repos~1.5k tokens
    Auto-check passed
  • Amazon Buy Box Monitor

    browser-act/skills

    Pulls Amazon product details, competing seller prices and seller ratings for a given ASIN through the BrowserAct API, without browser automation.

    6.1k GitHub starsUsed in 2 repos~1.6k tokens
    Auto-check passed
  • Analyzes a competitor's Amazon listing by ASIN with BrowserAct data extraction, then reports what it does well, where the market has gaps and opportunity points for your own listing.

    6.1k GitHub starsUsed in 2 repos~3.2k tokens
    Auto-check passed
  • Pulls structured Amazon search results (titles, ASINs, prices, ratings, specifications) for a keyword and brand through BrowserAct's Amazon Product API template.

    6.1k GitHub starsUsed in 2 repos~1.5k tokens
    Auto-check passed
  • Collects structured product data from Amazon search results for a keyword and optional brand, using a BrowserAct script, for market and competitor research.

    6.1k GitHub starsUsed in 2 repos~1.6k tokens
    Auto-check passed

Questions about Google Image API Skill

What does Google Image API Skill do?

This skill helps users automatically extract structured image data from Google Images via BrowserAct API. Google Image API Skill is an agent skill from browser-act/skills. This skill helps users automatically extract structured image data from Google Images via BrowserAct API.

When should I use Google Image API Skill?

Google Image API Skill fits situations like: tasks that involve Market research; tasks that involve Logo and visual identity.

How do I install Google Image API Skill in Claude Code?

Run `npx skills add browser-act/skills --skill google-image-api-skill -a claude-code`. Or copy the skill folder (solutions/search-research/google-image-api-skill in browser-act/skills) into .claude/skills/google-image-api-skill in your project. Claude Code loads it when a task matches its description.

How do I install Google Image API Skill in Codex?

Run `npx skills add browser-act/skills --skill google-image-api-skill -a codex`. Or copy the skill folder (solutions/search-research/google-image-api-skill in browser-act/skills) into .agents/skills/google-image-api-skill in your project. Codex loads it when a task matches its description.

Can I use Google Image API Skill in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add browser-act/skills --skill google-image-api-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/google-image-api-skill, .gemini/skills/google-image-api-skill, .github/skills/google-image-api-skill and .opencode/skills/google-image-api-skill in your project.

What does Google Image API Skill need to run?

Going by SKILL.md and its folder, Google Image API Skill needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named BROWSERACT_API_KEY. Our summary lists: Python 3; A credential in BROWSERACT_API_KEY.

Does Google Image API Skill access the network?

SKILL.md names 1 domain. As links in the text: browseract.com. This is read from the text; nothing was executed.

Is Google Image API Skill safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Google Image API Skill use?

Google Image API Skill is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Google Image API Skill use?

About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Google Image API Skill?

Skills that share tags, products or a category with Google Image API Skill: Creative Director (smixs/creative-director-skill, 246 stars), SEO Setup (alisamadiii/Portfolio, 180 stars), Scenario Brand Kit (scenario-labs/skills, 898 stars) and Od Brand Guidelines (criptogus/agent-evolve-network, 288 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Google Image API Skill?

browser-act (a GitHub organization) maintains it in browser-act/skills, which has 6,108 GitHub stars. The repository holds 87 skills in this directory. The repository was last updated on August 24, 2026.

Source: browser-act/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.