Agent skill

Gemini Count In Video

by benchflow-ai in benchflow-ai/skillsbench

Analyze and count objects in videos using Google Gemini API (object counting, pedestrian detection, vehicle tracking, and surveillance video analysis).

Apache-2.0Auto-check passed

Install Gemini Count In Video

skills CLI
$ npx skills add benchflow-ai/skillsbench --skill gemini-count-in-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/skillsbench gemini-count-in-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tasks-extra/pedestrian-traffic-counting/environment/skills/gemini-count-in-video .claude/skills/gemini-count-in-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gemini-count-in-video
GitHub stars
1.8k
Token cost
~2.7k tokens
SKILL.md length
484 words
Files
1
Skills in repo
189
Repo updated
First seen
Licence
Apache-2.0

At a glance

Analyze and count objects in videos using Google Gemini API (object counting, pedestrian detection, vehicle tracking, and surveillance video analysis).

  • SKILL.md covers Purpose, When to Use, Required Libraries and Input Requirements, plus 7 more sections
  • Needs GEMINI_API_KEY

What it does

Gemini Count In Video is an agent skill from benchflow-ai/skillsbench. Analyze and count objects in videos using Google Gemini API (object counting, pedestrian detection, vehicle tracking, and surveillance video analysis).

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It works with Google Gemini. The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is Apache-2.0.

Example prompts

  • “/gemini-count-in-video”

Requirements

  • Python 3
  • A credential in GEMINI_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • ai.google.dev
    • aistudio.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gemini Count In Video loads about 2.7k tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 484 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~43
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its Apache-2.0 licence (© benchflow-ai). 484 words, ~2,698 tokens.

Download SKILL.mdSave it as .claude/skills/gemini-count-in-video/SKILL.md (or your agent's skills folder).
name
gemini-count-in-video
description
Analyze and count objects in videos using Google Gemini API (object counting, pedestrian detection, vehicle tracking, and surveillance video analysis).

Gemini Video Understanding Skill

Purpose

This skill enables video analysis and object counting using the Google Gemini API, with a focus on counting pedestrians, detecting objects, tracking movement, and analyzing surveillance footage. It supports precise prompting for differentiated counting (e.g., pedestrians vs cyclists vs vehicles).

When to Use

  • Counting pedestrians, vehicles, or other objects in surveillance videos
  • Distinguishing between different types of objects (walkers vs cyclists, cars vs trucks)
  • Analyzing traffic patterns and movement through a scene
  • Processing multiple videos for batch object counting
  • Extracting structured count data from video footage

Required Libraries

The following Python libraries are required:

python
from google import genai
from google.genai import types
import os
import time

Input Requirements

  • File formats: MP4, MPEG, MOV, AVI, FLV, MPG, WebM, WMV, 3GPP
  • Size constraints:
    • Use inline bytes for small files (rule of thumb: <20MB).
    • Use the File API upload flow for larger videos (most surveillance footage).
    • Always wait for processing to complete before analysis.
  • Video quality: Higher resolution provides better counting accuracy for distant objects
  • Duration: Longer videos may require longer processing times; consider the full video length for accurate counting

Output Schema

For object counting tasks, structure results as JSON:

json
{
  "success": true,
  "video_file": "surveillance_001.mp4",
  "model": "gemini-2.0-flash-exp",
  "counts": {
    "pedestrians": 12,
    "cyclists": 3,
    "vehicles": 5
  },
  "notes": "Optional observations about the counting process or edge cases"
}
Field Descriptions
  • success: Whether the analysis completed successfully
  • video_file: Name of the analyzed video file
  • model: Gemini model used for the request
  • counts: Object counts by category
  • notes: Any clarifications or warnings about the count

Code Examples

Basic Pedestrian Counting (File API Upload)
python
from google import genai
import os
import time
import re

client = genai.Client(api_key=os.getenv("GEMINI_API_KEY"))

# Upload video (File API for >20MB)
myfile = client.files.upload(file="surveillance.mp4")

# Wait for processing
while myfile.state.name == "PROCESSING":
    time.sleep(5)
    myfile = client.files.get(name=myfile.name)

if myfile.state.name == "FAILED":
    raise ValueError("Video processing failed")

# Prompt for counting pedestrians with clear exclusion criteria
prompt = """Count the total number of pedestrians who are WALKING through the scene in this surveillance video.

IMPORTANT RULES:
- ONLY count people who are walking on foot
- DO NOT count people riding bicycles
- DO NOT count people driving cars or other vehicles
- Count each unique pedestrian only once, even if they appear in multiple frames

Provide your answer as a single integer number representing the total count of pedestrians.
Answer with just the number, nothing else.
Your answer should be enclosed in <answer> and </answer> tags, such as <answer>5</answer>.
"""

response = client.models.generate_content(
    model="gemini-2.0-flash-exp",
    contents=[prompt, myfile],
)

# Parse the response
response_text = response.text.strip()
match = re.search(r"<answer>(\d+)</answer>", response_text)
if match:
    count = int(match.group(1))
    print(f"Pedestrian count: {count}")
else:
    print("Could not parse count from response")
Batch Processing Multiple Videos
python
from google import genai
import os
import time
import re

def upload_and_wait(client, file_path: str, max_wait_s: int = 300):
    """Upload video and wait for processing."""
    myfile = client.files.upload(file=file_path)
    waited = 0
    
    while myfile.state.name == "PROCESSING" and waited < max_wait_s:
        time.sleep(5)
        waited += 5
        myfile = client.files.get(name=myfile.name)
    
    if myfile.state.name == "FAILED":
        raise ValueError(f"Video processing failed: {myfile.state.name}")
    if myfile.state.name == "PROCESSING":
        raise TimeoutError(f"Processing timeout after {max_wait_s}s")
    
    return myfile

client = genai.Client(api_key=os.getenv("GEMINI_API_KEY"))

# Process all videos in directory
video_dir = "/app/video"
video_extensions = {".mp4", ".mkv", ".avi", ".mov"}
results = {}

for filename in os.listdir(video_dir):
    if any(filename.lower().endswith(ext) for ext in video_extensions):
        video_path = os.path.join(video_dir, filename)
        
        print(f"Processing {filename}...")
        
        # Upload and analyze
        myfile = upload_and_wait(client, video_path)
        
        response = client.models.generate_content(
            model="gemini-2.0-flash-exp",
            contents=["Count pedestrians walking through the scene. Answer with just the number.", myfile],
        )
        
        # Extract count
        count = int(re.search(r'\d+', response.text).group())
        results[filename] = count
        print(f"  Count: {count}")

print(f"\nProcessed {len(results)} videos")
# Results dictionary can now be used for further processing or saving
Differentiating Object Types
python
# Count different categories separately
prompt = """Analyze this surveillance video and count:
1. Pedestrians (people walking on foot)
2. Cyclists (people riding bicycles)
3. Vehicles (cars, trucks, motorcycles)

RULES:
- Count each unique individual/vehicle only once
- If someone switches from walking to cycling, count them in their primary mode
- Provide counts as three separate numbers

Format your answer as:
Pedestrians: <number>
Cyclists: <number>
Vehicles: <number>
"""

response = client.models.generate_content(
    model="gemini-2.0-flash-exp",
    contents=[prompt, myfile],
)

# Parse multiple counts
text = response.text
pedestrians = int(re.search(r'Pedestrians:\s*(\d+)', text).group(1))
cyclists = int(re.search(r'Cyclists:\s*(\d+)', text).group(1))
vehicles = int(re.search(r'Vehicles:\s*(\d+)', text).group(1))
Using Answer Tags for Reliable Parsing
python
# Request structured output with XML-like tags
prompt = """Count the total number of pedestrians walking through the scene.

You should reason and think step by step. Provide your answer as a single integer.
Your answer should be enclosed in <answer> and </answer> tags, such as <answer>5</answer>.
"""

response = client.models.generate_content(
    model="gemini-2.0-flash-exp",
    contents=[prompt, myfile],
)

# Robust extraction
match = re.search(r"<answer>(\d+)</answer>", response.text)
if match:
    count = int(match.group(1))
else:
    # Fallback: try to find any number in response
    numbers = re.findall(r'\d+', response.text)
    count = int(numbers[0]) if numbers else 0
Show full SKILL.md (247 more words)Show less

Best Practices

  • Use the File API for all surveillance videos (typically >20MB) and always wait for processing to complete.
  • Be specific in prompts: Clearly define what to count and what to exclude (e.g., "walking pedestrians only, not cyclists").
  • Use structured output formats: Request answers in specific formats (like <answer>N</answer>) for reliable parsing.
  • Ask for reasoning: Include "think step by step" to improve counting accuracy.
  • Handle edge cases: Specify rules for partial appearances, people entering/exiting frame, and mode changes.
  • Use gemini-2.0-flash-exp or gemini-2.5-flash: These models provide good balance of speed and accuracy for object counting.
  • Test with sample videos: Verify prompt effectiveness on representative samples before batch processing.

Error Handling

python
import time

def upload_and_wait(client, file_path: str, max_wait_s: int = 300):
    """Upload video and wait for processing with timeout."""
    myfile = client.files.upload(file=file_path)
    waited = 0

    while myfile.state.name == "PROCESSING" and waited < max_wait_s:
        time.sleep(5)
        waited += 5
        myfile = client.files.get(name=myfile.name)

    if myfile.state.name == "FAILED":
        raise ValueError(f"Video processing failed: {myfile.state.name}")
    if myfile.state.name == "PROCESSING":
        raise TimeoutError(f"Processing timeout after {max_wait_s}s")

    return myfile

def count_with_fallback(client, video_path):
    """Count pedestrians with error handling and fallback."""
    try:
        myfile = upload_and_wait(client, video_path)
        
        prompt = """Count pedestrians walking through the scene.
        Answer with just the number in <answer></answer> tags."""
        
        response = client.models.generate_content(
            model="gemini-2.0-flash-exp",
            contents=[prompt, myfile],
        )
        
        # Try structured parsing first
        match = re.search(r"<answer>(\d+)</answer>", response.text)
        if match:
            return int(match.group(1))
        
        # Fallback to any number found
        numbers = re.findall(r'\d+', response.text)
        if numbers:
            return int(numbers[0])
        
        print(f"Warning: Could not parse count, defaulting to 0")
        return 0
        
    except Exception as e:
        print(f"Error processing video: {e}")
        return 0

Common issues:

  • Upload processing stuck: Use timeout logic and fail gracefully after max wait time
  • Ambiguous responses: Use structured output tags like <answer></answer> for reliable parsing
  • Rate limits: Add retry logic with exponential backoff for batch processing
  • Inconsistent counts: Be very explicit in prompts about counting rules and exclusions

Limitations

  • Counting accuracy depends on video quality, camera angle, and object size/distance
  • Very crowded scenes may have higher counting variance
  • Occlusion (objects blocking each other) can affect accuracy
  • Long videos require longer processing times (typically 5-30 seconds per video)
  • The model may occasionally misclassify similar objects (e.g., motorcyclist as cyclist)
  • For highest accuracy, use clear prompts with explicit inclusion/exclusion criteria

Version History

  • 1.0.0 (2026-01-21): Tailored for pedestrian traffic counting with focus on object counting, differentiation, and batch processing

Resources

© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in tasks-extra/pedestrian-traffic-counting/environment/skills/gemini-count-in-video of benchflow-ai/skillsbench.

Open the folder on GitHubat commit 9a1f4dd

Compare with similar skills

Gemini Count In Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gemini Count In Video compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gemini Count In Video this skillbenchflow-ai/skillsbench1.8k—~2.7kAutomated safety check: PassApache-2.0
NotebookLM Research AssistantPleasePrompto/notebooklm-skill7.8k14 repos~2.4kAutomated safety check: NotesMIT
Brand and Design Toolkitnextlevelbuilder/ui-ux-pro-max-skill135k1 repos~3.5kAutomated safety check: PassMIT
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
OpenCLI Smart Search Routerjackwener/OpenCLI30k2 repos~753Automated safety check: PassApache-2.0
NotebookLM Automationteng-lin/notebooklm-py20k—~4.1kAutomated safety check: PassMIT

Similar skills

  • NotebookLM Research Assistant

    PleasePrompto/notebooklm-skill

    Lets Claude Code ask questions of your Google NotebookLM notebooks through browser automation and return answers grounded in your uploaded sources.

    7.8k GitHub starsUsed in 14 repos~2.4k tokens
    Knowledge ManagementAuto-check: notes
  • Brand and Design Toolkit

    nextlevelbuilder/ui-ux-pro-max-skill

    Bundles design tasks behind one skill: brand identity, tokens, UI styling, logos, corporate identity mockups, slides, banners, icons and social images.

    135k GitHub starsUsed in 1 repo~3.5k tokens
    Media & CreativeAuto-check passed
  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Routes a search question to the most suitable opencli source, normally one AI site plus up to two specialist sites, with call limits per question and a recap at the end.

    30k GitHub starsUsed in 2 repos~753 tokens
    Productivity & AutomationAuto-check passed
  • NotebookLM Automation

    teng-lin/notebooklm-py

    Installs, authenticates and operates Gemini Notebook (NotebookLM) through the notebooklm-py CLI or its typed async Python API, for notebooks, sources, grounded chat and generated artifacts.

    20k GitHub stars~4.1k tokensUpdated yesterday
    Knowledge ManagementAuto-check passed
  • Nano Banana Pro Prompts Recommend Skill

    YouMind-OpenLab/nano-banana-pro-prompts-recommend-skill

    Recommend suitable prompts from 10,000+ Nano Banana Pro image generation prompts based on user needs.

    1.9k GitHub starsUsed in 1 repo~4.1k tokens
    Media & CreativeAuto-check passed

More from benchflow-ai/skillsbench

All 189 skills in this repo
  • Lean4 Memories

    benchflow-ai/skillsbench

    This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…

    1.8k GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Ac Branch Pi Model

    benchflow-ai/skillsbench

    AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.

    1.8k GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Civ6lib

    benchflow-ai/skillsbench

    Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • D3 Visualization

    benchflow-ai/skillsbench

    Build deterministic, verifiable data visualizations with D3.js (v6).

    1.8k GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Dc Power Flow

    benchflow-ai/skillsbench

    DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~717 tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about Gemini Count In Video

What does Gemini Count In Video do?

Analyze and count objects in videos using Google Gemini API (object counting, pedestrian detection, vehicle tracking, and surveillance video analysis). Gemini Count In Video is an agent skill from benchflow-ai/skillsbench. Analyze and count objects in videos using Google Gemini API (object counting, pedestrian detection, vehicle tracking, and surveillance video analysis).

How do I install Gemini Count In Video in Claude Code?

Run `npx skills add benchflow-ai/skillsbench --skill gemini-count-in-video -a claude-code`. Or copy the skill folder (tasks-extra/pedestrian-traffic-counting/environment/skills/gemini-count-in-video in benchflow-ai/skillsbench) into .claude/skills/gemini-count-in-video in your project. Claude Code loads it when a task matches its description.

How do I install Gemini Count In Video in Codex?

Run `npx skills add benchflow-ai/skillsbench --skill gemini-count-in-video -a codex`. Or copy the skill folder (tasks-extra/pedestrian-traffic-counting/environment/skills/gemini-count-in-video in benchflow-ai/skillsbench) into .agents/skills/gemini-count-in-video in your project. Codex loads it when a task matches its description.

Can I use Gemini Count In Video in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill gemini-count-in-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gemini-count-in-video, .gemini/skills/gemini-count-in-video, .github/skills/gemini-count-in-video and .opencode/skills/gemini-count-in-video in your project.

What does Gemini Count In Video need to run?

Going by SKILL.md and its folder, Gemini Count In Video needs credentials named GEMINI_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY.

Does Gemini Count In Video access the network?

SKILL.md names 2 domains. As links in the text: ai.google.dev and aistudio.google.com. This is read from the text; nothing was executed.

Is Gemini Count In Video safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gemini Count In Video use?

Gemini Count In Video is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gemini Count In Video use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gemini Count In Video?

Skills that share tags, products or a category with Gemini Count In Video: NotebookLM Research Assistant (PleasePrompto/notebooklm-skill, 7.8k stars), Brand and Design Toolkit (nextlevelbuilder/ui-ux-pro-max-skill, 135k stars), AI Image Generation and Editing (zhayujie/CowAgent, 47k stars) and OpenCLI Smart Search Router (jackwener/OpenCLI, 30k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gemini Count In Video?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,835 GitHub stars. The repository holds 189 skills in this directory. The repository was last updated on July 23, 2026.

Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.