This skill enables AI video generation from images AND text-to-speech voiceover generation using Fal.ai's API.

MITAuto-check: notesMedia & Creative

Install Fal AI

skills CLI
$ npx skills add mikeOnBreeze/cc-crossbeam --skill fal-ai -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mikeOnBreeze/cc-crossbeam fal-ai --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mikeOnBreeze/cc-crossbeam.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/fal-ai .claude/skills/fal-ai && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fal-ai
GitHub stars
293
Token cost
~1.9k tokens
SKILL.md length
566 words
Files
7 (incl. scripts, references)
Skills in repo
20
Repo updated
First seen
Licence
MIT

At a glance

This skill enables AI video generation from images AND text-to-speech voiceover generation using Fal.ai's API.

  • Works in 3 steps: Iterate with cheap models (Wan/Kling) -… → Test with Veo Fast - Verify quality… → Final render with Veo Standard - Premium…
  • The user asks to
  • SKILL.md covers Prerequisites, Primary Capability:…, Model Selection Strategy and Quick Model Reference (January…, plus 5 more sections
  • Runs Python scripts from its folder; calls python and pip; needs FAL_API_KEY

What it does

Fal AI is an agent skill from mikeOnBreeze/cc-crossbeam. This skill enables AI video generation from images AND text-to-speech voiceover generation using Fal.ai's API. Use this skill when the user asks to (1) generate videos from images (image-to-video), or (2) generate voiceovers/narration from text (text-to-speech via ElevenLabs). Works seamlessly with the nano-banana skill for image-to-video workflows. IMPORTANT: Check references/ for latest models and pricing - AI models change frequently.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `env.md`, `references/api_reference.md` and `references/tts_reference.md`).

It sits in Media & Creative, covering Text to speech and voice and AI video generation. It works with fal, Google Gemini and ElevenLabs. The repository describes itself as: CrossBeam Permits — AI-assisted building-permit plan review for cities and builders. The licence is MIT.

When your agent uses it

  • The user asks to
  • Generate videos from images (image-to-video)
  • Generate voiceovers/narration from text (text-to-speech via ElevenLabs)

Example prompts

  • “/fal-ai”

Requirements

  • Python 3
  • A credential in FAL_API_KEY

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Iterate with cheap models (Wan/Kling) - Perfect your prompts at $0.05-0.07/sec
  2. Test with Veo Fast - Verify quality improvement at $0.15/sec
  3. Final render with Veo Standard - Premium output at $0.40/sec

What it can do on your machine

Read from SKILL.md and the folder at commit cc5591e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • fal.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • FAL_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fal AI loads about 1.9k tokens when it runs, and up to ~5.8k if it reads all its reference files. Until then it costs about 112 tokens; SKILL.md has 566 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~112
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:16
    quires a `FAL_API_KEY` in the project's `.env` or `.env.local` file:

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from mikeOnBreeze/cc-crossbeam at commit cc5591e, republished under its MIT licence (© mikeOnBreeze). 566 words, ~1,896 tokens.

Download SKILL.mdSave it as .claude/skills/fal-ai/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
fal-ai
description
This skill enables AI video generation from images AND text-to-speech voiceover generation using Fal.ai's API. Use this skill when the user asks to (1) generate videos from images (image-to-video), or (2) generate voiceovers/narration from text (text-to-speech via ElevenLabs). Works seamlessly with the nano-banana skill for image-to-video workflows. IMPORTANT: Check references/ for latest models and pricing - AI models change frequently.

Fal.ai Media Generation

Generate AI videos from images and AI voiceovers from text using Fal.ai's API.

Capabilities:

  • Image-to-Video: Generate video clips from static images (see below)
  • Text-to-Speech: Generate voiceovers using ElevenLabs voices (see TTS section)

Prerequisites

This skill requires a FAL_API_KEY in the project's .env or .env.local file:

env
FAL_API_KEY=your_api_key_here

To obtain an API key, visit: https://fal.ai/dashboard/keys

Primary Capability: Image-to-Video

Generate video clips from static images using scripts/image_to_video.py:

bash
python scripts/image_to_video.py image.png \
  --prompt "camera slowly pans right, gentle motion"

Parameters:

  • image (required): Path to input image or URL
  • --prompt, -p: Motion/action description for the video
  • --model, -m: Video model to use (default: kling)
  • --output, -o: Output directory or file path (default: current directory)
  • --duration, -d: Video duration in seconds (model-dependent, typically 5-10)

Model Selection Strategy

Video generation is expensive. Follow this workflow:

  1. Iterate with cheap models (Wan/Kling) - Perfect your prompts at $0.05-0.07/sec
  2. Test with Veo Fast - Verify quality improvement at $0.15/sec
  3. Final render with Veo Standard - Premium output at $0.40/sec

Quick Model Reference (January 2026)

FlagModelPrice/sec5-sec CostBest For
wanWan 2.5$0.05$0.25Cheapest iteration
klingKling 2.5 Turbo Pro$0.07$0.35Best value (default)
veo-fastVeo 3.1 Fast$0.15$0.75Quality test
veoVeo 3.1 Standard$0.40$2.00Premium final

For detailed model comparison, strengths/weaknesses, and latest updates, see references/api_reference.md.

Usage Examples

Basic video generation (default Kling model):

bash
python scripts/image_to_video.py photo.png \
  --prompt "gentle breeze moves the leaves, soft lighting"

Budget iteration with Wan:

bash
python scripts/image_to_video.py photo.png \
  --prompt "camera zooms in slowly" \
  --model wan

Premium render with Veo:

bash
python scripts/image_to_video.py photo.png \
  --prompt "cinematic dolly shot, dramatic lighting" \
  --model veo \
  --output final_video.mp4

Specify duration:

bash
python scripts/image_to_video.py photo.png \
  --prompt "waves crash on the shore" \
  --duration 10

Workflow with Nano Banana

Generate an image, then create a video from it:

bash
# Step 1: Generate image with Nano Banana
python .claude/skills/nano-banana/scripts/generate_image.py \
  "a serene Japanese garden with cherry blossoms" \
  --output garden.png

# Step 2: Iterate with Kling (default, $0.35 for 5 sec)
python skills/fal-ai/scripts/image_to_video.py \
  garden.png \
  --prompt "gentle breeze moves cherry blossom petals, camera slowly pans right"

# Step 3: Final render with Veo when satisfied ($2.00 for 5 sec)
python skills/fal-ai/scripts/image_to_video.py \
  garden.png \
  --prompt "gentle breeze moves cherry blossom petals, camera slowly pans right" \
  --model veo

Text-to-Speech Voiceovers

Generate AI voiceovers using ElevenLabs voices via scripts/text_to_speech.py:

bash
python scripts/text_to_speech.py "Your text here" --voice george

Parameters:

  • text (required): Text to convert to speech
  • --voice, -v: Voice name (see casting guide below)
  • --model, -m: TTS model (default: eleven-v3)
  • --output, -o: Output directory or file path
  • --stability: Emotion control 0-1 (lower = more emotion)
  • --similarity: Voice matching 0-1
  • --style: Expression exaggeration 0-1
  • --speed: Speaking pace 0.7-1.2
  • --list-voices: Show all available voices
TTS Model Reference (January 2026)
FlagModelPrice/1K charsBest For
eleven-v3ElevenLabs Eleven v3$0.10Latest, audio tags [whispers] etc.
turboElevenLabs Turbo v2.5$0.05Fast iteration, low latency
multilingualMultilingual v2$0.10Best stability
Show full SKILL.md (231 more words)Show less
Voice Casting Quick Reference

When the user describes what they're looking for, match to these voices:

Female Voices:

VoiceBest For
rachelNarration, explainers, tutorials (calm, warm)
ariaConversational, podcasts (engaging, social)
sarahCorporate, professional (clear, neutral)
lauraMarketing, launches (upbeat, energetic)
charlottePremium brands (British, elegant)
lilyWellness, calm content (soft, gentle)

Male Voices:

VoiceBest For
georgeDocumentaries, serious narration (British, authoritative)
charlieCasual explainers (natural, relaxed)
rogerTrailers, announcements (deep, commanding)
ericNews-style, corporate (professional, clear)
chrisBrand voices, ads (warm, trustworthy)
brianEducational, history (mature, wise)

For full casting descriptions and parameter presets, see references/tts_reference.md.

TTS Usage Examples

Basic voiceover:

bash
python scripts/text_to_speech.py "Welcome to our product demo." --voice george

Voice casting (run multiple to compare):

bash
TEXT="Introducing the future of productivity."
python scripts/text_to_speech.py "$TEXT" --voice george -o casting_george.mp3
python scripts/text_to_speech.py "$TEXT" --voice eric -o casting_eric.mp3
python scripts/text_to_speech.py "$TEXT" --voice chris -o casting_chris.mp3

Documentary style (authoritative, slower):

bash
python scripts/text_to_speech.py "In the depths of the ocean..." \
  --voice george --stability 0.65 --speed 0.95

Conversational style (more emotion):

bash
python scripts/text_to_speech.py "Hey, check this out!" \
  --voice aria --stability 0.4 --style 0.3

With audio tags (eleven-v3 only):

bash
python scripts/text_to_speech.py "[whispers] This is a secret..." --voice rachel

Resources

scripts/
  • image_to_video.py - Image-to-video generation script
  • text_to_speech.py - Text-to-speech voiceover script
  • requirements.txt - Python dependencies (install with pip install -r requirements.txt)
references/
  • api_reference.md - Video model comparison, pricing, best practices
  • tts_reference.md - Voice casting guide, parameter presets, TTS best practices

Notes

  • Models evolve rapidly: Check reference docs dates. If >1 month old, research latest models on Fal.ai before generating
  • Video is expensive: Always be aware of costs. Iterate cheap, render expensive.
  • TTS is cheap: Run voice casting calls (~$0.02 for 3 samples) before committing to full narration
  • Queue-based API: Generation takes time. Scripts show progress updates.
  • Output formats: Videos = MP4, Audio = MP3
  • Duration limits: Video models typically support 5-10 seconds. Check api_reference.md.

© mikeOnBreeze, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in .claude/skills/fal-ai of mikeOnBreeze/cc-crossbeam.

  • SKILL.md
  • env.md
  • references/api_reference.md
  • references/tts_reference.md
  • scripts/image_to_video.py
  • scripts/requirements.txt
  • scripts/text_to_speech.py

Open the folder on GitHubat commit cc5591e

Compare with similar skills

Fal AI next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fal AI compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fal AI this skillmikeOnBreeze/cc-crossbeam293—~1.9kAutomated safety check: NotesMIT
Fal AI Mediaaffaan-m/ECC277k4 repos~1.9kAutomated safety check: PassMIT
Super Video MakerBomx/super-video-maker-skill310—~11kAutomated safety check: NotesNone
Hearyourvoicekillernay/HearYourVOICE140—~10kAutomated safety check: NotesMIT
AI Video Gencalesthio/OpenMontage66k—~3kAutomated safety check: PassAGPL-3.0
AI Video Gencalesthio/OpenMontage66k—~2.8kAutomated safety check: PassAGPL-3.0

Similar skills

  • Fal AI Media

    affaan-m/ECC

    Unified media generation via fal.ai MCP — image, video, and audio.

    277k GitHub starsUsed in 4 repos~1.9k tokens
    Media & CreativeAuto-check passed
  • Super Video Maker

    Bomx/super-video-maker-skill

    End-to-end AI video production skill for agentic frameworks.

    310 GitHub stars~11k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Hearyourvoice

    killernay/HearYourVOICE

    The repeatable workflow for short Thai documentary/explainer videos — one topic in, one finished MP4 out.

    140 GitHub stars~10k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • AI Video Gen

    calesthio/OpenMontage

    Generate AI videos from text prompts using multiple provider gateways.

    66k GitHub stars~3k tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • AI Video Gen

    calesthio/OpenMontage

    Generate AI videos from text prompts using multiple provider gateways.

    66k GitHub stars~2.8k tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • Suggest Sfx

    hassancs91/claude-youtube-editor

    Step 4 of the AI Video Editor pipeline — the SFX pass. An agent skill from hassancs91/claude-youtube-editor.

    328 GitHub stars~3k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

More from mikeOnBreeze/cc-crossbeam

All 20 skills in this repo
  • Adu PDF Extraction

    mikeOnBreeze/cc-crossbeam

    This skill extracts construction PDF plan binders into agent-consumable formats.

    293 GitHub stars~5k tokensUpdated 7 mo ago
    Auto-check passed
  • Adu Targeted Page Viewer

    mikeOnBreeze/cc-crossbeam

    Extracts construction plan PDFs into page PNGs, reads the sheet index to build a sheet-to-page manifest, and enables targeted viewing of specific sheets.

    293 GitHub stars~1.8k tokensUpdated 7 mo ago
    Auto-check passed
  • Nano Banana

    mikeOnBreeze/cc-crossbeam

    This skill enables image generation and editing using Google's Gemini Nano Banana models.

    293 GitHub stars~944 tokensUpdated 7 mo ago
    Auto-check: notes
  • Adu City Research

    mikeOnBreeze/cc-crossbeam

    Researches city-level ADU regulations, municipal codes, and standard details for any California city.

    293 GitHub stars~3.9k tokensUpdated 7 mo ago
    Auto-check passed
  • Adu Plan Review

    mikeOnBreeze/cc-crossbeam

    City-side ADU plan review — the flip side of adu-corrections-flow.

    293 GitHub stars~3.7k tokensUpdated 7 mo ago
    Auto-check passed
  • Crossbeam Ops

    mikeOnBreeze/cc-crossbeam

    Operations manual for the CrossBeam ADU Permit Assistant. An agent skill from mikeOnBreeze/cc-crossbeam.

    293 GitHub stars~534 tokensUpdated 7 mo ago
    Auto-check passed

Questions about Fal AI

What does Fal AI do?

This skill enables AI video generation from images AND text-to-speech voiceover generation using Fal.ai's API. Fal AI is an agent skill from mikeOnBreeze/cc-crossbeam.ai's API.

When should I use Fal AI?

Fal AI fits situations like: the user asks to; generate videos from images (image-to-video); generate voiceovers/narration from text (text-to-speech via ElevenLabs).

How do I install Fal AI in Claude Code?

Run `npx skills add mikeOnBreeze/cc-crossbeam --skill fal-ai -a claude-code`. Or copy the skill folder (.claude/skills/fal-ai in mikeOnBreeze/cc-crossbeam) into .claude/skills/fal-ai in your project. Claude Code loads it when a task matches its description.

How do I install Fal AI in Codex?

Run `npx skills add mikeOnBreeze/cc-crossbeam --skill fal-ai -a codex`. Or copy the skill folder (.claude/skills/fal-ai in mikeOnBreeze/cc-crossbeam) into .agents/skills/fal-ai in your project. Codex loads it when a task matches its description.

Can I use Fal AI in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mikeOnBreeze/cc-crossbeam --skill fal-ai -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fal-ai, .gemini/skills/fal-ai, .github/skills/fal-ai and .opencode/skills/fal-ai in your project.

What does Fal AI need to run?

Going by SKILL.md and its folder, Fal AI needs Python for the scripts in its folder, the command-line tools its instructions call (python and pip) and credentials named FAL_API_KEY. Our summary lists: Python 3; A credential in FAL_API_KEY.

Does Fal AI access the network?

SKILL.md names 1 domain. As links in the text: fal.ai. This is read from the text; nothing was executed.

Is Fal AI safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Fal AI use?

Fal AI is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fal AI use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.9k tokens, read only when the agent opens those files.

What are the alternatives to Fal AI?

Skills that share tags, products or a category with Fal AI: Fal AI Media (affaan-m/ECC, 277k stars), Super Video Maker (Bomx/super-video-maker-skill, 310 stars), Hearyourvoice (killernay/HearYourVOICE, 140 stars) and AI Video Gen (calesthio/OpenMontage, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fal AI?

mikeOnBreeze (a GitHub user) maintains it in mikeOnBreeze/cc-crossbeam, which has 293 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on March 2, 2026.

Source: mikeOnBreeze/cc-crossbeam on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.