---
name: fal-ai-media
description: Create images, video, and audio with the fal.ai MCP server. Use for text-to-image, image edits, text-to-video, image-to-video, speech, music, sound effects, and video-to-audio tasks.
origin: ECC
---

# fal.ai Media Creation

> **Original skill by ECC. Credit: ECC.**
>
> **This skill can change fast.** Model names, prices, inputs, output types, and tool names may change. Check the current model data before you promise a model, price, setting, or file type.

Use the fal.ai MCP server to make images, video, and audio.

## When to use this skill

Use this skill when the user wants to:

- Make an image from text
- Edit an image
- Make a video from text or an image
- Make speech, music, or sound effects
- Add sound to a video
- Make a thumbnail or other media file

Do not use this skill when the user only wants to edit media with local tools.

## Setup

The fal.ai MCP server must be set up first. Add this entry to `~/.claude.json`:

```json
{
  "fal-ai": {
    "command": "npx",
    "args": ["-y", "fal-ai-mcp-server"],
    "env": {
      "FAL_KEY": "YOUR_FAL_KEY_HERE"
    }
  }
}
```

Get an API key from [fal.ai](https://fal.ai).

Never print, save, or share the key in chat. If the server or key is missing, explain what is missing and stop.

## Tools

The server may offer these tools:

- `search`: Find models by task or key words.
- `find`: Get the current inputs and rules for a model.
- `generate`: Start a media job.
- `result`: Get the result of a job.
- `status`: Check if a job is done.
- `cancel`: Stop a job that is still running.
- `estimate_cost`: Get a cost guess.
- `models`: List common models.
- `upload`: Upload a local file for use as input.

Tool names may change. Use the names shown by the active server.

## Safe work flow

Follow these steps:

1. Ask for any key detail that is missing, such as size, length, style, or source file.
2. Use `search` if the right model is not clear.
3. Use `find` before the first run. Check the exact model ID and allowed inputs.
4. Upload local source files with `upload`.
5. Check cost before a costly job or when the user asks about cost.
6. Run `generate` with only inputs that the model accepts.
7. If the job is not done at once, save its job ID.
8. Use `status` or `result` until it is done or fails.
9. Give the user the output link and key job details.

Do not keep retrying a failed paid job. Read the error first. Fix the input, then ask before a new costly run.

Do not say that a seed will make the same result every time. It may help, but models and systems can change.

## Image creation

Search for a current text-to-image model:

```text
search(query: "text to image")
find(endpoint_ids: ["<model_id>"])
```

Example call after `find` confirms the inputs:

```text
generate(
  app_id: "<model_id>",
  input_data: {
    "prompt": "A small red boat on a calm lake at sunrise",
    "image_size": "landscape_16_9",
    "num_images": 1,
    "seed": 42
  }
)
```

Common image inputs may include:

| Input | Use |
|---|---|
| `prompt` | Says what the image should show |
| `image_size` | Sets the shape and size |
| `num_images` | Sets how many images to make |
| `seed` | May help make a similar result |
| `guidance_scale` | May set how closely the model follows the prompt |

These inputs are not shared by every model. Check with `find`.

### Image editing

Upload the source image first:

```text
upload(file_path: "/path/to/image.png")
```

Then use a model that supports image editing:

```text
generate(
  app_id: "<image_edit_model_id>",
  input_data: {
    "prompt": "Keep the same scene. Change the style to soft watercolor.",
    "image_url": "<uploaded_url>"
  }
)
```

State what must stay the same and what must change. If the model needs a mask, ask the user for one or make sure they approve how the edit area will be chosen.

## Video creation

Find and check a current video model:

```text
search(query: "text to video")
find(endpoint_ids: ["<model_id>"])
```

Then run it with confirmed inputs:

```text
generate(
  app_id: "<model_id>",
  input_data: {
    "prompt": "A slow drone view over a mountain lake at sunrise",
    "duration": "5s",
    "aspect_ratio": "16:9",
    "seed": 42
  }
)
```

For image-to-video, upload the image and pass its URL:

```text
generate(
  app_id: "<image_to_video_model_id>",
  input_data: {
    "prompt": "The camera moves back slowly. A light wind moves the trees.",
    "image_url": "<uploaded_image_url>",
    "duration": "5s"
  }
)
```

Common video inputs may include:

| Input | Use |
|---|---|
| `prompt` | Says what happens in the video |
| `duration` | Sets the video length |
| `aspect_ratio` | Sets the frame shape |
| `seed` | May help make a similar result |
| `image_url` | Gives the first image or source image |

Check whether the model makes sound. Do not promise sound unless the model data says it does.

A good video prompt states:

- The main subject
- The action
- The camera move
- The light and mood
- Any sound the model should make

## Audio creation

Find the right model for the audio task:

```text
search(query: "text to speech")
search(query: "text to music")
search(query: "sound effects")
find(endpoint_ids: ["<model_id>"])
```

Example speech call:

```text
generate(
  app_id: "<speech_model_id>",
  input_data: {
    "text": "Hello. Welcome to the demo.",
    "speaker_id": 0
  }
)
```

Example video-to-audio call:

```text
generate(
  app_id: "<video_to_audio_model_id>",
  input_data: {
    "video_url": "<uploaded_video_url>",
    "prompt": "Soft forest sounds with birds and light wind"
  }
)
```

Check language, voice, length, file type, and use rights before the run. Do not copy a real person's voice without clear permission.

Other audio tools, such as ElevenLabs or VideoDB, are outside this skill. Do not call them unless the user asks for them and they are set up.

## Cost checks

Use the active server's cost tool before a costly run:

```text
estimate_cost(
  estimate_type: "unit_price",
  endpoints: {
    "<model_id>": {
      "unit_quantity": 1
    }
  }
)
```

A cost result is only a guess. Tell the user what unit it covers. If cost data is missing, say so. Do not guess a price.

## Full usage example

User request:

> Make a five-second, wide video of ocean waves at sunset. Keep the cost low.

Work flow:

```text
search(query: "low cost text to video")
find(endpoint_ids: ["<chosen_model_id>"])
estimate_cost(
  estimate_type: "unit_price",
  endpoints: {
    "<chosen_model_id>": {
      "unit_quantity": 1
    }
  }
)
generate(
  app_id: "<chosen_model_id>",
  input_data: {
    "prompt": "Wide view of ocean waves at sunset. Warm orange light. The camera stays still. Natural wave motion.",
    "duration": "5s",
    "aspect_ratio": "16:9"
  }
)
status(request_id: "<job_id>")
result(request_id: "<job_id>")
```

If `find` shows different input names, use those names. If the model does not support five seconds or `16:9`, tell the user and offer the closest valid choice.

## Edge cases

- If no model fits, say what part is not supported.
- If a local file is missing, ask for the right path or file.
- If upload fails, check file size and file type.
- If a URL cannot be read, ask for a public URL or a local file.
- If the prompt breaks a model rule, explain the issue and ask for a safe change.
- If the job takes a long time, report that it is still running. Do not start a copy.
- If the job fails after a charge, keep the job ID and error text for the user.
- If the output is empty or broken, check the result data before trying again.
- If the user asks for many files, confirm the count and cost first.
- If the user asks to cancel, call `cancel` with the saved job ID.
- If the output link will expire, tell the user to save the file soon.
- If the user needs private media, warn them before upload and get clear approval.

## Tips

- Start with a low-cost draft when the user wants to test ideas.
- Use one image or a short video first.
- Keep prompts clear and short.
- For image edits, list what must not change.
- For video, focus on action and camera moves.
- Image-to-video often gives more control than text-to-video.
- Check cost before long video, high quality, or large batch jobs.

## Related skills

- `videodb`: Video work, edits, and streams
- `video-editing`: Video editing steps
- `content-engine`: Social media content creation