fal.ai Media Creation
Original skill by ECC. Credit: ECC.
This skill can change fast. Model names, prices, inputs, output types, and tool names may change. Check the current model data before you promise a model, price, setting, or file type.
Use the fal.ai MCP server to make images, video, and audio.
When to use this skill
Use this skill when the user wants to:
- Make an image from text
- Edit an image
- Make a video from text or an image
- Make speech, music, or sound effects
- Add sound to a video
- Make a thumbnail or other media file
Do not use this skill when the user only wants to edit media with local tools.
Setup
The fal.ai MCP server must be set up first. Add this entry to ~/.claude.json:
{
"fal-ai": {
"command": "npx",
"args": ["-y", "fal-ai-mcp-server"],
"env": {
"FAL_KEY": "YOUR_FAL_KEY_HERE"
}
}
}Get an API key from fal.ai.
Never print, save, or share the key in chat. If the server or key is missing, explain what is missing and stop.
Tools
The server may offer these tools:
search: Find models by task or key words.find: Get the current inputs and rules for a model.generate: Start a media job.result: Get the result of a job.status: Check if a job is done.cancel: Stop a job that is still running.estimate_cost: Get a cost guess.models: List common models.upload: Upload a local file for use as input.
Tool names may change. Use the names shown by the active server.
Safe work flow
Follow these steps:
- Ask for any key detail that is missing, such as size, length, style, or source file.
- Use
searchif the right model is not clear. - Use
findbefore the first run. Check the exact model ID and allowed inputs. - Upload local source files with
upload. - Check cost before a costly job or when the user asks about cost.
- Run
generatewith only inputs that the model accepts. - If the job is not done at once, save its job ID.
- Use
statusorresultuntil it is done or fails. - Give the user the output link and key job details.
Do not keep retrying a failed paid job. Read the error first. Fix the input, then ask before a new costly run.
Do not say that a seed will make the same result every time. It may help, but models and systems can change.
Image creation
Search for a current text-to-image model:
search(query: "text to image")
find(endpoint_ids: ["<model_id>"])Example call after find confirms the inputs:
generate(
app_id: "<model_id>",
input_data: {
"prompt": "A small red boat on a calm lake at sunrise",
"image_size": "landscape_16_9",
"num_images": 1,
"seed": 42
}
)Common image inputs may include:
| Input | Use |
|---|---|
prompt |
Says what the image should show |
image_size |
Sets the shape and size |
num_images |
Sets how many images to make |
seed |
May help make a similar result |
guidance_scale |
May set how closely the model follows the prompt |
These inputs are not shared by every model. Check with find.
Image editing
Upload the source image first:
upload(file_path: "/path/to/image.png")Then use a model that supports image editing:
generate(
app_id: "<image_edit_model_id>",
input_data: {
"prompt": "Keep the same scene. Change the style to soft watercolor.",
"image_url": "<uploaded_url>"
}
)State what must stay the same and what must change. If the model needs a mask, ask the user for one or make sure they approve how the edit area will be chosen.
Video creation
Find and check a current video model:
search(query: "text to video")
find(endpoint_ids: ["<model_id>"])Then run it with confirmed inputs:
generate(
app_id: "<model_id>",
input_data: {
"prompt": "A slow drone view over a mountain lake at sunrise",
"duration": "5s",
"aspect_ratio": "16:9",
"seed": 42
}
)For image-to-video, upload the image and pass its URL:
generate(
app_id: "<image_to_video_model_id>",
input_data: {
"prompt": "The camera moves back slowly. A light wind moves the trees.",
"image_url": "<uploaded_image_url>",
"duration": "5s"
}
)Common video inputs may include:
| Input | Use |
|---|---|
prompt |
Says what happens in the video |
duration |
Sets the video length |
aspect_ratio |
Sets the frame shape |
seed |
May help make a similar result |
image_url |
Gives the first image or source image |
Check whether the model makes sound. Do not promise sound unless the model data says it does.
A good video prompt states:
- The main subject
- The action
- The camera move
- The light and mood
- Any sound the model should make
Audio creation
Find the right model for the audio task:
search(query: "text to speech")
search(query: "text to music")
search(query: "sound effects")
find(endpoint_ids: ["<model_id>"])Example speech call:
generate(
app_id: "<speech_model_id>",
input_data: {
"text": "Hello. Welcome to the demo.",
"speaker_id": 0
}
)Example video-to-audio call:
generate(
app_id: "<video_to_audio_model_id>",
input_data: {
"video_url": "<uploaded_video_url>",
"prompt": "Soft forest sounds with birds and light wind"
}
)Check language, voice, length, file type, and use rights before the run. Do not copy a real person's voice without clear permission.
Other audio tools, such as ElevenLabs or VideoDB, are outside this skill. Do not call them unless the user asks for them and they are set up.
Cost checks
Use the active server's cost tool before a costly run:
estimate_cost(
estimate_type: "unit_price",
endpoints: {
"<model_id>": {
"unit_quantity": 1
}
}
)A cost result is only a guess. Tell the user what unit it covers. If cost data is missing, say so. Do not guess a price.
Full usage example
User request:
Make a five-second, wide video of ocean waves at sunset. Keep the cost low.
Work flow:
search(query: "low cost text to video")
find(endpoint_ids: ["<chosen_model_id>"])
estimate_cost(
estimate_type: "unit_price",
endpoints: {
"<chosen_model_id>": {
"unit_quantity": 1
}
}
)
generate(
app_id: "<chosen_model_id>",
input_data: {
"prompt": "Wide view of ocean waves at sunset. Warm orange light. The camera stays still. Natural wave motion.",
"duration": "5s",
"aspect_ratio": "16:9"
}
)
status(request_id: "<job_id>")
result(request_id: "<job_id>")If find shows different input names, use those names. If the model does not support five seconds or 16:9, tell the user and offer the closest valid choice.
Edge cases
- If no model fits, say what part is not supported.
- If a local file is missing, ask for the right path or file.
- If upload fails, check file size and file type.
- If a URL cannot be read, ask for a public URL or a local file.
- If the prompt breaks a model rule, explain the issue and ask for a safe change.
- If the job takes a long time, report that it is still running. Do not start a copy.
- If the job fails after a charge, keep the job ID and error text for the user.
- If the output is empty or broken, check the result data before trying again.
- If the user asks for many files, confirm the count and cost first.
- If the user asks to cancel, call
cancelwith the saved job ID. - If the output link will expire, tell the user to save the file soon.
- If the user needs private media, warn them before upload and get clear approval.
Tips
- Start with a low-cost draft when the user wants to test ideas.
- Use one image or a short video first.
- Keep prompts clear and short.
- For image edits, list what must not change.
- For video, focus on action and camera moves.
- Image-to-video often gives more control than text-to-video.
- Check cost before long video, high quality, or large batch jobs.
Related skills
videodb: Video work, edits, and streamsvideo-editing: Video editing stepscontent-engine: Social media content creation