Specifications
- Input
- Text prompt or single start image
- Output
- MP4 video download
- Duration
- 4s, 6s, 8s, or 10s
- Quality options
- Standard or high fidelity
- Credits
- From 90 credits · scales with duration
Flexible clip length
Gemini Omni supports 4, 6, 8, or 10 second clips from text or a single start image — Google's multimodal stack inside the same Meta Image account.

Clip duration
Pick your duration
Choose 4s for quick hooks, or up to 10s when the scene needs more room to breathe.
Text-to-video
Describe motion, camera, and mood in plain language — no start frame required.
Image-to-video
Upload one start image and animate it into a short MP4 clip.
Example outputs

Cinematic text-to-video
Try →“Wide establishing shot with slow push-in, moody atmosphere, cinematic color grade”

Blog hero motion
Try →“Abstract gradient motion background for a tech blog hero section, subtle parallax”

Social clip concept
Try →“Short upbeat social clip, bold colors, quick camera movement”
What Gemini Omni excels at

Flexible duration — 4, 6, 8, or 10 second clips depending on your needs.

Text-to-video — describe motion, camera, and mood in plain language.

Image-to-video — animate a single start frame into a short clip.

Multimodal workflows — fits alongside Muse Image and other engines in one studio.
Best for
Flexible short clips, Google multimodal workflows, and quick text-to-video when you do not need end-frame control.
Related
Does Gemini Omni support end frames?
No. Gemini Omni accepts text prompts or a single start image. For end-frame control, use Kling 3.0 or Seedance 2.0.
How does Gemini Omni compare to Kling?
Kling leans cinematic with native audio and end-frame support. Gemini Omni offers flexible duration and Google's multimodal stack — pick based on your clip style.
Third-party models
Model names refer to third-party AI integrations available through Meta Image. Meta Image is an independent platform — not affiliated with or endorsed by Meta Platforms, Inc. or any model provider.