# Google Gemini Omni Flash 1.1 Text to Video - Core.Today AI API > Google Gemini Omni Flash 1.1 text-to-video via Fal.AI. Generates 3-10 second video with synchronized native audio from a text prompt, grounded in Gemini's real-world knowledge and physics understanding, with cinematic camera control in natural language — 16:9 or 9:16 at 360p to 4K. - **Provider**: Google - **Model ID**: google/gemini-omni-flash/v1.1/text-to-video - **Category**: Video Generation - **Credits**: 1860 per 8s 720p video — billed per second (360p 70/s, 720p 233/s, 1080p 349/s, 4k 698/s) - **Speed**: Fast - **Quality**: High ## Features - Synchronized native audio (dialogue, music, ambience) from the prompt - Physics-aware motion grounded in Gemini's world knowledge - Cinematic camera control in natural language - Four resolution tiers from 360p to 4K, billed per second ## Use Cases - Short ads and social clips with sound - Cinematic establishing shots and b-roll - Fast 360p concept previews before a high-resolution render - Explainer scenes that need plausible real-world physics ## API Endpoint Base URL: https://api.core.today/v1 Create Prediction: POST /predictions Get Status: GET /predictions/{job_id} ## Authentication Header: X-API-Key: YOUR_API_KEY ## Input Parameters ### Required - **prompt**: string - The text prompt describing the video. Describe scene, camera movement, lighting, mood and audio. ### Optional - **duration**: integer (default: 8) - The duration of the generated video in seconds (3-10), billed per second. - **aspect_ratio**: string (default: 16:9) - The aspect ratio of the generated video. Options: 16:9, 9:16 - **resolution**: string (default: 720p) - The resolution of the generated video. Higher resolutions cost more per second. Options: 360p, 720p, 1080p, 4k ## Examples ### Lighthouse at dusk A cinematic wide shot at the default 720p and 8 seconds. ```json { "model": "google/gemini-omni-flash/v1.1/text-to-video", "input": { "prompt": "A cinematic wide shot of a lighthouse on a rocky cliff at dusk, waves crashing below, the beam sweeping across the dark sea.", "duration": 8, "aspect_ratio": "16:9", "resolution": "720p" } } ``` ### Vertical street-food clip A 5-second 9:16 clip with ambient sound at 1080p. ```json { "model": "google/gemini-omni-flash/v1.1/text-to-video", "input": { "prompt": "Close-up of sizzling skewers on a night-market grill, smoke rising, vendors chatting in the background, handheld camera.", "duration": 5, "aspect_ratio": "9:16", "resolution": "1080p" } } ``` ## Response Format ```json { "job_id": "abc123", "status": "pending | processing | completed | failed", "result": "URL or data (when completed)" } ``` ## Usage Flow 1. POST /predictions with model and input -> receive job_id 2. GET /predictions/{job_id} -> poll until status is completed or failed 3. Result contains output URL(s) ## Tips - Describe audio (dialogue, music, ambience) directly in the prompt — it is generated with the video. - Draft at 360p (70/s), then re-render the keeper at 720p, 1080p or 4k. - Billing is duration x per-second rate, so shorter clips cost proportionally less. - The same model is also available as google/gemini-omni-1.1 with editing and reference modes. ## Documentation https://fal.ai/models/google/gemini-omni-flash/v1.1/text-to-video