Masonry Logo
AI & Technology

Introducing Gemini Omni Flash: Fast, Affordable AI Video on Masonry

Google's fastest, cheapest video model turns a sentence into a ten-second clip with sound in about thirty seconds. It's live on Masonry for text-to-video, image-to-video, and editing a clip by describing the change.

Gaurav BisenGaurav Bisen
5 min read

A ten-second video, with sound, from one sentence, in about the time it takes to read this paragraph. That's Gemini Omni Flash, Google's cheapest and fastest video model, which landed this week next to Nano Banana 2 Lite. It's on Masonry today.

It works three ways: prompt it and it makes a clip, hand it a still and it animates it, or hand it a clip and describe a change and it re-renders the scene. The last one is the new part, and it's why this launch matters more than another fast model.

What it is

Omni Flash is the budget tier of Google's video lineup, and Google is blunt about it: a "high quality, cost-efficient model for video generation and conversational editing," at $0.10 per second of output (blog.google). That matches Veo 3.1 Fast on price. What you're buying is throughput, not the last five percent of fidelity.

In practice you get a ten-second clip at 720p with an ambient audio track, back in about half a minute. Everything in this post came straight out of Omni Flash from a plain prompt, with no upscaling or cleanup. Tap the unmute button on any clip to hear it.

Text-to-video from one sentence: 'A golden retriever puppy bounding through a sunlit meadow of wildflowers, slow motion, shallow depth of field, warm cinematic light.'

Look at the light and the motion. Shallow depth of field, sun coming from one direction, the puppy's weight landing as it runs, all held for the full clip off one line of text.

Fast enough to iterate

The speed is the point. At thirty seconds a clip and a few cents each, you can try a direction, watch it, and try another before you lose the thread of what you wanted.

Lifestyle b-roll: a barista pouring steamed milk into a latte, warm morning light in a cozy cafe, close-up, cinematic, with ambient audio.

Short atmospheric b-roll is the obvious use: the clip behind a headline, or the one that drops into an ad cut.

Aerial nature shot: a drone pass over ocean waves crashing on a rugged coastline at golden hour, seabirds gliding, cinematic.

It doesn't come apart on bigger scenes, either. This one tracks the camera over moving water and light, again in one pass.

Animate a still

Text isn't the only way in. Give Omni Flash an image and a short instruction and it puts the scene into motion. When you already have the frame you want, this is the quickest route to a clip.

Image-to-video: starting from a single still frame, prompted to 'finish the pour and let the leaf pattern settle, gentle steam rising.'

Edit a clip by describing it

Editing is where it gets interesting. Give Omni Flash a clip you already have, tell it what to change in words, and it re-renders the scene to match. No masks, no keyframes, no timeline. You say the edit, it makes the edit.

Here's the puppy-in-the-meadow clip from the top of this post, followed by the same clip after one instruction: "change the season to winter with falling snow." Watch the two back to back.

Before: the original summer clip, straight from a text prompt.
After: one instruction, 'change the season to winter with falling snow.' Same puppy, same run, same camera, new season.

The subject and the motion survive; everything around them gets restyled. Swap day for night, recolor a product, or push a clip's whole grade somewhere new, all from a sentence and the clip you already made.

The same trick works from a prompt alone. Ask for a candle flame in the dark, then hand that clip back and say "turn the flame blue," and it recolors the fire while the flicker and the shadows stay put. It's the workflow the Masonry agent runs for you: generate a draft, then edit it in place instead of starting over.

Editing is happiest on short clips. You can reach it three ways: right on the canvas (select a video, hit Edit video, and type the change), by asking the Masonry agent to edit a clip in plain language, or through the API.

Where it sits next to Veo

None of this replaces Veo. Omni Flash is the tier below it, and the three cover different jobs:

  • Gemini Omni Flash — fastest and cheapest, ten seconds at 720p with ambient sound. For drafts, exploring directions, and generating in volume.
  • Veo 3.1 and Veo 3.1 Fast — higher fidelity, and the only choice when you need spoken dialogue and lip-sync (Hindi and other languages included), vertical or longer clips, or hero-shot quality.
  • Kling — silent, longer, cinematic, with tighter camera and per-object control.

In practice: draft on Omni Flash, then step up a tier for the final cut. Every model lives on one hub, so going from an Omni Flash draft to a Veo final is one click, no re-export into a different tool.

What it can't do yet

It's a preview, and it shows. Output is locked at about ten seconds, 720p, 16:9, with no way to set duration, aspect ratio, or resolution through the API yet. The audio is ambient, not speech, so talking heads and lip-sync belong on Veo 3.1. For vertical, longer, or higher-resolution work today, use Veo or Kling. Inside those lines, nothing beats it on speed per dollar.

Using it

It's live now. Pick Gemini Omni Flash in the model picker and prompt it, or attach an image to animate. Or just tell the Masonry agent you want a fast, cheap video draft and it will route there, then help you step up when the draft is worth finishing.

Fast, cheap, and good enough covers most of the work on most days. Omni Flash gives you motion and sound in about thirty seconds at the lowest price on the platform, and now it edits too. Try it on Masonry.

Share: