Ground Truth.
AI, checked against the source.

News · 2026-08-27

Gemini Omni 1.1 Flash can extend a scene instead of restarting it

Google released Gemini Omni 1.1 Flash, an update to its generative video model whose main new capability is continuing an existing clip while reading up to ten seconds of what came before -- a jump from previous models that referenced only the final second. The release also adds first-and-last-frame control, 360p draft generation at roughly a third the cost of 720p, 4K upscaling, and the ability to supply up to three seconds of reference video for character consistency.

Key facts

The context window on scene extension is the substantive change, and it is easy to under-rate. Generative video models produce short clips, and the standard trick for making something longer is to feed the last frame back in and generate onward. That works about as well as writing a novel where each chapter begins by looking only at the final sentence of the previous one. Characters drift, lighting shifts, a jacket changes colour. Giving the model ten seconds of prior footage means it is continuing a shot rather than guessing from a still.

The keyframe feature attacks the same problem from the other end. Specify a starting frame and an ending frame and the model generates the movement between them, which is how you get a camera orbit that actually returns to where it started, or a loop that closes cleanly. Anyone who has tried to art-direct a generative video model by prompt alone will recognise why pinning both ends of a shot is more useful than another adjective.

The pricing tier is the other half of the story and probably the more consequential half. Google's pricing page lists $1.50 per million input tokens covering text, image, video and audio, and $17.50 per million tokens of video output, which Google's own footnote translates to roughly ten cents per second of 720p video. The 360p draft mode exists so you do not pay that rate to discover a shot does not work. The intended workflow is explicit in the announcement: generate three or four cheap variations, vary one thing at a time, compare them side by side, then render the keeper at 4K.

That is a production pipeline, not a demo, and the customers Google names back it up. Adobe has integrated the model into Firefly. "Gemini Omni Flash is one of the strongest video models available in Figma Weave, where the canvas helps creative teams build on every generation," said Itay Schiff, Creative Director at Figma Weave, adding that the new controls take teams "beyond generating videos to truly directing them."

Why it matters: the competitive question in generative video has shifted from fidelity to controllability and unit cost. A model that produces a beautiful clip you cannot extend, loop or match to an existing shot is a toy for social posts. Ten seconds of context, keyframe endpoints and a cheap draft tier are the boring features that let the output enter an edit timeline. It also arrives a day after Google's transcription model that edits what you said, continuing a pattern of shipping the unglamorous production plumbing rather than the headline demo.

The caveats are real. Forty seconds total is still short, output runs 3 to 10 seconds per generation at 24 frames per second, and there is no downloadable checkpoint -- this is API-only through Google AI Studio, the Gemini Enterprise Agent Platform, Google Flow and the Gemini app. The Hacker News discussion, which drew 198 points and 146 comments, is engaged but pointed: commenters note the model still cannot sync generated video to supplied audio, and that at ten cents a second the economics remain rough for anything casual. For a thirty-second finished spot with a normal number of takes, that is a real bill -- which is exactly why the 360p draft tier exists.


Primary source, verified: read the paper →

Key questions

How long a video can it make?

Clips extend in ten-second increments to a cumulative total of 40 seconds, with each extension able to read up to ten seconds of prior footage for consistency.

What does it cost?

Google's pricing lists $1.50 per million input tokens and $17.50 per million tokens of video output, which the company's own footnote works out to roughly ten cents per second of 720p video.

Are the weights available?

No. Gemini Omni 1.1 Flash is API-only, reachable through Google AI Studio, the Gemini Enterprise Agent Platform, Google Flow and the Gemini app, with no downloadable checkpoint.
Cite this

APA

Ground Truth. (2026, August 27). Gemini Omni 1.1 Flash can extend a scene instead of restarting it. Ground Truth. https://groundtruth.day/news/gemini-omni-1-1-flash-can-extend-a-scene-instead-of-restarting-it.html

BibTeX

@misc{groundtruth:gemini-omni-1-1-flash-can-extend-a-scene-instead-of-restarting-it,
  title  = {Gemini Omni 1.1 Flash can extend a scene instead of restarting it},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {aug},
  url    = {https://groundtruth.day/news/gemini-omni-1-1-flash-can-extend-a-scene-instead-of-restarting-it.html}
}

Topics: google · video · models · api · creative-tools · multimodal · generative-media

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.