Tasting Notes: Omni
Where the FLORA team reviews the latest text, image, and video models
Most video model reviews ask the same question: how pretty is the output? By that measure, Omni, Google’s new multimodal video model, will disappoint you. Prompt it cold for a moody campaign film and you'll get something pretty competent, but a little generic.
We took a different approach in our testing. We ran Omni through an internal evaluation across eight production categories, scored by two human annotators (one an AI PM with multimodal eval experience, one a creative workflow specialist from video production), against the highest-volume workflows we see in FLORA: advertising, fashion campaigns, e-commerce product visualization, and VFX compositing. Omni lost the same way every time and won the same way every time. Weak when generating footage from nothing. Strong, sometimes really strong, when changing footage that already exists.
Four places it earned a spot in our workflows.
Multi-element composition straight to video
This was the workflow result that surprised us most. We connected three separate inputs on the canvas: a model portrait, a product shot of the sneaker, and an outdoor scene with a stone pedestal. Instead of compositing them into a single image first, we prompted Omni to generate the video directly from all three. One node, one step. Both annotators scored it 4 out of 5. The comparison model needed the intermediate composite and still lost the shoe's design details in motion.
Omni is an editor, not a filmmaker. Use it accordingly
Ai • Articles