Google has released Gemini Omni 1.1 Flash with a product surface that looks less like a one-shot generator and more like an editing session. A developer can create a draft, pass its previous_interaction_id into another request, revise it through natural language, extend a scene, interpolate between first and last frames, or upscale an accepted result. Google presents 360p as a faster drafting surface and 4K as a finishing option.
That changes the product question. A team is no longer deciding whether one prompt produced an impressive clip. It is deciding whether a chain of generated clips can survive review, branching, correction, approval, retention, and publication without losing the reason each version exists.
The interaction ID helps the API find prior state. It does not tell your team which visual details were supposed to stay fixed, what changed unexpectedly, whether an earlier branch was better, which reference asset was licensed, what a reviewer actually approved, or whether the final download still carries the expected provenance signal. In other words, server-side state is not creative version control.
This field note proposes an 18-case branch-and-acceptance experiment and a reusable video-version receipt. We did not call Gemini Omni 1.1 Flash, generate the fixtures, reproduce Google’s demonstrations, or compare video providers. The workflow, thresholds, and blank result fields below are a protocol for a small team to run, not Y Build test results.
What launched, and what remains unproved
Google’s launch announcement describes Gemini Omni 1.1 Flash as a production-ready update available through the Gemini API. It highlights scene extension in ten-second increments up to a cumulative 40 seconds, first-and-last-frame control, 360p drafts, and upscaled 4K output. For extension, Google says the model can analyze as much as ten seconds of prior context instead of only the last second used by earlier models.
The corresponding Gemini Omni API guide documents text-to-video, image-to-video, reference-to-video, edit, and extend tasks. It also exposes important boundaries: extension appends to the end rather than inserting into the middle; some final input frames may be edited to smooth a transition; uploaded talking video cannot currently be extended with additional dialogue; voice editing is unsupported; multi-video reasoning is unsupported; and English is fully supported while other languages have not been evaluated.
Those are useful product facts, not proof that a particular brand character will remain consistent through four edits, that dialogue will preserve every approved word, that 4K upscaling will repair a bad composition, or that a 40-second chain will satisfy a campaign brief. The launch examples are selected demonstrations. A small team still needs evidence shaped like its own deliverable.
The bounded conclusion is simple: iterative control is now accessible enough to test as a workflow. It has not made generated video deterministic, mergeable, rights-cleared, or publication-ready by default.
An interaction chain is a state pointer, not a commit graph
The Interactions API documentation says previous_interaction_id preserves earlier inputs and outputs so a later request can continue from that history. Other request parameters remain interaction-scoped and must be specified again when needed. That distinction matters. A child interaction can inherit creative context without inheriting every configuration or product rule your application assumes.
A software commit normally binds content to an identifier, records ancestry, and allows an earlier state to be retrieved. A creative review needs more:
- the exact source assets and their hashes;
- the model identifier, task, resolution, region, and client version;
- the prompt and negative requirements expressed in ordinary prompt text;
- an explicit list of invariants that must not change;
- the output file hash and durable local location;
- reviewer observations by timecode;
- the accepted purpose, channel, territory, and expiry;
- the relationship between draft, branch, extension, upscale, and published export.
Do not call a sequence v1, v2, and final-final and assume the history is recoverable. Save every candidate as an immutable node. Represent a new edit or extension as a child. If two reviewers explore different directions, create two children from the same accepted parent instead of letting the newest server interaction silently become the only path.
There is no conventional merge operation for two generated videos. A prompt that asks the model to combine the best parts is another generation step, not a deterministic merge. It must create a new node with two declared ingredients and its own review.
Start with one publishable job and six locked invariants
Imagine a four-person product team preparing a 20-second vertical launch clip for a fictional calendar assistant called Cedar Calendar. The approved concept shows a red paper calendar on a wooden desk, a hand circling a date, the room shifting from morning to evening, and a final empty area where the product team will add a real interface capture in post-production.
The team deliberately avoids generating the product UI or brand text. Generated typography is easy to misread, and a fictional interface can overpromise product behavior. The generated clip is an ingredient, not the completed advertisement.
Before the first request, lock six invariants:
- the calendar stays red and retains the same shape;
- the circled date remains the 14th;
- the hand has no visible jewelry or identifying marks;
- the camera stays in one continuous overhead shot;
- the lower-right safe area remains visually quiet for a real UI overlay;
- no logo, readable product claim, public figure, or recognizable private person appears.
Then define editable dimensions: lighting, camera speed, background props, sound texture, and the transition from morning to evening. This separation gives reviewers something better than “looks close.” They can reject an output because a locked invariant drifted even if the new lighting is attractive.
For another product, the invariants may be packaging geometry, garment color, ingredient count, accessibility contrast, an instructional hand position, or the absence of a medical claim. The point is not to freeze creativity. It is to make intentional change distinguishable from accidental mutation.
Build an 18-case branch-and-acceptance fixture
Use six content briefs and run three workflow arms for each. Eighteen cases will not estimate a general model success rate. It is small enough for the whole team to inspect every frame transition, prompt, output, and decision.
| Brief family | Deliberate stress | Locked evidence |
|---|---|---|
| Product object | Color, geometry, orientation, and negative space | Reference image plus invariant list |
| Human action | Hands, object contact, motion sequence | Keyframes and timecoded action rubric |
| Camera move | Dolly, orbit, or overhead continuity | Start/end composition and motion path |
| Scene transition | Lighting or location changes over time | Transition boundary and unchanged subjects |
| Dialogue/audio | Exact short line, speaker continuity, ambience | Approved script and audio review sheet |
| Localization | Non-English on-screen context or spoken direction | Fluent reviewer and locale-specific brief |
Each brief gets the same three arms:
Arm A — draft and upscale. Generate at 360p, select one candidate, then request the final resolution. This tests whether a cheap draft predicts the accepted high-resolution output rather than merely reducing initial cost.
Arm B — stateful edit branch. Create two children from the same saved parent: one changes an editable dimension, while the other asks for the same change plus a locked-invariant reminder. This tests whether the explicit invariant list reduces unintended change.
Arm C — extension. Extend an accepted short clip once, then extend the resulting child again. Review both the new material and the final seconds of the parent because the API guide warns that frames near the join can be modified for continuity.
Freeze the briefs, rubrics, reference hashes, reviewer assignments, and decision rules before running the candidate. Do not replace a failed prompt after seeing the output and still count the repaired version as the original case. Keep repair attempts, blocked requests, and discarded generations in the ledger.
Review the delta, not only the newest clip
A reviewer watching only the newest output is likely to miss identity drift, background substitution, an altered prop, or a softened product constraint. Every child needs a parent-versus-child review.
Use six separate dimensions:
| Dimension | Review question | Evidence |
|---|---|---|
| Requested change | Did the intended edit occur? | Prompt-specific pass/fail with timecode |
| Invariant preservation | Did every locked property survive? | One result per invariant, no averaging |
| Temporal continuity | Are subject, background, motion, and audio coherent? | Join review plus full-play review |
| Factual/physical plausibility | Does the scene violate obvious object or motion constraints? | Human note; specialist check when relevant |
| Usability | Can the clip be placed in the intended layout and channel? | Crop, safe-area, duration, caption, and sound-off checks |
| Repair burden | How much human and generation work remains? | Minutes, extra calls, edits, and discarded outputs |
This decomposition follows the spirit of research benchmarks without pretending that a public benchmark is a campaign acceptance test. VBench separates subject consistency, background consistency, flicker, motion smoothness, aesthetics, imaging quality, and multiple semantic dimensions. VideoScore was trained from human multi-aspect ratings covering visual quality, temporal consistency, dynamic degree, text alignment, and factual consistency. Both are useful reminders that “video quality” is not one scalar.
For a small product workflow, use automatic metrics only as triage. They can flag flicker or large visual changes, but the release decision should include a human who knows the brief. The reviewer must be able to reject an attractive clip that violates the product truth, rights boundary, or safe-area requirement.
The reusable video-version receipt
Create one receipt per output node, not one receipt for the entire session.
video_version_receipt:
node_id: "cedar-calendar-brief-03-edit-b"
parent_nodes: ["cedar-calendar-brief-03-draft-accepted"]
purpose:
campaign: "calendar-launch"
channel: "vertical-social-draft"
territory: "record-intended-territory"
expires_at: null
provider:
model: "pin-exact-model-id"
interaction_id: ""
previous_interaction_id: ""
task: "text_to_video|image_to_video|reference_to_video|edit|extend"
resolution: "360p|720p|1080p|4k"
region: ""
client_version: ""
ingredients:
- id: "calendar-reference"
sha256: ""
license_receipt: "rights/calendar-reference.md"
consent_receipt: null
instruction:
prompt_sha256: ""
requested_change: ""
locked_invariants: []
output:
file_sha256: ""
duration_seconds: null
generated_at_utc: ""
local_archive: ""
synthid_check: "not_run|detected|not_detected|unclear"
content_credentials_check: "not_run|valid|invalid|absent"
review:
requested_change: "pass|fail"
invariant_results: []
continuity_findings: []
rights_findings: []
repair_minutes: null
generation_calls_to_accept: null
decision:
state: "candidate|accepted_parent|approved_export|rejected|withdrawn"
approved_by: null
approved_at_utc: null
allowed_edits_after_approval: []
Blank fields are intentional. The receipt must not imply that rights, consent, provenance, or review passed merely because the API returned a video. Record not_run, absent, or unknown rather than converting missing evidence into a green check.
Hash the downloaded file, not only the interaction response. Save the prompt as a separate versioned artifact if it contains sensitive campaign information. When the final asset is trimmed, captioned, color-corrected, or combined with real footage, create a new export receipt that names the generated node as an ingredient.
Draft resolution is an experiment arm, not a promise
Google recommends 360p previews for faster and cheaper iteration, and the API guide labels 1080p and 4K outputs as upscaled. That suggests a rational funnel: explore cheaply, then pay to finish only the accepted direction. It does not establish that selection at 360p predicts acceptance at 4K.
Small details can change the decision. A hand deformation may be invisible in a draft. Texture, type, edge artifacts, facial features, or background objects can become more noticeable after upscaling. Conversely, a reviewer may reject a rough-looking 360p candidate whose composition would have survived at final resolution.
Measure the funnel directly:
- 360p generations per brief;
- minutes to choose a draft;
- share of selected drafts still accepted after final-resolution review;
- new defects first visible at final resolution;
- total generation calls and provider cost per approved export;
- human repair minutes after the final generation.
The current Gemini API pricing page lists a paid Gemini Omni Flash preview surface and explains video billing in output-token terms. Treat that as a planning reference, not a guaranteed Omni 1.1 quote. Capture the actual model, billed units, account tier, and invoice evidence during the experiment. “Cheap draft” is a hypothesis until the accepted-export ledger includes retries and repair.
Stateful editing creates a retention decision
The convenience of previous_interaction_id depends on stored state. Google’s Interactions API guide says interactions default to store=true; it currently documents 55-day retention for paid-tier projects and one day for free-tier projects, with shorter paid-tier settings available. It also says store=false prevents subsequent use of previous_interaction_id and is incompatible with background execution.
The zero-data-retention guide makes the trade-off explicit: stateful interactions, uploaded files, logs, and caches have separate handling, and File API objects must be deleted or allowed to expire according to their own lifecycle. Paid-service training restrictions do not mean every operational copy disappears immediately.
Before uploading a customer’s product, face, unreleased package, or campaign footage, choose one mode:
| Mode | Creative capability | Evidence obligation |
|---|---|---|
| Stateless exploration | No server-side branch continuation | Resend approved context; verify application logs and file handling |
| Stateful private draft | Iterative edits with stored interaction history | Approved data class, retention window, access owner, and deletion test |
| Publication pipeline | Accepted node enters downstream editing/export | Durable local archive, rights receipt, provenance checks, withdrawal path |
Do not hide this choice inside an SDK default. Product design should expose when a creative session is stored, how long it remains retrievable, who can see it, and what breaks after deletion. Run one deletion fixture and confirm the interaction, uploaded ingredients, application cache, review preview, and derived export follow the declared lifecycle.
Watermarks and provenance do different jobs
The Omni API guide says generated videos include SynthID. Google’s SynthID documentation describes an imperceptible watermark designed to remain detectable through common changes such as cropping, filtering, frame-rate changes, and lossy compression. Google’s verification guidance also warns that a negative or unclear result does not prove a file was not AI-generated.
SynthID is an origin signal for Google-generated media. It is not a license receipt, consent record, prompt history, complete edit log, or proof that a clip is factually safe to publish.
C2PA 2.4 provides a complementary standard for signed Content Credentials, ingredients, and actions such as creation and editing. A team can preserve the generated node as an ingredient when it adds captions, product footage, or human editing. NIST’s Generative AI Profile likewise treats provenance tracking and synthetic-content detection as useful transparency mechanisms while emphasizing their limitations and organizational context.
The release gate should therefore ask four different questions:
- Is the downloaded file the exact node that was approved?
- Can available watermark tooling identify Google-generated material?
- Does a valid provenance manifest describe the ingredients and later edits?
- Do separate rights and consent records authorize this use?
No single yes answers the other three.
Promotion rules should be branch-specific and non-averaged
Run the 18 cases with the exact API surface you plan to ship. Pin the model identifier and client version, record region and safety behavior, and have at least two people review high-value or identity-bearing assets. A beautiful average score must not erase one prohibited claim or one unlicensed likeness.
Promote a workflow only if every hard gate passes:
| Gate | Pass condition |
|---|---|
| Ancestry | Every approved export resolves to immutable parent and ingredient hashes |
| Requested change | The accepted child performs the requested edit or extension |
| Invariants | Zero launch-critical invariant failures in approved outputs |
| Branch recovery | An earlier accepted node can be retrieved without relying on the newest interaction |
| Rights and consent | Every reference and recognizable person has scoped evidence or is excluded |
| Provenance | File hash, watermark check, and manifest status are recorded without overclaiming |
| Retention | Storage mode, deletion trigger, and one end-to-end deletion test are documented |
| Economics | Calls, blocked runs, discarded outputs, and repair minutes are included per approved export |
Choose among explore, internal-draft, limited-publish, or hold. Promotion belongs to a specific brief family, locale, reference type, duration, region, and output channel. Passing product-object clips does not authorize recognizable-person advertising. Passing silent scene extension does not authorize dialogue-heavy localization.
Retest when the model alias, interaction API, safety behavior, region, resolution path, watermark tooling, reference policy, or editing application changes.
Failure modes to force before launch
The newest interaction overwrites the best branch. The team cannot retrieve the previously approved composition without regenerating it.
The requested edit succeeds while a locked detail drifts. Lighting improves, but the product color, date, hand, or safe area changes.
The join rewrites accepted footage. Reviewers inspect only the new ten seconds and miss changes near the extension boundary.
360p approval is treated as 4K approval. A final-resolution defect appears after the creative decision has already been signed.
Prompt history is mistaken for rights history. A reference URL is present, but ownership, consent, territory, and expiry are unknown.
A watermark becomes a truth certificate. Detection establishes a Google AI signal, not factuality, permission, or an unedited chain.
Stored state is omitted from the privacy diagram. Interaction history, uploaded media, review previews, or local exports outlive the stated purpose.
A blocked generation disappears from cost reporting. The successful clip looks inexpensive because retries, safety blocks, and human repair were excluded.
Localization is judged by an English-speaking reviewer. The API documentation says non-English languages have not been evaluated; fluent review is replaced by visual polish.
Where this protocol applies—and where it does not
Use this protocol when a small team is integrating iterative generated video into marketing drafts, product storytelling, social variants, onboarding visuals, concept previews, or creative tooling. It is most valuable once a user can edit or extend earlier output and the product needs to explain what was accepted.
It does not certify Gemini Omni 1.1 Flash, establish a provider ranking, measure rare harms, or reproduce Google’s claims. It is not sufficient for news evidence, political communication, biometric use, children’s content, medical instruction, legal proof, safety-critical training, or any workflow where a synthetic scene could be mistaken for a real event. Those uses require specialist review and stronger controls.
The fixture also does not prove copyright ownership or resolve likeness, trademark, music, union, advertising, or platform-disclosure obligations. Rights depend on the ingredients, instructions, output, contract, jurisdiction, and intended use. Consult qualified counsel where the stakes require it.
A 48-hour Build Lab plan
Hours 0–4: choose one publishable video job. Freeze six content briefs, locked invariants, editable dimensions, rights exclusions, and intended channels.
Hours 4–10: implement immutable node storage and the video-version receipt. Hash source assets, prompts, outputs, and local archives. Make parent and ingredient links visible in the review UI.
Hours 10–20: run the draft/upscale, stateful-edit, and extension arms. Keep every failed, blocked, discarded, and repaired result. Do not use production customer assets.
Hours 20–30: conduct parent-versus-child review. Inspect extension joins, final-resolution output, sound-off viewing, captions, crop safety, factual claims, recognizable people, and all locked invariants.
Hours 30–36: calculate calls, billed units, latency, reviewer time, and repair minutes per approved export. Separate provider generation from downstream human editing.
Hours 36–42: verify file hashes, SynthID status, Content Credentials status, ingredient rights, and consent records. Run one withdrawal and one deletion fixture through every named store.
Hours 42–48: sign explore, internal-draft, limited-publish, or hold for each brief family. Record exact model, API, region, language, reference type, resolution path, reviewer, and retest triggers.
Gemini Omni 1.1 Flash makes iterative video easier to build into a product. That convenience should make the creative history more explicit, not less. Save each output as evidence, compare every child with its parent, and approve the asset your audience will actually see—not the interaction chain that happened to produce it.
References
- Google, Build with Gemini Omni 1.1 Flash
- Google AI for Developers, Generate and edit videos with Gemini Omni Flash
- Google AI for Developers, Interactions API
- Google AI for Developers, Zero data retention in the Gemini Developer API
- Google AI for Developers, Gemini Developer API pricing
- Google DeepMind, SynthID
- Coalition for Content Provenance and Authenticity, C2PA Technical Specification 2.4
- NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile
- Huang et al., VBench: Comprehensive Benchmark Suite for Video Generative Models
- He et al., VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation
- Google Gemini Apps Help, Verify AI-generated images, videos, and audio