Gemini Omni in Google Vids: Generate and Edit One Clip Step by Step — Agent Lab Journal
Agent Lab Journal
Guides
Glossary
Practical lab · Google Vids · Generative video
Gemini Omni in Google Vids: Generate and Edit One Clip Step by Step
Level: intermediate
Reading and lab time: 60 minutes
Outcome: a test clip, prompt log, and stability scorecard
The useful question is not whether a generative video editor can produce one attractive shot. It is whether the editor preserves the same shot after several independent instructions. In this lab, you will generate one controlled clip, replace its background, change its lighting, add an atmospheric effect, and record exactly what changes unintentionally at every step.
Contents
The question this lab answers
Confirm the available editing mode
Concrete case
Define acceptance criteria
Prepare the workspace and log
Generate the baseline clip
Apply three sequential edits
Create a direct control version
Verify and score the clips
Failure cases and diagnosis
Limitations
Final report template
The question this lab answers
Gemini Omni in Google Vids can be used through natural-language editing controls when the feature is available to the account. The interface, limits, supported inputs, and exact labels may change. Record what your account actually displays instead of assuming that every Google Vids account exposes the same workflow.
The test focuses on five properties:
Composition stability: the subject’s position, shot size, perspective, camera angle, and surrounding negative space remain stable.
Character consistency: the face, age, hair, clothing, body proportions, and distinctive objects do not change between versions.
Edit locality: a request to replace the background does not also replace the subject, motion, or camera.
Reference adherence: the generated subject retains the visible attributes defined by the supplied reference image.
Cumulative stability: repeated edits do not produce increasing generation drift away from the baseline.
What this article does not claim
No scores or product conclusions are supplied in advance. The blank score fields are intentional. Your exported clips and recorded observations determine the result.
Confirm the available editing mode
Open Google Vids with the account you intend to test. Find the AI video creation panel and check whether it offers generation, animation, editing, or another relevant operation. Write down the exact wording visible in the interface.
Before spending a generation, record:
the product and feature name shown in the interface;
the account type or workspace context, without recording credentials;
the test date and interface language;
whether an existing uploaded video can be edited;
whether only AI-generated clips from the current history can be edited;
whether image references can be attached;
the available aspect ratios and output settings;
whether the current session keeps previous generations;
whether a fixed seed or other reproducibility control is exposed.
If the editor does not accept an uploaded clip, use an AI-generated clip from the same session if that workflow is available. If no edit operation exists, stop the sequential test and record “editing mode unavailable.” Recreating the scene from text alone is a different experiment because it does not test preservation of an existing clip.
Preserve every version immediately
Generation history may be tied to the current session. Insert each accepted clip into the project and export or otherwise preserve it before closing the tab or starting the next edit.
Concrete case: a courier in a studio
Use a scene in which unintended changes are easy to detect: one person, one prop, one short action, and a static camera. Avoid crowds, readable labels, brands, elaborate choreography, or rapid camera movement.
A young courier in a dark navy windbreaker stands in the center of a bright neutral studio. The courier holds a small square cardboard box with a plain white label in the left hand, takes two calm steps toward the camera, stops, and raises the box slightly. The camera remains static in a full-body wide shot.
The scene contains seven observable anchors:
exactly one visible person;
a dark navy windbreaker;
a small square box;
the box remains in the left hand;
two steps followed by a small lift of the box;
the subject stays near the center of the frame;
the camera remains static.
Use a synthetic, licensed, or self-created reference that you are permitted to upload. Do not use a private customer asset, an identifiable stranger, confidential footage, or a celebrity. The reference should have a clear face, visible clothing, and a neutral background. Do not include text that must remain legible.
Define acceptance criteria before generating
A successful export is not automatically a successful edit. Set the conditions before seeing the output so that an attractive result does not hide unwanted changes.
Area
Condition
Evidence
Subject
Face, hair, proportions, and clothing remain recognizable
Matched frames from every version
Prop
The same square box remains in the left hand
Start, lift, and final frames
Action
The two-step movement and box lift remain present
Action timecodes
Camera
No new pan, zoom, cut, or angle change appears
Full playback and frame overlay
Requested edit
The named background, light, or effect is visibly applied
Before-and-after comparison
Unrequested edits
Untargeted elements remain substantially unchanged
Deviation log
Temporal quality
No new flicker, warping, disappearing objects, or discontinuities
Normal-speed and frame-level review
Use a four-point score for each criterion:
Score
Meaning
4
Preserved or executed accurately; no meaningful deviation is visible.
3
A small deviation is visible but does not prevent practical use.
2
A clear deviation requires another edit or a manual correction.
1
The instruction is only partly followed or an important element changes.
0
The requested edit fails or a critical element is lost.
Prepare the workspace and prompt log
Create a directory outside the Google Vids project. Its purpose is to preserve the version lineage: which clip was edited to create each later clip.
mkdir -p google-vids-omni-test/{reference,clips,frames,notes}
cd google-vids-omni-test
touch prompts.csv scores.csv notes/environment.txt
Use this file structure:
google-vids-omni-test/
├── reference/
│ └── 00-reference.png
├── clips/
│ ├── 01-baseline.mp4
│ ├── 02-background.mp4
│ ├── 03-light.mp4
│ ├── 04-effect.mp4
│ └── 05-control-direct.mp4
├── frames/
├── notes/
│ └── environment.txt
├── prompts.csv
└── scores.csv
Enter this header in prompts.csv:
step,file,parent,operation,prompt,reference,settings,attempt,result_notes
1,01-baseline.mp4,,generate,"...",00-reference.png,"fill in",1,""
2,02-background.mp4,01-baseline.mp4,edit_background,"...",00-reference.png,"same",1,""
3,03-light.mp4,02-background.mp4,edit_light,"...",00-reference.png,"same",1,""
4,04-effect.mp4,03-light.mp4,edit_effect,"...",00-reference.png,"same",1,""
5,05-control-direct.mp4,01-baseline.mp4,edit_all,"...",00-reference.png,"same",1,""
Copy every prompt exactly, including punctuation, negative constraints, and line order. If you retry a step, add another row instead of replacing the failed attempt.
Record the environment in a neutral manifest:
test_date: YYYY-MM-DD
product_label: fill_in
feature_label: fill_in
account_context: personal_or_workspace
interface_language: fill_in
source_video_upload: available_or_unavailable
reference_images: available_or_unavailable
aspect_ratio: fill_in
requested_duration: fill_in
actual_duration: fill_in
export_resolution: fill_in
audio: enabled_or_disabled
seed_control: available_or_unavailable
generation_limit_shown: fill_in_or_not_shown
notes:
Generate the baseline clip
Start a new AI video generation. Attach 00-reference.png if reference inputs are supported. Request a short landscape clip and keep the action simple.
Create one realistic short video clip in landscape format.
One young courier wearing a dark navy windbreaker stands in the center
of a bright, neutral studio. The courier holds a small square cardboard
box with a plain white label in the left hand.
The courier takes exactly two calm steps toward the camera, stops,
and raises the box slightly. Use a full-body wide shot. Keep the camera
completely static, with no zoom, pan, cut, or change of angle.
Use soft neutral lighting and one continuous shot. Preserve the visible
appearance and clothing from the attached reference image.
Do not add other people, logos, readable text, extra props, camera
movement, or scene transitions.
If the interface separates the main request from exclusions, place the final paragraph in the exclusion field without changing its meaning. If duration is controlled outside the prompt, select the available short-clip duration and record the actual value.
Baseline gate
Do not continue with a baseline that is unsuitable for comparison. It should satisfy at least five of the seven scene anchors and must not contain a critical facial, hand, or box deformation.
Exactly one person is visible.
The subject remains near the center.
The windbreaker is dark navy.
The box is held in the left hand.
The action includes the steps and box lift.
The camera remains static.
The face, hands, and box are usable for comparison.
Set a retry limit before generating, such as three baseline attempts. Log every attempt. Choose the baseline by compliance with the test anchors, not by cinematic appeal. Save the accepted version as 01-baseline.mp4.
Apply three sequential edits
The main branch must follow this exact chain:
01-baseline.mp4
↓
02-background.mp4
↓
03-light.mp4
↓
04-effect.mp4
Each step starts from the immediately preceding version. Do not return to the baseline inside this branch.
Step 1: replace only the background
Select the baseline clip as the edit source and submit:
Replace only the background.
Change the bright neutral studio into a quiet city street at night
after rain. Add wet asphalt and softly blurred shop windows in the
distance.
Do not change the courier, face, hair, body proportions, dark navy
windbreaker, box, left hand, action, timing, subject position, shot
size, perspective, focal length, or static camera.
Keep the existing neutral lighting on the courier. Do not add people,
vehicles, logos, readable text, or new foreground objects.
Save the result as 02-background.mp4. Before continuing, record whether the subject’s size, first-frame pose, walking path, clothing, face, hand, or box changed.
Step 2: change only the lighting
Use 02-background.mp4 as the source:
Change only the lighting on the subject and the scene.
Add a cool blue rim light from camera right and a soft warm shop-window
light from camera left. Keep the face naturally exposed and retain
visible detail and the original dark navy color of the windbreaker.
Do not change the rainy night street, wet asphalt, shop windows,
courier, face, hair, clothing, box, action, timing, object positions,
camera angle, shot size, duration, or static camera.
Save the result as 03-light.mp4. Check whether the editor created directional light or merely placed a uniform color cast over the entire frame. Record flicker, clipped highlights, unnatural skin color, and changes to the background.
Step 3: add only an atmospheric effect
Use 03-light.mp4 as the source:
Add only a subtle atmospheric effect.
Add sparse, fine raindrops in the air and a very light layer of steam
close to the wet asphalt. Keep the effect realistic and restrained.
It must not cover the courier's face, hands, windbreaker, or box.
Do not change the background, lighting directions, subject appearance,
clothing, box, action, composition, camera angle, shot size, duration,
or static camera. Do not add lightning, heavy fog, strong wind, people,
vehicles, text, or cuts.
Save the result as 04-effect.mp4. Look for new temporal artifacts: rain attached to the subject, steam crossing foreground objects incorrectly, particles appearing for one frame, edge warping, or unstable brightness.
Create a direct control version
The sequential branch cannot reveal whether a failure was caused by cumulative editing or ordinary generation variability. Create a control branch directly from 01-baseline.mp4 and request all three edits in one operation.
Preserve the original courier, face, hair, proportions, dark navy
windbreaker, square box in the left hand, two-step action, box lift,
subject position, shot size, perspective, duration, and static camera.
Replace the studio with a quiet city street at night after rain,
including wet asphalt and softly blurred shop windows in the distance.
Add a cool blue rim light from camera right and a soft warm shop-window
light from camera left. Keep the face naturally exposed.
Add sparse fine raindrops and very light steam close to the asphalt.
The effects must not cover the face, hands, windbreaker, or box.
Do not add people, vehicles, logos, readable text, lightning, heavy
fog, camera movement, cuts, or extra foreground objects.
Save this version as 05-control-direct.mp4.
Branch
Parent
Transformations
Question answered
Sequential
Previous edited version
Three separate edits
Does drift accumulate across turns?
Direct control
Baseline
One combined edit
Does one composite instruction preserve the baseline better?
Verify and score the clips
Do not rely only on sequential playback. Compare the same moments in every clip: the first frame, the beginning of the first step, the stop, the box lift, and the final frame.
1. Check technical comparability
Record the actual duration, resolution, frame rate if available, and aspect ratio of every export. A changed duration can shift events and make visual comparison misleading.
If FFmpeg is installed, inspect each file without modifying it:
for file in clips/*.mp4; do
echo "$file"
ffprobe -v error \
-show_entries stream=width,height,r_frame_rate \
-show_entries format=duration \
-of default=noprint_wrappers=1 "$file"
done
2. Extract matched frames
Choose time points that exist in every clip. The following example extracts frames at one, three, and five seconds; adjust the values to the actual action timing:
mkdir -p frames
for file in clips/*.mp4; do
name=$(basename "$file" .mp4)
ffmpeg -i "$file" \
-vf "select='eq(t,1)+eq(t,3)+eq(t,5)'" \
-vsync vfr "frames/${name}-%02d.png"
done
If FFmpeg is unavailable, use the Google Vids timeline or another player that can pause consistently. Record timecodes and capture screenshots manually.
3. Fill the scorecard
Criterion
02 background
03 light
04 effect
05 control
Face and apparent age
Hair and body proportions
Clothing and color
Box and left-hand continuity
Reference adherence
Subject position
Angle and shot size
Action and timing
Camera stability
Accuracy of requested edit
Absence of unrequested changes
Flicker and deformation control
There are 12 criteria, each worth up to four points. The maximum score for one version is 48:
overall_stability_percent = points_received / 48 × 100
Also calculate three narrower measurements:
Character stability: average of face, hair and proportions, clothing, box continuity, and reference adherence.
Scene stability: average of subject position, angle and shot size, action timing, and camera stability.
Edit precision: average of requested-edit accuracy and absence of unrequested changes.
4. Record deviations with timecodes
A numeric score is not enough for diagnosis. Add one row for every visible deviation:
file,timecode,area,severity,expected,observed
03-light.mp4,00:00:03.200,face,fill_in,"same face","fill in"
04-effect.mp4,00:00:05.000,box,fill_in,"box in left hand","fill in"
5. Interpret the branches cautiously
If scores decline from version 02 to 04, the run shows possible cumulative drift.
If version 04 is weaker than version 05, the direct composite edit preserved this baseline better in this run.
If version 05 preserves the subject but misses one requested effect, separate edits may offer better command precision.
If both branches fail differently, repeat the test before preferring either strategy.
If the result looks attractive but changes the face, action, or box, it did not fully satisfy the editing task.
Use narrow conclusions
Write “In this scene and run, the sequential branch scored…” rather than “Gemini Omni always preserves” or “Gemini Omni cannot preserve.” One scene cannot establish universal product behavior.
Failure cases and diagnosis
The editor rebuilds the entire scene
Symptoms: the first-frame pose, subject scale, walking path, facial structure, or box shape changes after a local request.
Diagnosis: confirm that the previous video was supplied as the edit source. If only the text was reused, the operation may have been a new generation rather than an edit.
Next attempt: place the requested change first, shorten secondary prose, and list the protected elements in a separate “Do not change” block.
Changing the background also changes subject lighting
Symptoms: the face becomes blue, the windbreaker changes hue, or new shadows appear during the background step.
Diagnosis: score the background replacement and the unrequested lighting change separately. A plausible night scene is not evidence of a local edit.
Next attempt: explicitly preserve the existing light on the subject and delay the intended lighting change until the next step.
“Change only the lighting” becomes a global color filter
Symptoms: skin, clothing, box, and background receive the same blue tint without directional shadows or highlights.
Next attempt: name the light sources, their directions, their relative softness, and the areas whose local colors must remain natural.
The atmospheric effect covers protected details
Symptoms: rain or steam obscures the face, fingers, clothing edges, or box.
Next attempt: reduce density, constrain steam to the ground plane, and explicitly prohibit particles over the protected regions.
The character changes only during motion
Symptoms: the first and last frames look correct, but the face, fingers, or box deform during the steps or lift.
Diagnosis: inspect the transition around the action frame by frame. Start and end screenshots alone cannot detect short-lived failures.
The instruction is ignored
Symptoms: the previous background, lighting, or effect remains substantially unchanged.
Next attempt: verify that an edit operation is active, move the requested change to the first sentence, and remove decorative language that competes with the instruction.
The clip duration or camera changes
Symptoms: the action is faster, slower, cropped, zoomed, or interrupted by a cut.
Diagnosis: compare actual durations and matched event timecodes, not only the beginning of the files.
Next attempt: protect duration, action timing, shot size, focal perspective, and static camera as separate constraints.
The run cannot be reproduced
Possible causes: no seed control, a changed model, lost session history, a modified reference, or undocumented settings.
Response: preserve source files, prompts, dates, feature labels, settings, and exports. Run the same branch more than once when the generation budget permits, then report the range rather than presenting one score as deterministic.
Editing is restricted for the current account or location
Symptoms: the edit tab is missing, uploading a source video is unavailable, or only clips from generation history can be selected.
Response: do not substitute another workflow without documenting it. Record the exact boundary and either test an eligible generated clip or classify the lab as blocked at the availability check.
Limitations
One courier scene does not represent every visual style, duration, camera movement, or number of characters.
Without a fixed seed, command effects cannot be completely separated from generation randomness.
A single run can reveal a failure but cannot estimate how frequently it occurs.
Reference adherence is partly subjective unless several independent reviewers use the same rubric.
Export compression can resemble generation flicker or remove small details.
Matched timestamps are imperfect when an edit changes clip duration or action speed.
Feature names, eligibility, quotas, and supported inputs may change after publication.
If every step regenerates from text instead of editing the previous clip, the experiment measures prompt repeatability rather than state preservation.
The score weights are designed for this lab. A commercial production may treat face, product shape, text, or brand color as automatic-failure criteria.
Visual comparison must remain about visible attributes; it should not be used to infer identity or other sensitive personal information.
Final report template
Complete this report only after preserving all clips and filling the scorecard:
TEST ID:
DATE:
PRODUCT LABEL:
FEATURE LABEL:
ACCOUNT CONTEXT:
EDIT SOURCE TYPE:
REFERENCE USED:
SEED CONTROL:
ASPECT RATIO:
ACTUAL EXPORT SETTINGS:
BASELINE ATTEMPTS:
ACCEPTED BASELINE FILE:
SEQUENTIAL BRANCH
Character stability: __ / 4
Scene stability: __ / 4
Edit precision: __ / 4
Overall score: __ / 48 (__%)
DIRECT CONTROL BRANCH
Character stability: __ / 4
Scene stability: __ / 4
Edit precision: __ / 4
Overall score: __ / 48 (__%)
First visible deviation:
Timecode:
Unrequested changes:
Critical artifacts:
Requested edits that were missed:
Branch that performed better in this run:
Evidence supporting that conclusion:
What must be repeated:
Test limitations:
Completion checklist
The reference and accepted baseline are preserved.
All three sequential versions are preserved.
The direct control version is preserved.
Every prompt and retry is recorded verbatim.
Environment and export settings are recorded.
Matched frames or timecodes have been reviewed.
The scorecard contains no guessed values.
The conclusion is limited to the observed scene and run.
The lab is complete even if the final video is unusable. A failed clip is still a valid experimental result when the prompt log and version chain reveal exactly which edit introduced the failure.
Continue with the journal’s practical guides, or review the evaluation and generative-video terms in the Agent Lab Journal glossary.
© Agent Lab Journal
Top comments (0)