The Full Breakdown - I Think We Have a Problem

ToolsGPT Image 2.5 ยท Nano Banana Pro ยท Seedance 2.5
Length7 seconds, one take, 1080ร—1920

Everything used in that video: the tools, the exact order I run them in, and every prompt.

Nothing here is theory. This is the same pipeline, step by step, with the prompts.

Try it here โ†’ Higgsfield AI

The whole video comes from two images and one video prompt. No editing and no cuts. The images do most of the work, so that's where I spend most of my time.

The final video. Turn the sound on.

The pipeline in one look

StepToolWhat it gives you
1GPT Image 2.5 SunburstThe first frame. The street, the guy at the car, and that old consumer-camera look.
2Nano Banana ProA character sheet for the person filming, who only shows up at the end.
3Seedance 2.5Turns both images into one 7-second handheld take with voice.

The more real your images, the more real your video.

Seedance copies the look of your first frame into every second of the clip. If the image looks glossy and AI-made, the video will too. If it looks like a real photo from a cheap camera, the video reads as real footage. So don't rush the references.


Step 1: The first frame

This image becomes frame 0.0 of the video. It sets the location, the guy at the car, the parked cars, the light and how far away the camera is. I make it vertical so it matches the final video.

ModelQualitySize
GPT Image 2.5 SunburstHigh1520ร—2688
A man pulling on the door handle of a black sedan across a quiet Manhattan street, framed by two parked cars in the foreground

The prompt:

Plain Text
A man trying to open a parked car on the opposite side of a quiet Lower Manhattan street. The black sedan is parallel-parked tightly against the opposite curb. The man is standing entirely ON THE SIDEWALK, in the narrow space between the parked car and the apartment buildings โ€” never in the roadway.

He is seen almost completely from behind. His back faces the camera, his body turned only slightly toward the car. He stands directly beside the driver's door and leans forward toward it. One hand is clearly wrapped around the exterior driver's door handle, actively pulling it, while his other hand rests against the upper door near the window for support. The door remains closed.

His head is angled downward toward the lock and handle. He is completely focused on the car and has not noticed the camera. His face is mostly hidden.

The camera is positioned far away on the opposite sidewalk, looking completely across the width of the street. The roadway is clearly visible between the camera and the man. He and the parked car remain relatively small in the middle distance.

Two nearby parked cars partially enter the bottom left and bottom right edges of the foreground, naturally framing the distant action without blocking it.

Behind the man is the opposite sidewalk and a row of old Lower Manhattan red-brick apartment buildings with narrow stoops, black iron railings, fire escapes, window air conditioners and mature leafy street trees. More parallel-parked cars continue along the curb in both directions.

The composition should immediately read as:
foreground parked cars โ†’ open roadway โ†’ opposite parked black sedan โ†’ man standing on the sidewalk side of that sedan โ†’ apartment buildings behind him.

Real-life early-2000s candid photography. Warm natural color temperature, slightly washed everyday colors, soft imperfect contrast and realistic uneven exposure. Bright areas are gently overexposed rather than perfectly recovered, with warm luminous skin and soft highlight clipping. Shadows remain natural and slightly murky rather than heavily graded.

Fine visible organic grain across the entire image, subtle digital noise in darker areas, gentle softness in facial and fabric detail, slightly imperfect consumer-camera sharpness and mild color noise. Natural skin texture without modern computational processing. The image should feel tactile, warm and imperfect.

Beautiful but completely unpolished โ€” like a striking accidental frame captured on an early-2000s consumer camera during a real moment, not a modern cinematic color grade, not glossy editorial photography and not hyper-clean digital imagery.

What matters here:


Step 2: The character sheet

The second guy is the one filming. You don't see him until the camera swings around at the end, so the video model needs a clear face to land on. A three-panel sheet gives it the face, the body and the back of the head.

ModelQualitySize
Nano Banana Pro1K1376ร—768
Character sheet of the same curly-haired man: face close-up, full body front, full body back

The prompt:

Plain Text
Create an ultra-photorealistic professional character reference sheet for one fictional male character.

The image must contain exactly three panels arranged horizontally on one clean neutral light-gray studio background.

CHARACTER:
A highly recognizable 24-year-old Southern European man with a tall, lean athletic build and strong natural charisma. Thick dark-brown curly hair with a distinctive loose curl pattern, striking light green eyes, thick straight dark eyebrows, warm olive skin, prominent cheekbones, a defined natural jawline, and a strong slightly Roman-shaped nose.

He has a subtle five-oโ€™clock shadow and naturally expressive eyes. His face has strong recognizable proportions and subtle natural asymmetry, with realistic pores, faint under-eye texture and small skin imperfections.

He should have the kind of face people easily remember after seeing it once โ€” charismatic, recognizable and attractive without looking like a perfect fashion model or AI-generated person.

WARDROBE:
Faded dark navy heavyweight T-shirt with a relaxed fit, straight washed charcoal jeans, simple dark sneakers. No jewelry, tattoos, glasses or hat.

PANEL 1 โ€” CLOSE-UP:
Tight head-and-shoulders portrait facing directly toward camera. Neutral relaxed expression. Eyes clearly visible. Curly hair, eyebrows, nose shape, ears, jaw and natural skin texture clearly readable.

PANEL 2 โ€” FULL BODY FRONT:
The exact same person standing naturally facing directly toward camera. Entire body visible from head to shoes. Relaxed neutral stance, arms resting naturally beside the body. Accurate realistic human proportions.

PANEL 3 โ€” FULL BODY BACK:
The exact same person photographed directly from behind. Entire body visible from head to shoes. Same distinctive curly hairstyle, clothing and body proportions.

IDENTITY CONSISTENCY:
All three panels depict exactly the same individual at the same age, with identical facial structure, skull shape, hairstyle, hair color, skin tone, body proportions and wardrobe.

LIGHTING:
Professional neutral reference photography. Soft even daylight-balanced studio lighting, approximately 5600K. Near-shadowless. Natural eye catchlights. No dramatic or cinematic lighting.

CAMERA:
Highly realistic professional photography. Natural perspective and realistic skin detail. No beauty filter, excessive skin smoothing, stylization or exaggerated depth of field.

COMPOSITION:
Clean technical character turnaround sheet. Equal visual spacing between panels. Consistent scale between both full-body views. Plain seamless light-gray background. No text, labels, graphics, borders or decorative elements.

The final image should look like photographs of a real person taken during a professional casting and VFX reference session, not concept art.

What matters here:


Step 3: The video

Both images go into Seedance as references. I write the prompt like a shot list: first what each image is for, then the action split into timed moments, then camera, physics, hard rules and audio.

ModelQualityBitrateSize
Seedance 2.51080pHigh1080ร—1920

References:

The prompt:

Plain Text
=== REFERENCE MAP ===

<<<image_1>>> (street scene) โ†’
This is the EXACT first frame of the clip.

It defines:
the location,
the suspicious man,
the black parked sedan,
all parked cars,
the architecture,
the street,
the trees,
the lighting,
the framing,
the camera position,
the distance across the street,
and the photographic character.

Frame 0.0 is exactly <<<image_1>>>.

The take grows directly out of this image.

<<<image_2>>> (camera operator) โ†’
Identity reference ONLY for the unseen person filming.

Use only the man's identity:
his face,
facial proportions,
hair,
eyes,
eyebrows,
nose,
jawline,
skin tone,
age and natural facial features.

Ignore the studio background, character-sheet layout, multiple poses and studio lighting.

The man from <<<image_2>>> is the person filming from the very beginning.

He is BEHIND THE VIEWPOINT and therefore invisible at the beginning.

He becomes visible only near the end when the viewpoint physically swings around toward his face.

The man beside the black sedan in <<<image_1>>> and the man from <<<image_2>>> are two completely different people.


=== CHARACTER ELEMENT ===

<<<SUSPICIOUS_MAN>>> โ†’
The exact man already standing beside the black sedan in <<<image_1>>>.

Preserve his appearance and clothing.

<<<OPERATOR>>> โ†’
The exact man from <<<image_2>>>.

He is the unseen person filming.

He is never visible before the final camera turn.


=== CAMERA / VIEWPOINT RULE ===

The camera is a VIEWPOINT ONLY.

The recording device itself is NEVER visible in the image.

No phone.
No camera body.
No lens.
No camera rig.
No hands holding a camera.
No reflection of the camera.
No filming equipment visible anywhere in the frame.

The viewer sees THROUGH the recording device.

The physical recording device always remains outside the visible image.


=== PART 1 โ€” SHOT BREAKDOWN
7 seconds, one continuous handheld take


=== MOMENT 1 (0.0โ€“1.8s) โ€” LOCKED DOOR HANDLE ===

Frame 0.0 is exactly <<<image_1>>>:
same location,
same man,
same parked cars,
same architecture,
same lighting,
same framing,
same crop,
same distance,
same photographic character.

The take grows directly out of this frame with no reframing or settle-in.

<<<SUSPICIOUS_MAN>>> remains beside the black parked sedan across the street, completely unaware that he is being filmed.

He repeatedly tries the locked driver's door handle with his right hand.

He gives it TWO short, believable pulls โ€” not exaggerated, just quick frustrated tugs from the wrist and forearm.

PULL 1:
his fingers grip the handle,
his wrist pulls,
his elbow bends slightly,
his shoulder reacts,
the handle reaches its physical limit,
the locked door stays shut,
his hand relaxes slightly.

A tiny natural pause.

PULL 2:
he grips and pulls again,
slightly more frustrated,
his forearm tightens,
his shoulder and upper torso respond,
his weight shifts subtly through his legs.

The door stays completely shut and rigid.

The car does not move or deform.

His hand remains physically connected to the door handle during each pull.

His body reacts naturally to each pull:
slight movement through the shoulder and upper torso,
small movement through the elbow,
a subtle shift of weight through his legs.

His attention stays completely focused on the door.

At the same time, the unseen <<<OPERATOR>>> quietly says from behind the viewpoint:

"What is this guy doing?"

The voice sounds spontaneous and slightly confused.

He speaks quietly under his breath rather than performing for an audience.

The voice is recorded naturally by the same device filming the scene.

Natural city ambience continues underneath:
distant traffic,
subtle street noise,
faint footsteps,
quiet residential ambience.

No music.


=== MOMENT 2 (1.8โ€“2.3s) โ€” HE HEARS SOMETHING ===

Immediately after the SECOND pull, <<<SUSPICIOUS_MAN>>> suddenly freezes.

A brief natural pause.

His hand remains on the door handle.

His fingers are still physically touching it.

His body stops before his head moves.

It feels as though he has just heard the person filming.

No instant full-body turn.

No exaggerated reaction.


=== MOMENT 3 (2.3โ€“3.4s) โ€” HE FINDS THE CAMERA ===

<<<SUSPICIOUS_MAN>>> quickly turns his head over his shoulder toward the viewpoint.

His eyes search across the street.

His eyes find the viewpoint.

A fraction of a second later, his shoulders and upper torso rotate partially toward it while one hand remains near the car door.

Natural head-first reaction:

eyes and head first,
shoulders second,
upper torso afterward.

The reaction is sharp and instinctive, as if he has suddenly realized somebody is secretly recording him.

He does not smile.

He does not pose.

He does not gesture.

He does not overreact.

His expression is restrained and believable:
alert,
suspicious,
slightly startled.

For a brief uncomfortable moment he looks directly toward the viewpoint from across the street.

The unseen person filming instinctively reacts with a small abrupt handheld jerk backward and slightly downward, followed by an imperfect correction.


=== MOMENT 4 (3.4โ€“4.3s) โ€” HE STARTS COMING OVER ===

<<<SUSPICIOUS_MAN>>> releases the door handle.

His hand visibly leaves the handle.

He turns fully away from the car.

He immediately begins walking quickly toward the person filming.

The movement begins from his EXACT existing position beside the sedan.

He takes quick, deliberate steps away from the car and toward the curb.

He keeps looking directly toward the viewpoint.

He is NOT running.

His pace is fast, purposeful and slightly confrontational.

Natural acceleration from standing.

Natural weight transfer.

Natural leg movement.

Natural arm swing.

His distance decreases ONLY because he physically walks toward the viewpoint.

No teleportation.

No sudden scale change.

No artificial camera movement toward him.

The person filming becomes visibly more nervous through the handheld movement.

The tiny hand tremor becomes slightly stronger.


=== MOMENT 5 (4.3โ€“5.1s) โ€” VIEWPOINT SWINGS TO OPERATOR ===

The frightened <<<OPERATOR>>> quickly swings the VIEWPOINT approximately 180 degrees toward himself.

IMPORTANT:

The recording device itself remains completely invisible.

We see only the resulting changing VIEW.

No camera body enters the frame.
No lens enters the frame.
No phone enters the frame.
No filming equipment enters the frame.

<<<OPERATOR>>> does NOT turn his body around.

His feet and torso remain facing the approaching <<<SUSPICIOUS_MAN>>>.

He does NOT turn his back to him.

Only the arm holding the unseen recording device moves.

The viewpoint physically sweeps through the surrounding environment.

For a brief moment we see a fast, messy sweep of:
roadway,
parked cars,
sidewalk,
trees,
brick buildings.

The movement is fast, frightened and imperfect.

Strong natural directional motion blur during the fastest part.

Temporary horizon tilt.

Subtle rolling-shutter distortion.

Brief autofocus uncertainty.

Small automatic exposure reaction.

Irregular human acceleration and deceleration.

This is NOT a camera switch.

This is NOT a cut.

This is NOT a digital flip.

This is NOT a transition.

It is one continuous physical handheld recording.


=== MOMENT 6 (5.1โ€“5.4s) โ€” OPERATOR REVEAL ===

The viewpoint lands on <<<OPERATOR>>>.

He is exactly the man from <<<image_2>>>.

Same recognizable face,
same facial proportions,
same dark curly hair,
same eyes,
same eyebrows,
same nose,
same jawline,
same skin tone,
same natural facial features.

Do NOT reproduce the studio background or character-sheet composition from <<<image_2>>>.

He appears naturally inside the SAME real Manhattan street environment.

The original lighting, exposure, texture and photographic character continue naturally.

The viewpoint lands slightly imperfectly on his face.

He is slightly off-center for a fraction of a second.

Tiny overshoot.

One instinctive correction.

Autofocus briefly catches and locks onto his face.

The framing remains handheld and imperfect.

Because <<<OPERATOR>>>'s BODY NEVER TURNED AROUND, <<<SUSPICIOUS_MAN>>> is physically in front of him and therefore BEHIND THE VIEWPOINT.

<<<SUSPICIOUS_MAN>>> must NOT appear behind <<<OPERATOR>>>.


=== MOMENT 7 (5.4โ€“7.0s) โ€” "I THINK WE HAVE A PROBLEM" ===

<<<OPERATOR>>> looks genuinely frightened.

His eyes are widened.

His eyebrows are naturally raised.

His lips are already slightly parted.

His breathing is faster.

There is subtle tension through his jaw and face.

He looks like someone who has just realized that the stranger he was secretly filming is now coming directly toward him.

His eyes briefly flick past the side of the viewpoint toward the approaching man.

Then immediately back into the lens.

He quickly says:

"I think we have a problem."

The entire sentence comes out FAST, almost in one breath.

Slightly breathless.

Quiet but urgent.

He is frightened and does not want the approaching man to hear him clearly.

No dramatic pause between:
"I think"
"we have"
"a problem."

It is one quick natural sentence.

His hand trembles subtly while he speaks, creating tiny irregular movement in the framing.

His faster breathing creates tiny natural movement through his shoulders and the viewpoint.

Natural blinking.

Natural eye movement.

Natural lip and jaw movement.

No theatrical fear.

No screaming.

No exaggerated facial expression.

No influencer-style delivery.

Hard end immediately after the line.


=== CAMERA BEHAVIOR ===

The viewpoint stays physically on the opposite sidewalk throughout the entire shot.

Raw handheld movement from a real person filming unsupported.

Constant tiny hand tremor.

Subtle breathing movement.

Slight irregular horizontal and vertical drift.

Occasional tiny framing corrections.

When <<<SUSPICIOUS_MAN>>> suddenly notices the filming, the person filming instinctively reacts:
a small abrupt viewpoint jerk backward and slightly downward,
followed by an imperfect correction.

When <<<SUSPICIOUS_MAN>>> starts approaching, the handheld movement becomes slightly more nervous.

During the 180-degree physical turn:
strong natural directional motion blur,
temporary horizon tilt,
subtle rolling shutter,
brief autofocus uncertainty,
small exposure adaptation,
irregular human acceleration and deceleration.

After landing on <<<OPERATOR>>>, the framing remains subtly unstable.

No stabilization.
No gimbal.
No smooth floating motion.
No artificial push-in.
No zoom.
No camera movement toward <<<SUSPICIOUS_MAN>>>.
No cuts.
No transitions.


=== PHYSICAL REALISM ===

Natural human timing and weight transfer.

The hand remains physically connected to the door handle while pulling.

Exactly TWO visible door-handle pulls.

The locked door does not open.

The black sedan remains completely stationary.

Realistic wrist, elbow and shoulder mechanics.

Natural head-first reaction when <<<SUSPICIOUS_MAN>>> notices the filming:
eyes and head first,
shoulders afterward.

He starts approaching from his real existing position.

He becomes closer only by physically walking.

<<<OPERATOR>>> remains physically in the same location.

<<<OPERATOR>>> does not rotate his body when turning the viewpoint toward himself.

Only his arm and the unseen recording device move.

Natural blinking.

Natural eye movement.

Natural frightened facial behavior.

Clothing reacts subtly to body movement.

All parked vehicles remain stationary.

Background pedestrians continue naturally without reacting.


=== HARD RULES ===

Total duration is exactly 7 seconds.

One single continuous handheld take.

Frame 0.0 is EXACTLY <<<image_1>>>.

Same camera position.
Same framing.
Same crop.
Same man.
Same black sedan.
Same parked cars.
Same architecture.
Same lighting.

The take grows directly out of <<<image_1>>>.

The recording device is NEVER visible.

NO visible camera.
NO visible lens.
NO visible phone.
NO visible camera body.
NO visible filming rig.
NO visible hand holding a camera.
NO camera reflection.

The viewer always sees THROUGH the recording device.

<<<SUSPICIOUS_MAN>>> and <<<OPERATOR>>> are TWO DIFFERENT PEOPLE.

Never swap them.

Never morph one into the other.

Never replace <<<SUSPICIOUS_MAN>>> with <<<OPERATOR>>>.

<<<OPERATOR>>> is invisible before the viewpoint turns toward him.

<<<image_2>>> defines ONLY <<<OPERATOR>>>'s identity.

Exactly TWO short door-handle pulls happen before <<<SUSPICIOUS_MAN>>> freezes.

Do not skip the pulls.

Do not simplify them into touching the handle.

The hand physically pulls the handle twice.

The door remains locked and closed.

<<<SUSPICIOUS_MAN>>> does not notice the filming until AFTER the second pull and AFTER "What is this guy doing?"

He turns his head first.

Then his shoulders.

Then he begins approaching.

No teleportation.

No sudden close-up.

The camera never crosses the street.

The camera never moves next to the black sedan.

<<<OPERATOR>>> never turns his body around.

Only the unseen recording device swings around toward his face.

After the turn, <<<SUSPICIOUS_MAN>>> is behind the viewpoint and is NOT visible behind <<<OPERATOR>>>.

No new people.

No new foreground objects.

No new vehicles.

No unexplained objects.

No cuts.

No scene changes.

No transitions.

No time skips.

No stabilization.

No zoom.

No artificial camera movement.

Preserve the exact real-world lighting, shadows, reflections, film grain, texture and exposure character of <<<image_1>>> throughout the entire take.


=== AUDIO ===

Everything sounds like raw sound captured by the same recording device.

Natural quiet city ambience:
distant traffic,
faint tire noise,
subtle footsteps,
light residential street ambience.

At the beginning, <<<OPERATOR>>> quietly says:

"What is this guy doing?"

Spontaneous.
Slightly confused.
Under his breath.

At the end, the SAME voice says:

"I think we have a problem."

Fast.
Frightened.
Slightly breathless.
Quiet but urgent.
Almost one breath.

No music.

No cinematic sound design.

No whoosh during the camera turn.

No artificial impact sounds.

No narrator.

No additional dialogue.

The shot, second by second

TimeMomentWhat happens
0.0โ€“1.8sLocked doorTwo short pulls on the handle. From behind the camera: "What is this guy doing?"
1.8โ€“2.3sHe hears somethingHe freezes with his hand still on the handle.
2.3โ€“3.4sHe finds the cameraHead first, then shoulders. The camera flinches back and down.
3.4โ€“4.3sHe comes overLets go of the car and walks fast toward the camera. He doesn't run.
4.3โ€“5.1sThe swingThe camera spins 180ยฐ with motion blur, a tilted horizon and focus hunting.
5.1โ€“5.4sThe revealIt lands a bit off-center on the operator and autofocus locks on his face.
5.4โ€“7.0sThe line"I think we have a problem." Fast, quiet, almost one breath. Hard end.

Why the video prompt works


Recap

  1. Make the first frame in the final aspect ratio. Put as much work into how the photo looks as into what's in it.
  2. Make a character sheet for anyone who appears later. Use a memorable face, real skin and flat light.
  3. Write the video prompt as a shot list: references, timed moments, camera, physics, hard rules, audio.
  4. Repeat the rules that matter in more than one place. The model follows what it reads more than once.

And once more: the more real your references are, the more real your video will be.


Follow me on Instagram to see more content like this.

@morfl_ow