The Full Breakdown - Miami Convertible

ToolsSoul 2.0 ยท GPT Image 2.5 ยท Seedance 2.5
Length5 seconds, one take, 1920ร—1080

Everything used in that video: the tools, the exact order I run them in, and every prompt.

Nothing here is theory. This is the same pipeline, step by step, with the prompts.

Try it here โ†’ Higgsfield AI

Two images and one video prompt. A first frame that looks like a real early-2000s photo, a character sheet for the second person, and a video model that brings both together.

The final video. Turn the sound on.

The pipeline in one look

StepToolWhat it gives you
1Soul 2.0The first frame. Miami street, the girl, the convertible, and that harsh on-camera flash look.
2GPT Image 2.5A three-panel character sheet of the guy who walks in, built from face photos.
3Seedance 2.5Turns the frame into a talking shot and brings the guy in, with sound.

The more real your images, the more real your video.

The video model copies the look of your first frame into every second of the clip. A glossy, AI-looking frame gives you an AI-looking video. A frame that looks like a real photo from a cheap camera gives you footage that feels real.


Step 1: The first frame

This image becomes the first frame of the video. It sets the street, the girl, the car, the light and that flash-photo look for the whole clip.

ModelSize
Soul 2.02048ร—1152
A young woman in a white tank top leaning on a vintage silver convertible on a palm-lined Miami Beach street, lit by direct flash

The prompt:

Plain Text
A stylish young woman standing casually beside a vintage silver convertible on a quiet palm-lined side street in Miami Beach โ€” positioned slightly off-center in the near foreground, one hand resting loosely on the open car door, the other hanging naturally by her side, her body turned slightly away while her face looks back toward the camera with a relaxed, effortless expression. She wears a simple fitted white tank top, loose low-rise cream linen trousers and minimal jewelry, her hair slightly messy from the warm coastal breeze.

Setting: a dreamy residential street in Miami Beach lined with tall palm trees, pastel Art Deco apartment buildings and lush tropical greenery, a few parked vintage cars along the curb, with a narrow glimpse of pale blue ocean visible at the far end of the street. Warm late-afternoon sunlight falls between the buildings, creating long soft shadows and patches of golden light across the pavement. The environment feels quiet, expensive and lived-in rather than touristy or staged.

Shot on an early-2000s point-and-shoot digicam with harsh direct on-camera fill flash that strongly illuminates the woman, her skin, the white tank top and the metallic edge of the convertible, creating sharp specular highlights and slightly flattened facial features against the softer ambient background. The Miami street behind her remains visible but hazy, warm and slightly underexposed relative to the flash.

Warm, slightly saturated CCD tones, sun-faded pastel colors, luminous warm skin, strong bloom and halation around bright highlights and sunlit edges, subtle chromatic imperfections, minor digital noise, slightly crushed soft shadows and imperfect consumer-camera sharpness. A faint warm haze hangs in the distance between the palms and buildings.

Candid amateur photograph, early-2000s Miami summer, Y2K fashion editorial, effortless luxury, unpolished but stylized, intimate and accidentally cinematic. 4:5.

What matters here:


Step 2: The character sheet

The guy who walks into the shot needs a consistent face. A three-panel sheet gives the video model his face, his outfit from the front and from behind.

ModelSize
GPT Image 2.52688ร—1520

I'm not sharing my character sheet here.

It's built from my own face photos. Attach photos of yourself or of the person you want, and the prompt below does the rest.

The prompt:

Plain Text
Generate a completely new, original three-panel studio photoshoot of a man. The attached photos are a face reference ONLY: take his exact facial features from them โ€” face shape, eyes, brows, nose, lips, jawline, ears, skin tone, hairline, hair colour and hair texture โ€” and nothing else. Ignore everything else in the references: their pose, head angle, expression, crop, framing, lighting, background, clothing, camera and image quality. Do not reproduce or edit the reference photos; create new images from scratch where this same person is photographed fresh.

Panel 1: portrait from mid-chest up, space above his head so his full hair is visible, head turned very slightly to a three-quarter angle, eyes to the camera, a calm, easy expression with a faint natural smile in the eyes.
Panel 2: full body from the front, relaxed natural stance, weight on one leg, left hand casually in his trouser pocket, right arm loose.
Panel 3: full body from directly behind, the same stance: left hand in his left trouser pocket, right arm loose.
His face must be the same person in all panels and match the references.

Outfit, identical in all panels: a relaxed blue linen shirt, top buttons open, sleeves loosely rolled, half-tucked; loose white linen trousers with natural wrinkles; brown suede loafers, no socks; a thin leather-strap watch on his left wrist. Casual old money look.

Lighting: soft, natural studio light from a large source at the front-left, so his face and body have gentle, real shadows on one side and soft falloff. The light rakes lightly across the skin and fabric, revealing their real texture. Light-grey seamless background with a soft natural gradient, no props, no text.

Realism, the most important part: everything must look like an unretouched photo of a real, living person taken on a real camera, never like a render, a 3D model, a painting or an AI image.
- Skin: alive and real, with visible pores, slight uneven tone, a little natural redness on the nose, cheeks and ears, small natural imperfections, fine facial hairs and a very light stubble shadow. A soft natural sheen on the forehead and nose, matte elsewhere. Warm light glowing slightly through the ear edges.
- Face and eyes: natural subtle asymmetry, relaxed real muscles, moist eyes with detailed irises and small natural catchlights, individual lashes and brow hairs, natural lip texture.
- Hair: real individual strands, flyaways and natural messiness, real volume and soft shine.
- Hands: correct anatomy with knuckle creases, visible veins and real nails.
- Materials: real linen with visible weave, slubs and soft creases where the body bends; real suede nap on the loafers; real metal and leather on the watch. Every material reacts to the light naturally, with correct reflections and soft shadows in the folds.
- Photo: natural depth, true-to-life colour, very fine natural grain, crisp focus on the face in every panel.

Avoid: smooth, waxy or plastic skin, airbrushing, beauty filters, clay-like or doll-like faces, over-sharpening, HDR, perfect symmetry, CGI or render look.

What matters here:


Step 3: The video

Both images go into the video model as references, with a prompt for the action and the lines.

ModelSize
Seedance 2.51920ร—1080

References:

The prompt:

Plain Text
=== REFERENCE MAP ===
<<<image_1>>> (Miami convertible) โ†’ the exact first frame of the clip. The young woman in the white tank top and cream linen trousers standing in front of the vintage cream-silver convertible with the top down, between the car body and its open driver's door, her right hand resting on the car body, her left hand in her trouser pocket. The pink vintage car at the left, the terracotta and pink apartment buildings, the palms, the old cars along the right curb, the pale ocean at the far end of the street and the low sun glowing from the upper right. This image defines her face and identity, her outfit, the street, the light, the colours and the camera look for the whole take.
<<<image_2>>> (man sheet) โ†’ the man: face, identity, hair, body type and wardrobe only. The grey studio background, the lighting and the poses in this image are NOT used. His skin, hair and fabrics must look fully real in the clip, never smooth or plastic as they may appear in the reference.

=== CHARACTERS ===
THE WOMAN โ€” from <<<image_1>>>: early 20s, dark brown wavy shoulder-length hair with loose strands, warm light-tan skin, dark brown eyes, full brows, full natural lips. White ribbed tank top, loose cream linen trousers, layered thin gold necklaces, gold hoop earrings, thin rings, a light-blue string bracelet and a thin gold bracelet on her right wrist, a gold bangle on her left wrist, small fine-line tattoos on her right forearm and near her collarbone.
THE MAN โ€” from <<<image_2>>>: mid-20s, light skin, tousled medium-length brown hair swept up and back, grey-green eyes, straight brows, defined jawline, slim athletic build. Relaxed blue linen button-up shirt with the top buttons open and sleeves loosely rolled, half-tucked; loose white linen trousers; brown suede loafers without socks; a thin leather-strap watch on his left wrist. He never holds or touches the camera. He enters from the right side of the frame at 2.9s.

PART 1 โ€” SHOT BREAKDOWN (5s, one continuous handheld take)

SHOT 1 โ€” one continuous handheld take, 5 seconds, no cuts. The camera is a small early-2000s digital camera held at chest height by an unseen person standing about 2.5 metres in front of her, on foot, reacting like a real person filming a friend.

MOMENT (0.0โ€“1.9s) โ€” Frame 0.0 is <<<image_1>>>, Line 1 starts immediately
EFFECT: handheld step-in (one real step forward, no zoom) + light sway
Frame 0.0 is exactly <<<image_1>>>: same framing, same light, same pose. The take grows out of the still with no settle-in and no morph. There is NO pause at the start: she begins speaking in the very first frames, her lips already moving into the first word, as if the camera caught her mid-thought.
She looks into the lens and says in a soft, low, slightly husky American English voice, at natural conversational speed: "This AI shot looks very cinematic, right?"
Intonation: relaxed, quietly amused and a little proud, as if showing a friend something beautiful.
- "This AI shot looks" โ€” easy, low, almost flat, a tiny playful weight on "AI".
- "very CINEMATIC" โ€” gentle stress on "cinematic", her brows lift slightly and a small smile starts in her eyes.
- "right?" โ€” a soft, light, rising question with a small head tilt and a small warm smile.
Her movement: on "cinematic" she pushes lightly off the car with her right hand, her hand lifting off the car body, and she straightens up, shifting her weight toward the camera. Her left hand slides out of her trouser pocket and hangs relaxed. On "right?" her shoulders lift in a small shrug. She blinks once naturally and a few strands of hair move in the breeze.
The person filming takes one slow, natural step forward during the line, so she grows slightly in the frame through real movement, not a zoom. The frame bobs slightly with the step and settles, never fully still.
Natural ambient sound: her voice about 2 metres from the camera mic, a quiet golden-hour street, palm fronds rustling, the distant hush of surf, the soft scuff of the operator's step on the asphalt.

MOMENT (1.9โ€“3.2s) โ€” Line 2, no gap
EFFECT: handheld sway, the camera holds its position
Straight after "right?", with only a tiny breath, she says: "Want to know how to make it?"
Intonation and emotion: a teasing, inviting question, a little lower and more confidential, like she is about to let the viewer in on a secret. Stress on "HOW", a light rising lift on "make it?".
Her movement: she leans her head and shoulders slightly toward the lens with a knowing half-smile, one brow slightly raised. On "how" her right hand comes up at waist height, palm half open, a small candid gesture, and drops back down. She does not point at anything and keeps her eyes on the lens.
Natural ambient sound: her voice, the breeze, and from about 2.8s quick relaxed footsteps on asphalt approaching from the right.

MOMENT (2.9โ€“3.7s) โ€” He walks into the frame from the right
EFFECT: handheld half-step back + small pan right (reactive, slightly late) + autofocus hunt-and-lock on him (about half a second)
While she is finishing "make it?", the man walks into the frame from the right side, coming from the street past the open car door, with a quick, relaxed, confident walk. He enters by walking in: first his shoulder and arm at the right edge, then his whole upper body. He stops beside her, on her left side (frame right), slightly closer to the camera, his body angled toward the lens, already looking into the lens.
The person filming reacts like a real operator: a quick half-step back and a small pan to the right, a beat late, to fit both of them in a medium two-shot, both visible from mid-thigh up. The autofocus hunts briefly and locks onto him in about half a second, silently.
She does not turn to him and does not look at him. She keeps looking into the lens with the same knowing half-smile, as if she expected him.
Natural ambient sound: his footsteps, the rustle of his linen shirt, the operator's step back, the breeze.

MOMENT (3.7โ€“4.9s) โ€” His line
EFFECT: light handheld sway
He looks straight into the lens and says: "Then I'll tell you."
Intonation and emotion: very charismatic, calm and confident, cool, with a playful spark. A warm, low, slightly raspy voice in natural English with a light natural accent, unhurried, like it is no effort at all for him. Stress on "I'LL", as a promise, with a short, sure drop in pitch on "tell you".
His movement: on "Then" he gives a small upward nod toward the lens with a relaxed half-smile; on "I'll" one brow lifts; on "tell you" his half-smile widens slightly. He does not point at himself or at her. He stays fully facing the camera.
She stays still in place beside him, looking into the lens with her small smile, breathing naturally.
Natural ambient sound: his voice about 1.5โ€“2 metres from the mic, a little clearer than hers, the breeze, a faint gull far over the water.

MOMENT (4.9โ€“5.0s) โ€” Hard out
EFFECT: hard cut (out-point, no trailing beat)
The clip ends right after the last syllable of "you", while both are still looking into the lens. No glance between them, no reaction, no laugh, no hold, no fade, no black frame.

PART 2 โ€” EFFECTS LIST
Everything is handheld or in-lens behaviour. No added effects exist in this clip.
Light handheld sway โ€” continuous, 0.0โ€“5.0s. Never locked, never still.
Handheld step-in โ€” 1x (0.0โ€“1.9s). One real slow step forward by the operator, no zoom.
Handheld half-step back + small late pan right โ€” 1x (2.9โ€“3.7s). The operator makes room for him.
Autofocus hunt-and-lock โ€” 1x, on the man, about half a second as he enters. Silent, in-lens.
Hard cut โ€” 1x (5.0s), the out-point.
Not used anywhere: other cuts, zooms, punch-ins, whips, smears, slow motion, speed ramps, stabilisation, colour grading, light changes, flashes, transitions, whoosh sounds, music, on-screen text.

PART 3 โ€” EFFECT DENSITY BY TIME
0.0โ€“1.9s = MEDIUM (instant speech, step-in, push off the car, hand out of the pocket, shrug)
1.9โ€“3.2s = LOW (lean in, hand gesture, footsteps approaching)
2.9โ€“3.7s = HIGH (his entrance, step back and late pan, focus hunt โ€” the peak)
3.7โ€“5.0s = LOW (his line, nod, brow lift, hard out)

PART 4 โ€” MOTION FLOW
Opening: no warm-up. She is already talking as the clip begins, and the camera steps in toward her.
Build: the question turns into a teasing invitation, and footsteps are heard from the right before we see him.
Resolution: he walks into the frame like he owns the answer, the camera scrambles back to fit him in, and he delivers "Then I'll tell you" straight into the lens. Cut on the last word.

=== HARD RULES ===
Total duration is exactly 5 seconds. One single continuous handheld take in real time. The only cut is the out-point at 5.0s.
Frame 0.0 is exactly <<<image_1>>>.
She starts speaking immediately at 0.0. No pause, no silent moment, no settling before her first word.
No dead air anywhere: her two lines flow with only a tiny breath between them, and his line starts as soon as he is in the frame.
Both characters keep their identities from their references for the whole clip: her face from <<<image_1>>>, his face, hair and outfit from <<<image_2>>>. No face changes.
Skin is never smooth on either of them: visible pores, fine texture, small natural unevenness and redness, fine facial hairs, a natural sheen where the light hits. His skin especially must look like real skin in warm sunlight, never waxy, plastic, clay-like or like the flat studio look in <<<image_2>>>. No retouching, no beauty filter.
Real hair: individual strands and flyaways, both of their hair moving in the breeze. Real linen with visible weave and wrinkles that move with their bodies.
Nobody stands frozen: both keep breathing, blinking and making small natural movements in every frame they are in.
Every movement has a visible beginning, middle and end. He enters the frame by walking into it from the right; he never appears suddenly.
They never look at each other. Both speak to the lens.
The camera moves only by real human steps and hand movements: no zoom, no smooth glide, no stabilised look.
Dialogue is spoken exactly once, exactly as written, in this order: her "This AI shot looks very cinematic, right?" / her "Want to know how to make it?" / his "Then I'll tell you." No added words, no repeats. Precise lip sync for both.
Neither of them points at themselves. The camera operator is never seen and never speaks. No other people appear.
No light in the scene ever changes. No flash firing, no white frame, no shutter sound, no click.

=== LOOK ===
Matches <<<image_1>>>: early-2000s consumer digital camera video. A hard, constant on-camera light gently fills their faces, her white tank top, his blue shirt and the cream-silver car, while the street glows in warm low late-afternoon sun from the upper right, catching the edges of their hair and shoulders. Balanced exposure: highlights keep detail, the sky and the sun glow are softly blown with a warm milky haze.
Warm, slightly saturated CCD colour, pastel pinks and terracottas, deep tropical greens, gentle bloom and halation around bright edges and chrome, a constant soft warm flare in the upper right corner, visible fine noise and grain, softly lifted shadows, imperfect consumer-camera sharpness. Candid, amateur, quietly cinematic. Aspect ratio 16:9.

=== AUDIO ===
Raw on-camera microphone sound, thin and unprocessed, light auto-gain, no denoising, no music, no sound effects.
Sound palette: a quiet golden-hour street, palm fronds, distant surf, a light breeze on the mic, the camera operator's steps, the man's footsteps and linen rustle as he walks in, one faint distant gull.

=== VOICE ===
Her: soft, low, slightly husky American English, calm and natural, starting mid-thought with no warm-up. Line 1 is quietly amused and admiring; line 2 is a teasing, confidential invitation. Human imperfections: a tiny breath between the lines, a little vocal fry at the end of "cinematic".
Him: warm, low, slightly raspy, natural English with a light natural accent. Very charismatic, cool and confident, unhurried, with a playful grin in his voice; "Then I'll tell you" lands as an effortless promise, never loud or theatrical. The same camera-mic rawness as hers, slightly closer to the mic.

What matters here:


Recap

  1. Make the first frame look like a real photo: a candid pose, a lived-in place, and the flaws of a cheap flash camera.
  2. Build a character sheet for anyone who enters the shot. Use your face photos as a face reference only.
  3. Give the video model both images and a prompt for the action and the lines.

And once more: the more real your references are, the more real your video will be.


Follow me on Instagram to see more content like this.

@morfl_ow