Don't Fall In Love (Official Music Video) made with comfyui minimax h3.
100% MINIMAX H3 • 100% COMFYUI • FULL MUSIC VIDEO
I wanted to see if I could create an entire AI-generated music video using Minimax H3 locally in ComfyUI — without paying for the big video generators like Seedance, Kling, or even the official Minimax website.
So I did it.
🎬 100% of the AI video generation: Minimax H3 + ComfyUI
🤖 ChatGPT: prompting / planning
✂️ Camtasia: editing/ upscaled with topaz ai.
🎵 Final result: a complete AI-generated music video
This was mainly an experiment to see just how far I could push H3 running through ComfyUI.
Minimax H3 Comfyui — 1536x864 8 (25 minutes) vs 40 STEPS (over an hour) and also REALISM LORA TEST.
Pc Specs: Laptop 5090 24 vram 64 ram
***got this workflow ( I added a sage to it) from aiseek youtuber (link provided). I included both videos and also the starting image I made if you want to test it yourself.***
https://t.co/nqtvswc4Ob
Prompt:
15-second cinematic single continuous handheld shot. One uninterrupted take from start to finish. No cuts, no transitions, no camera resets.
The <Picture 0> is exactly the first frame.
Scene:
A fictional original young woman in her early twenties sits at the counter of a cozy daytime sushi bistro. Preserve the exact facial identity and physical appearance of the supplied character reference throughout the entire shot. She has fair natural skin with visible freckles, bright blue eyes, and long, naturally curly copper-red hair worn loose around her shoulders. Her hair has realistic individual strands, natural volume, and soft movement. She wears soft natural makeup, muted rose lips, and a chunky polished silver collar necklace.
She wears a vivid green sheer mesh long-sleeve top falling slightly off one bare shoulder over a matching green inner top, exactly matching the starting image. No visible brands or logos.
The sushi bistro has a long polished dark-wood counter with a woven rattan runner, warm cream walls, blank decorative scrolls with no readable writing, softly glowing cabinets, and a blurred open kitchen with sushi chefs moving naturally in the background. A red orchid sits nearby. A large window fills the restaurant with bright midday daylight, with lush green trees visible outside.
In front of her is a ceramic plate with assorted sushi, a small soy sauce dish, pink pickled ginger, a napkin, and wooden chopsticks.
Composition:
Keep the woman in the RIGHT THIRD of the frame for the entire shot. Never center her.
Shoot from slightly below her eye level, across the countertop. The long counter and blurred kitchen extend through the left two-thirds of the frame.
The foreground sushi plate and woven counter runner remain slightly out of focus along the bottom edge.
Maintain a constant 10–12 degree Dutch angle throughout the entire video. Never level the horizon.
Camera movement:
Playful natural handheld cinematography. The camera operator appears to be resting their elbows near the counter, creating subtle breathing sway, tiny micro-jitters, and gentle imperfect movement.
Throughout the shot, slowly and slightly slide the camera closer along the counter toward her face.
Never become perfectly static.
Soft window light occasionally creates a subtle natural lens flare.
Action and emotion timeline:
0–1 seconds:
The action starts immediately on the very first frame.
She already has a whole shrimp nigiri held with her chopsticks: pale pink butterflied shrimp over a small rice pillow.
She quickly lifts it and puts the entire shrimp nigiri into her mouth in one bite.
No hesitation and no slow buildup.
1–4 seconds:
The sour flavor hits her INSTANTLY with exaggerated comic intensity.
Her eyes squeeze tightly shut.
Her lips pucker dramatically around her full cheeks.
Her entire face contracts from the sourness.
Her shoulders shudder.
She quickly shakes her head side to side, causing her long curly red hair to bounce, whip and sway naturally around her face and shoulders.
Her chopsticks slip from her hand and clatter onto the ceramic plate.
She lightly pounds one fist against the counter once or twice while continuing to chew.
The reaction must be immediate, expressive, and physically readable.
4–8 seconds:
The sour shock turns into genuine disgust.
Her expression suddenly becomes flatter and deeply repulsed.
She chews slower and slower.
Her nose wrinkles.
She slightly gags.
Her tongue pushes the food toward one cheek as if she cannot tolerate the taste.
She grabs the napkin and brings it toward her mouth, clearly preparing to spit the food out.
Her eyes show disappointment and betrayal.
8–11 seconds:
With the napkin touching or hovering directly in front of her lips, she suddenly freezes.
Her eyebrows jump upward.
Her eyes open wide.
Something about the flavor unexpectedly changes.
She slowly lowers the napkin about an inch.
She chews once.
Pause.
She chews a second time.
Her expression changes from disgust → confusion → suspicious curiosity.
Make this transformation clearly visible through her eyes, eyebrows, cheeks, and mouth.
11–15 seconds:
Her expression gradually explodes into pure delight.
Her eyes close softly in pleasure.
A huge genuine smile spreads across her face while she continues chewing.
Her shoulders relax completely.
She leans and melts backward slightly on the stool.
She lightly kicks her feet beneath the counter from happiness.
She makes a subtle pleased humming expression without speaking.
Still smiling, she reaches forward again with her chopsticks toward another shrimp nigiri on the plate, clearly wanting another bite.
End the shot while she happily reaches for the next piece.
Performance:
Prioritize expressive facial acting and physically readable emotion changes.
The emotional progression must be extremely clear:
instant sour shock → disgust → about to spit it out → sudden surprise → suspicious curiosity → intense pleasure
Natural blinking, eye movement, cheek movement, jaw chewing, breathing, shoulder movement, hand gestures, and realistic loose curly-hair physics.
IDENTITY LOCK: Preserve the woman's exact facial identity from the supplied character reference and starting frame throughout the entire video. Do not alter her facial structure, freckles, eye color, skin tone, age, or copper-red curly hair. Do not change her hairstyle into a ponytail or straight hair.
Visual style:
Cinematic Japanese indie slice-of-life feeling.
35mm film texture.
Shallow depth of field.
Soft natural skin texture with her freckles clearly retained.
Warm daylight mixed with a subtle teal-and-warm film color cast.
Fine soft grain.
Natural highlight rolloff.
Cozy intimate restaurant atmosphere.
Realistic human motion and realistic food interaction.
Important constraints:
One continuous 15-second shot.
No cuts.
No scene changes.
No time jumps.
No slow motion.
No freeze-frame effect.
Do not center the woman.
Maintain the Dutch angle.
No dialogue.
No captions.
No subtitles.
No readable text.
Blank scrolls and blank menus only.
No logos.
No brands.
No recognizable characters.
No mascot figurines.
No copyrighted characters.
No real people.
MiniMax H3 COmfyui Cinematic Prompting Test | 768p • 8 Steps • No Cinematic LoRA
Prompt:
Man reference <image 0>
Woman reference <image 1>
Use the two supplied character references to preserve the exact identity and appearance of the red-haired woman and dark-haired man. EXACTLY TWO PEOPLE are present in this scene: ONE woman and ONE man. NEVER duplicate either character.
Late at night inside an old, slightly worn roadside American diner during heavy rain. The woman is already sitting alone on ONE SIDE of a booth, beside a rain-covered window, holding a ceramic cup of coffee. The man arrives and sits DIRECTLY ACROSS THE TABLE FROM HER on the OPPOSITE bench. He NEVER sits beside her. Once seated, both characters remain on their respective opposite sides of the table for the entire scene. Maintain this exact spatial relationship through every camera cut.
The visual treatment is moody, imperfect, atmospheric live-action feature-film photography, not pristine digital video. Slightly underexposed faces, deep soft shadows, low-key practical lighting, warm dim tungsten diner lamps mixed with cold blue-gray rainy window light, subdued colors, restrained saturation, gentle highlight bloom and halation around practical lights, subtle organic 35mm-style grain, slight lens softness, natural skin texture, atmospheric haze, imperfect shadow detail, soft optical depth of field and natural motion blur. Rain streaks distort distant neon and headlights outside into soft colored bokeh. Avoid clean clinical lighting, beauty lighting, glossy commercial imagery, HDR, razor-sharp edges, perfect digital clarity, plastic skin and polished AI aesthetics. It should look photographed through a real cinema lens in a dark location, with texture and imperfection.
[00:00–00:02] Moody medium-wide shot from behind the man's shoulder as he approaches the booth. The woman is visible ALONE on the opposite side of the table, looking down sadly at her coffee. He slides into the empty bench directly across from her. CUT immediately once he is seated.
[00:02–00:04] Tight close-up of the woman from the man's seated perspective. His shoulder may appear ONLY as a soft out-of-focus foreground edge. Do not show another man anywhere in frame. She raises her eyes toward him, disappointed, and quietly says: “You're late.”
[00:04–00:06] Reverse tight close-up of the man from the woman's seated perspective. Her red hair may appear ONLY as a soft blurred foreground edge. He says quietly: “I'm sorry... but I had a good reason.”
[00:06–00:08] CUT to a close tabletop insert. His hand places a small OPEN engagement-ring box beside her coffee. A beautiful engagement ring is clearly visible inside. Warm practical light catches the ring naturally. Only his hand enters the shot.
[00:08–00:10] CUT directly to an extreme close-up of the woman's face. She freezes in disbelief. Her eyes widen and begin filling with tears; her lips part, then slowly form a stunned, emotional smile. Cold rainy-window bokeh glows softly behind her. She looks across the table toward him. End on her face.
CRITICAL CONTINUITY: Exactly ONE woman and ONE man. Never clone, duplicate or create another version of either character. The woman remains on her original side of the booth. The man sits ONLY on the opposite side. He NEVER sits beside her. An over-the-shoulder foreground shoulder is part of the SAME person whose viewpoint the camera represents and must NEVER become another visible character.
NO MUSIC. Only quiet rain, subdued diner ambience and dialogue.
Expanding my old 2:3 portrait AI video "When I close My Eyes" into full 16:9 cinematic widescreen With H3 Minimax . Side by side comparision
( Special note: One thing I learned is to chop the video into individual scenes before putting them into H3 (it is labeled in the example video). If you give H3 a longer sequence with multiple shots, it can stretch certain scenes out longer than they should be, which throws off the timing and completely messes with the beat of the music.)
Prompt Used: Use the supplied video <video 1> exactly as the source and convert it to 16:9 widescreen by outpainting ONLY the missing areas on the left and right sides. Preserve the original video in the center exactly as it is, including the woman, her movement, body direction, pose, timing, camera angle, background, lighting, colors and composition. Do not recreate or reinterpret the scene, do not change the camera angle or make the woman move differently, and do not crop, zoom, stretch or reposition the original footage. Simply extend the existing environment naturally to the left and right to create a seamless 16:9 video. Maintain the original visual quality with no added noise, grain, flicker or warping. No Music.
Can AI actually dance to the beat?
Testing Minimax H3 Comfyui with original song “Don’t Fall in Love” to see how well it can make my AI character actually FEEL the music and move in rhythm.
Let’s see how she does… 👀
****(1280 x736 14 seconds. This was more about the movement than the quality so you will see some face morphing in video)****
Prompt: integrated_multimodal_description:
<Image 1> is the CHARACTER APPEARANCE REFERENCE for femal dancer.
<audio 1> use for music reference.
integrated_multimodal_description:
Create a 15-second realistic cinematic dance-performance test using the supplied female character image as the exact visual reference and starting frame.
Use the supplied music clip as the ONLY audio <audio 1>.
CRITICAL MUSIC SYNCHRONIZATION:
The woman actively listens and dances to the actual rhythm of the supplied music.
Analyze the music’s tempo, percussion, bass line, and strongest beat accents.
Her choreography must remain synchronized with the music throughout the entire video.
Her steps, hip movements, shoulder movements, arm gestures, head turns, and changes in direction land naturally on the musical beats.
Do NOT create random dancing that is disconnected from the music.
Do NOT change the speed of the song.
Do NOT generate additional music, vocals, dialogue, clapping, footsteps, or sound effects.
PERFORMANCE AND PERSONALITY:
She is DANCING AND HAVING FUN.
She genuinely enjoys the music and looks like she is feeling the rhythm rather than mechanically performing choreography.
She has lively eyes, natural smiles, playful expressions, confident energy, and spontaneous little reactions to the music.
Her facial expression naturally changes throughout the dance. She can smile wider when a strong beat hits, briefly laugh to herself, give the camera a playful look, or show that she is simply enjoying the song.
Her personality should come through in her face AND body language.
CRITICAL: NO DEAD, BLANK, EMPTY, BORED, EMOTIONLESS, MODEL-POSE EXPRESSION.
She must never look like an AI character mechanically executing dance movements.
Her expression and body language should feel connected: when she smiles, her eyes and posture participate naturally.
She performs a confident, playful, feminine dance suitable for an energetic pop-rock song.
The movement should feel natural and realistic—not like a rehearsed professional dance routine and not exaggerated social-media dancing.
She begins with a subtle bounce and shoulder movement as she catches the rhythm.
She then adds controlled side-to-side steps, natural hip movement, relaxed arm gestures, and occasional confident turns toward the camera.
Every fourth beat receives a slightly stronger movement, such as a sharper step, shoulder hit, hip accent, hair movement, or brief turn.
She remains energized throughout the full 15 seconds without repeating the exact same movement continuously.
CAMERA:
Begin with a clear full-body shot so her feet and entire body remain visible.
Use a mostly stable camera with a very slow cinematic push inward.
Do not use rapid editing, spinning cameras, extreme camera movement, or cuts that hide whether she is following the beat.
Maintain realistic body mechanics, balance, weight shifts, foot placement, and momentum.
CHARACTER CONTINUITY:
Preserve the woman’s exact facial identity, hairstyle, hair color, age, body proportions, clothing, and accessories from the supplied reference.
Do not alter her outfit.
Do not change her face or body during movement.
Keep both hands anatomically correct.
No extra fingers, duplicated limbs, warped joints, sliding feet, floating feet, or unnatural body motion.
She does NOT sing, speak, or lip-sync.
Her mouth remains naturally expressive because she is enjoying herself, but she NEVER mouths or sings the lyrics.
The final result should look like a real woman who loves the song, is genuinely having fun, and naturally dances precisely in rhythm with its beat.
NO GENERATED MUSIC. NO DIALOGUE. NO LIP-SYNC. SUPPLIED SONG AUDIO ONLY.
Kling 3.0...Prompt:Scene Description:A playful blonde woman sitting on a modern wooden table in a bright high-rise apartment with floor-to-ceiling windows and a city skyline in the background. She’s wearing a fitted white crop tank, denim shorts, and knee-high white socks. She has a fun, lively expression with an open, cheerful https://t.co/EVodL0kM7k & Motion:Fast-moving, smooth cinematic camera.The camera starts in front of her at mid-body https://t.co/Dofyr9evub quickly pushes forward toward her with energetic motion.Without stopping, the camera continues into a smooth 360-degree arc around her body.The motion is fluid, confident, and upbeat — not slow.The camera completes the full arc and lands back in front of her.Performance:She maintains a bright, playful smile the entire https://t.co/JyLknz5Gtf the camera returns to the starting position, she lifts her hand and gives a quick, fun wave toward the camera.She leans slightly forward as she waves, playful and charismatic.She mouths “bye” silently with a wink (optional if you want more personality).Style:Natural daylight.Modern lifestyle commercial look.Crisp focus.Soft depth of field.Clean color grading.Upbeat, energetic https://t.co/J2diEpMhRw slow motion.Ending:She finishes the wave as the screen quickly fades to black. Music Playing in background Upbeat pop-funk at 120 BPMWith a strong beat hit right when the camera finishes the arc and she waves.
Krea 2 turbo and Minimax H3 in Comfyui (Upscaled in Topaz AI)....integrated_multimodal_description:
[10-second image-to-video generation]
Use the supplied image as the starting frame and preserve the established characters, tavern environment, clothing, robot designs, bar layout, and warm cinematic lighting.
STYLE / TONE:
Live-action cinematic science-fiction pirate adventure with playful Hollywood-style banter. This is a cozy local tavern in a sci-fi world where human and robot pirates casually drink together.
The conversation is playful and familiar. The robot pirate is extremely boastful, arrogant, swaggering, and proud of his reputation. He genuinely believes he is an exceptional pirate. The bartender finds his exaggerated stories amusing and enjoys teasing him.
CRITICAL ROBOT SPEECH RULE:
ALL robot faces remain mechanically rigid.
ROBOTS NEVER MOVE THEIR MOUTHS OR JAWS WHILE SPEAKING.
NO robotic lip-sync.
Robot speech comes from internal electronic speakers.
When a robot speaks, ONLY its glowing blue eyes subtly brighten, pulse, or blink in rhythm with the spoken words.
[00:00–00:04]
The human bartender stands behind the bar holding her drink and looks directly at the robot pirate across from her.
The robot pirate (S1) looks directly toward her with confident pirate swagger. He gives a small proud tilt of his head.
His metal mouth and jaw remain COMPLETELY MOTIONLESS.
His glowing blue eyes pulse subtly as he boasts in an arrogant, self-satisfied mechanical male voice:
(S1) <d>[English] He drew his blaster. I drew mine faster.</d>
[00:04–00:07]
The bartender (S2) raises an eyebrow and gives him a playful, skeptical grin. She maintains eye contact with HIM while speaking.
Because she is human, her lips move naturally and accurately with her dialogue:
(S2) <d>[English] You don't even carry a blaster.</d>
[00:07–00:10]
Cut slightly closer on the robot pirate.
Brief comedic pause.
He confidently leans back just slightly, completely unfazed by being caught exaggerating.
His mouth and jaw NEVER MOVE.
His blue eyes brighten with smug confidence as he replies:
(S1) <d>[English] Exactly. And I still won.</d>
The bartender immediately laughs and shakes her head at his ridiculous arrogance.
End on the robot sitting proudly as though his story made perfect sense.
BACKGROUND ROBOTS:
The other robot pirates remain background patrons. Their mechanical mouths and jaws remain rigid at ALL times. They may make tiny natural head movements or subtle eye-light changes, but they do not speak during this clip.
overall_soundscape:
Warm local-tavern ambience, quiet distant patron chatter, occasional glass clinks, subtle mechanical servo noises, natural bartender laughter, and clean internal-speaker robot dialogue. Robot eye lights subtly pulse during their speech.
non_diegetic_music:
N/A