Hah. Yeah I mean it's definitely not a "one secret prompt" kind of thing, so in that sense it would be hard to recreate without having a decent amount of experience with After Effects and Comfy.
That said it's not necessarily all that complicated, it's just a lot of elements and a fairly long process (8-10 hours maybe), which would obviously be a lot longer if you're figuring it out at the same time
But here's the process/recipe:
- Generate a bunch of key frames with GPT Image 2 and Fable, by giving Fable detailed instructions and image references to write a prompt for GPT (which doesn't get to see the image refs). This video has 6 of those key frames, but I'd usually make around 20-30 to pick from.
- First frame/last frame Wan 2.2 I2V gens between each image to create a loop, using a transition/morphing LoRA I've been tinkering with since Wan 2.1 came out (still not perfect). Wan is obviously not the only model that you could do this with, but I like how predictable it is and the short clip duration tends to help when working in this way
- Smooth out the joins with Vace (one gen per intersection)
- Low denoise upscale of the full loop with Wan 2.2 T2V (helps smooth out the GPT image artifacts)
- Final Vace smoothing pass to stitch together the end/start of the loop as the V2V pass generally ends up making them quite different
- Then once I have a solid looping video I create a bunch of different elements in Comfy, including depth/normals/masks etc, as well as stuff created with a bunch of different nodes I've made, such as the time slice effect, text labels, coloured BBOXes - essentially a bunch of different ways of visualising tracking data
- Once I have all the pieces I bring them all into After Effects and play around with layering, timing, masking etc until it feels right (which might be one of the harder parts to replicate as it's basically just a taste thing)
- I usually generate 100+ tracks with Suno for each video, select 10-15, lay them over the rough pass and go through each one to see if there is a section that fits the mood. Then I rejig the edit a bit to try to hit some of the beats
But yeah - I guess it's a bit like asking if anyone can train a LoRA...which they can, but it helps to have some experience doing it if you want it to be good 😅
Callipeg Studio is coming soon to macOS and Windows: including features like the inbetween assist, a great way to place your next drawing to get the best animation. (animation by Inamura Takeshi @altamontagna170 )
Seedance 2.0で短編を制作中なのですが、クオリティを保つために試しているワークフローです
① GPT image2で6コマ絵コンテを作る
② キャラ部分を単色で塗りつぶす
③ その絵コンテをキャラシートと一緒にSeedanceへ渡す
④「塗りつぶした部分を、キャラシートの人物に置き換えてください」と指示する
GPT image2で絵コンテを作り続けると、絵柄がGPT image2っぽく収斂していくことがあります。
なのでMidjourneyなどで作ったキャラクターデザインを維持したまま長めの尺を作りたい場合は、一度絵コンテ上のキャラを単色で塗りつぶす工程を挟んでいます。
Gemini 3.1 Flash TTS is our most controllable text-to-speech model yet.
With new Audio Tags, you can easily direct vocal style, delivery, and pace through text commands. 🧵