Bark Text-to-Audio Model
Full Text Input: "Why was six afraid of seven?"
Ignore Bark's "I'm done with this input" token and tell Bark to just keep generating more audio anyway.
I wanted to imagine how we’d better use #stablediffusion for video content / AR.
A major obstacle, why most videos are so flickery, is lack of temporal & viewing angle consistency, so I experimented with an approach to fix this
See 🧵 for process & examples