After examining it a bit, I couldn't understand why this was necessary if it was captioned correctly in the structure of flux, maybe we can't make our captions correctly for now and this method improves training. But good results are already obtained without caption, maybe this is related to the mask, as I said, I'm just trying to understand. We can't know until we try, but there are so many things to try.
flux looks really good, I'm already excited about its potential. I hope the community embraces this potential and @bfl_ml can be more generous with the license