@Selen7005717917 If removing some intermediate layers doesn’t hurt the result, it may mean the model wasn’t trained well. What we want is a model where every layer is useful. Adding adaptive loops is a way to further increase or decrease efficient computational depth.
@AlexiGlad Interesting work! Beyond hard/soft-min, have you considered a more general aggregation over the (K) losses, such as top-(m)? It could offer a useful trade-off between specialization and mode coverage.