New feature alert in the @huggingface ecosystem!
Flash Attention 2 natively supported in huggingface transformers, supports training PEFT, and quantization (GPTQ, QLoRA, LLM.int8)
First pip install flash attention and pass use_flash_attention_2=True when loading the model!