Azure Databricks GPU集群无法安装Flash Attention以运行Hugging Face模型的问题求助
大家好,我遇到了一个在Azure Databricks GPU集群上安装Flash Attention的难题,想请各位帮忙排查下:
问题背景
我在Databricks的CPU集群上可以顺利运行以下Hugging Face模型代码:
import torch import transformers model = transformers.AutoModelForCausalLM.from_pretrained( "mosaicml/mpt-7b", trust_remote_code=True, torch_dtype=torch.bfloat16, )
在CPU集群上,我预先安装了这些依赖:
- PyTorch 2.0.1
- Transformers 4.28.1
- einops 0.6.1
GPU集群出现的运行错误
但把完全相同的代码放到GPU集群执行时,直接抛出了导入错误:
ImportError: This modeling file requires the following packages that were not found in your environment: flash_attn. Run pip install flash_attn
安装Flash Attention的尝试及失败情况
第一次尝试:直接执行pip安装
我在GPU集群上运行pip install flash-attn,但安装过程中编译环节失败,关键错误信息如下:
In file included from /tmp/pip-install-nluf5697/flash-attn_79c1dfb03cc4482ba86c435b2db7a8b6/csrc/flash_attn/src/fmha.h:39, from /tmp/pip-install-nluf5697/flash-attn_79c1dfb03cc4482ba86c435b2db7a8b6/csrc/flash_attn/src/fmha_bwd_launch_template.h:6, from /tmp/pip-install-nluf5697/flash-attn_79c1dfb03cc4482ba86c435b2db7a8b6/csrc/flash_attn/src/fmha_bwd_hdim32.cu:5: /local_disk0/.ephemeral_nfs/cluster_libraries/python/lib/python3.10/site-packages/torch/include/ATen/cuda/CUDAContext.h:6:10: fatal error: cusparse.h: No such file or directory 6 | #include <cusparse.h> | ^~~~~~~~~~~~ compilation terminated. ... RuntimeError: Error compiling objects for extension
最终安装流程终止,提示legacy-install-failure。
第二次尝试:升级PyTorch后重新安装
我把GPU集群上的PyTorch版本升级到了和CPU集群一致的2.0.1,之后再次尝试安装flash-attn,但还是遇到了完全相同的编译失败问题。
我的初步猜测
我怀疑问题根源在GPU集群的预装环境:CPU集群没有预装旧版本的PyTorch,我手动安装2.0.1后一切正常;但GPU集群本身预装了更早版本的PyTorch,可能和Flash Attention的编译要求存在兼容性冲突,导致安装始终失败。
有没有朋友遇到过类似的场景?或者有什么可行的解决办法能让我在Databricks GPU集群上成功安装Flash Attention并运行mpt-7b模型?
备注:内容来源于stack exchange,提问作者Padhraig

