安装DGL 2.1.0+cu118时进程意外终止问题求助
之前安装了不支持CUDA的DGL 2.1.0,卸载后尝试安装DGL 2.1.0+cu118版本,执行pip install dgl -f https://data.dgl.ai/wheels/cu118/repo.html命令时,安装包下载完成后进程直接被标注Killed,重复操作、重启终端都无法解决,完整终端输出如下:
(graduation)root@xxx:~/venv/graduation# pip install dgl -f https://data.dgl.ai/wheels/cu118/repo.html
Looking in indexes: http://mirrors.aliyun.com/pypi/simple
Looking in links: https://data.dgl.ai/wheels/cu118/repo.html
Collecting dgl
Downloading https://data.dgl.ai/wheels/cu118/dgl-2.1.0%2Bcu118-cp38-cp38-manylinux1_x86_64.whl (748.2 MB)
|████████████████████████████████| 748.2 MB 589 kB/s eta 0:00:01Killed
以下是可行的解决方法:
排查内存/交换空间不足问题
进程被Killed最常见的原因是系统内存不够,大体积安装包解压和安装时会占用大量内存。用free -h查看内存和swap使用情况,若swap空间不足,可临时新增swap分区:- 创建2G大小的swap文件(可按需调整数值):
sudo fallocate -l 2G /swapfile - 设置文件权限:
sudo chmod 600 /swapfile - 启用swap:
sudo mkswap /swapfile && sudo swapon /swapfile - 安装完成后如需清理,执行:
sudo swapoff /swapfile && sudo rm /swapfile
- 创建2G大小的swap文件(可按需调整数值):
使用pip内存优化参数
给pip添加--no-cache-dir参数,避免缓存占用额外内存,修改后的安装命令:pip install --no-cache-dir dgl -f https://data.dgl.ai/wheels/cu118/repo.html手动下载后本地安装
跳过pip自动下载环节,手动获取安装包后本地安装:- 下载安装包:
wget https://data.dgl.ai/wheels/cu118/dgl-2.1.0%2Bcu118-cp38-cp38-manylinux1_x86_64.whl - 本地安装:
pip install ./dgl-2.1.0+cu118-cp38-cp38-manylinux1_x86_64.whl
- 下载安装包:
调高进程资源限制
用ulimit -a查看当前用户的进程资源限制,若虚拟内存限制过低,可临时在当前终端调高:ulimit -v unlimited
内容的提问来源于stack exchange,提问作者Boyun

