使用Poetry运行单元测试时all-distilroberta-v1路径缺失问题求助
问题:Poetry虚拟环境中缺失sentence-transformers模型文件导致单元测试失败
项目结构
dataanalysis ├── config │ └── config.env ├── coverage.xml ├── Dockerfile ├── poetry.lock ├── pylibs │ └── analyzer-1.0.0.tar.gz ├── pylint-report.txt ├── pyproject.toml ├── requirements.txt ├── result.xml ├── scripts │ ├── __init__.py │ ├── common.py │ ├── analyze_data.py │ ├── analyze_main.py │ └── analyze_scheduler.py └── tests ├── __init__.py │ ├── __init__.cpython-38.pyc │ ├── conftest.cpython-38-pytest-7.4.0.pyc ├── conftest.py ├── test_analyze_data.py ├── test_env_vars.py └── test_analyze_main.py
Docker环境运行情况
项目已打包为Docker镜像,通过以下命令安装依赖可正常运行:
cp ./scripts /home/app_user/scripts cp ./pylibs /home/app_user/pylibs cp requirements.txt /home/app_user pip install --trusted-host pypi.org --trusted-host files.pythonhosted.org --no-cache-dir -r ./requirements.txt pip install --no-cache-dir /home/app_user/pylibs/analyzer-1.0.0.tar.gz
Poetry配置及问题重现
使用Poetry创建虚拟环境运行单元测试,pyproject.toml配置如下:
#pyproject.toml [tool.poetry] name = "dataanalysis" version = "1.0.0" .... [tool.poetry.dependencies] python = "^3.8" pycron = "3.0.0" swifter = "1.3.4" Levenshtein = "0.20.9" numpy = "1.22.0" pandas = "1.1.5" scikit-learn = "0.24.1" scipy = "1.5.3" sklearn = "0.0" mysql-connector-python = "8.0.32" analyzer = {path = "pylibs/analyzer-1.0.0.tar.gz"} textdistance = "^4.5.0" sentence_transformers="^2.2.2" # added this explicitly as unit test fails specifying this is module not found. (which not case in production) [tool.poetry.group.test.dependencies] pytest = "^7.3" pytest-coverage = "^0.0" python-dotenv = "^1.0.0" [build-system] requires = ["poetry-core"] build-backend = "poetry.core.masonry.api"
执行poetry run pytest时,先出现sentence_transformers模块未找到的错误,添加依赖后又出现以下错误:
***ValueError: Path /Users/myuser/Library/Caches/pypoetry/virtualenvs/dataanalysis-uwOjF3tf-py3.8/lib/python3.8/site-packages/analyzer/files/all-distilroberta-v1 not found ../../Library/Caches/pypoetry/virtualenvs/dataanalysis-uwOjF3tf-py3.8/lib/python3.8/site-packages/sentence_transformers/SentenceTransformer.py:77: ValueError***
该模型文件在手动创建的虚拟环境中存在,但Poetry环境中缺失,且sentence_transformers被用于analyzer-1.0.0.tar.gz内的脚本。
解决方案
1. 手动在Poetry环境中下载模型
进入Poetry虚拟环境,执行模型下载命令:
poetry shell python -c "from sentence_transformers import SentenceTransformer; model = SentenceTransformer('all-distilroberta-v1')"
执行完成后,模型会自动下载到Poetry环境的对应目录,再运行单元测试即可。
2. 检查并修复analyzer包的打包配置
问题根源是analyzer-1.0.0.tar.gz未将all-distilroberta-v1模型文件包含在包内,手动环境中模型是运行时自动下载的,但Poetry安装流程可能未触发下载。
- 解压
analyzer-1.0.0.tar.gz,查看其setup.py或pyproject.toml,确认是否配置了包含files目录的规则。如果没有,需修改配置重新打包:
在setup.py中添加:
或者在package_data={ 'analyzer': ['files/all-distilroberta-v1/**/*'], }, include_package_data=True,pyproject.toml中添加:[tool.poetry] include = ["analyzer/files/all-distilroberta-v1/**/*"]
3. 添加Poetry post-install钩子自动下载模型
在项目的pyproject.toml中添加post-install脚本配置:
[tool.poetry.scripts] post-install = "scripts.post_install:download_model"
然后在scripts目录下创建post_install.py文件:
from sentence_transformers import SentenceTransformer def download_model(): # 下载并缓存模型 SentenceTransformer('all-distilroberta-v1') print("Model all-distilroberta-v1 downloaded successfully") if __name__ == "__main__": download_model()
执行poetry install时,会自动运行该脚本下载模型,确保测试环境中模型文件存在。
4. 确认analyzer包的依赖声明
检查analyzer包的配置文件,确认是否已将sentence_transformers声明为依赖。如果没有,需添加到其pyproject.toml或setup.py中,这样Poetry安装analyzer时会自动处理依赖,避免手动添加到主项目配置。
内容的提问来源于stack exchange,提问作者Prasad
相关产品推荐
相关产品推荐

