You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Poetry运行单元测试时all-distilroberta-v1路径缺失问题求助

问题:Poetry虚拟环境中缺失sentence-transformers模型文件导致单元测试失败

项目结构

dataanalysis
├── config
│  └── config.env
├── coverage.xml
├── Dockerfile
├── poetry.lock
├── pylibs
│  └── analyzer-1.0.0.tar.gz
├── pylint-report.txt
├── pyproject.toml
├── requirements.txt
├── result.xml
├── scripts
│  ├── __init__.py
│  ├── common.py
│  ├── analyze_data.py
│  ├── analyze_main.py
│  └── analyze_scheduler.py
└── tests
    ├── __init__.py
    │  ├── __init__.cpython-38.pyc
    │  ├── conftest.cpython-38-pytest-7.4.0.pyc
    ├── conftest.py
    ├── test_analyze_data.py
    ├── test_env_vars.py
    └── test_analyze_main.py

Docker环境运行情况

项目已打包为Docker镜像,通过以下命令安装依赖可正常运行:

cp ./scripts  /home/app_user/scripts
cp ./pylibs  /home/app_user/pylibs
cp requirements.txt /home/app_user
pip install --trusted-host pypi.org --trusted-host files.pythonhosted.org --no-cache-dir -r ./requirements.txt
pip install --no-cache-dir /home/app_user/pylibs/analyzer-1.0.0.tar.gz

Poetry配置及问题重现

使用Poetry创建虚拟环境运行单元测试,pyproject.toml配置如下:

#pyproject.toml
[tool.poetry]
name = "dataanalysis"
version = "1.0.0"
....

[tool.poetry.dependencies]
python = "^3.8"
pycron = "3.0.0"
swifter = "1.3.4"
Levenshtein = "0.20.9"
numpy = "1.22.0"
pandas = "1.1.5"
scikit-learn = "0.24.1"
scipy = "1.5.3"
sklearn = "0.0"
mysql-connector-python = "8.0.32"
analyzer = {path = "pylibs/analyzer-1.0.0.tar.gz"}
textdistance = "^4.5.0"
sentence_transformers="^2.2.2" # added this explicitly as unit test fails specifying this is module not found. (which not case in production)

[tool.poetry.group.test.dependencies]
pytest = "^7.3"
pytest-coverage = "^0.0"
python-dotenv = "^1.0.0"

[build-system]
requires = ["poetry-core"]
build-backend = "poetry.core.masonry.api"

执行poetry run pytest时,先出现sentence_transformers模块未找到的错误,添加依赖后又出现以下错误:

***ValueError: Path /Users/myuser/Library/Caches/pypoetry/virtualenvs/dataanalysis-uwOjF3tf-py3.8/lib/python3.8/site-packages/analyzer/files/all-distilroberta-v1 not found

../../Library/Caches/pypoetry/virtualenvs/dataanalysis-uwOjF3tf-py3.8/lib/python3.8/site-packages/sentence_transformers/SentenceTransformer.py:77: ValueError***

该模型文件在手动创建的虚拟环境中存在,但Poetry环境中缺失,且sentence_transformers被用于analyzer-1.0.0.tar.gz内的脚本。

解决方案

1. 手动在Poetry环境中下载模型

进入Poetry虚拟环境,执行模型下载命令:

poetry shell
python -c "from sentence_transformers import SentenceTransformer; model = SentenceTransformer('all-distilroberta-v1')"

执行完成后,模型会自动下载到Poetry环境的对应目录,再运行单元测试即可。

2. 检查并修复analyzer包的打包配置

问题根源是analyzer-1.0.0.tar.gz未将all-distilroberta-v1模型文件包含在包内,手动环境中模型是运行时自动下载的,但Poetry安装流程可能未触发下载。

  • 解压analyzer-1.0.0.tar.gz,查看其setup.py或pyproject.toml,确认是否配置了包含files目录的规则。如果没有,需修改配置重新打包:
    在setup.py中添加:
    package_data={
        'analyzer': ['files/all-distilroberta-v1/**/*'],
    },
    include_package_data=True,
    
    或者在pyproject.toml中添加:
    [tool.poetry]
    include = ["analyzer/files/all-distilroberta-v1/**/*"]
    

3. 添加Poetry post-install钩子自动下载模型

在项目的pyproject.toml中添加post-install脚本配置:

[tool.poetry.scripts]
post-install = "scripts.post_install:download_model"

然后在scripts目录下创建post_install.py文件:

from sentence_transformers import SentenceTransformer

def download_model():
    # 下载并缓存模型
    SentenceTransformer('all-distilroberta-v1')
    print("Model all-distilroberta-v1 downloaded successfully")

if __name__ == "__main__":
    download_model()

执行poetry install时,会自动运行该脚本下载模型,确保测试环境中模型文件存在。

4. 确认analyzer包的依赖声明

检查analyzer包的配置文件,确认是否已将sentence_transformers声明为依赖。如果没有,需添加到其pyproject.toml或setup.py中,这样Poetry安装analyzer时会自动处理依赖,避免手动添加到主项目配置。

内容的提问来源于stack exchange,提问作者Prasad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 15:03:11