Python项目跨文件夹导入函数报错,求解决方案与最佳实践
解决Python爬虫项目模块导入错误及最佳实践
一、解决"No module named Utils"错误的方法
1. 临时添加项目根目录到Python路径
在tourney_scraper.py的最开头加入以下代码,让Python能定位到上层的Utils文件夹:
import sys from pathlib import Path # 将当前脚本的父级父目录(项目根目录)添加到Python模块搜索路径 sys.path.append(str(Path(__file__).parent.parent))
添加后再执行from Utils.test import test即可正常导入。
2. 将文件夹转为Python包(推荐方案)
给Utils和Scrapers文件夹都创建空的__init__.py文件,标记它们为Python可识别的包。
方式A:从项目根目录直接运行脚本
打开终端切换到项目根目录,执行:
python Scrapers/tourney_scraper.py
此时from Utils.test import test可直接生效。
方式B:使用相对导入(适合包内模块调用)
将导入语句改为相对路径格式:
from ..Utils.test import test
然后从项目根目录以包模块的方式运行:
python -m Scrapers.tourney_scraper
二、爬虫项目的最佳实践
1. 规范项目目录结构
采用分层结构,清晰区分通用工具、业务爬虫和入口文件,便于维护扩展:
爬虫项目/ ├── __init__.py ├── Utils/ │ ├── __init__.py │ ├── http_utils.py # 封装请求、代理、重试等网络操作 │ ├── parse_utils.py # 封装HTML解析、数据提取逻辑 │ └── storage_utils.py # 封装数据库存储、文件导出等工具 ├── Scrapers/ │ ├── __init__.py │ ├── tourney_scraper.py │ └── news_scraper.py ├── main.py # 统一入口,管理各爬虫的启动、调度 └── requirements.txt # 记录项目依赖包版本
2. 使用统一入口启动爬虫
在项目根目录创建main.py,集中管理爬虫启动逻辑,彻底规避路径问题:
from Scrapers.tourney_scraper import run_tourney_scraper from Scrapers.news_scraper import run_news_scraper if __name__ == "__main__": # 按需启动指定爬虫 run_tourney_scraper() # run_news_scraper()
运行时只需在项目根目录执行:
python main.py
3. 用虚拟环境隔离依赖
为项目创建独立虚拟环境,避免与全局Python环境的依赖冲突:
# 创建虚拟环境 python -m venv venv # Windows激活虚拟环境 venv\Scripts\activate # Linux/Mac激活虚拟环境 source venv/bin/activate # 安装依赖包 pip install requests beautifulsoup4 # 导出依赖到文件 pip freeze > requirements.txt
4. 配置可安装项目(进阶)
通过pyproject.toml将项目配置为可安装包,实现全局任意位置导入模块:
在项目根目录创建pyproject.toml:
[build-system] requires = ["setuptools>=61.0"] build-backend = "setuptools.build_meta" [project] name = "crawler-collection" version = "0.1.0"
执行以下命令以可编辑模式安装项目:
pip install -e .
之后在任何位置都能直接导入项目模块,比如from Utils.http_utils import send_request。
内容的提问来源于stack exchange,提问作者Jesper van Beemdelust
相关产品推荐
相关产品推荐

