You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy 1.5.2运行爬虫时遭遇ModuleNotFoundError问题求助

Troubleshooting "ModuleNotFoundError" for Your Scrapy Pipeline

Hey there! Let's work through this module error you're hitting with your Scrapy pipeline. I've been in your shoes when I was new to Scrapy, so let's break down the most likely fixes step by step:

1. Make sure your pipelines folder is a proper Python package

Scrapy requires the pipelines directory to be recognized as a Python package, which means it must contain an empty __init__.py file. If you don't have this file, Python won't treat entreprises.pipelines as a package, hence the error.

  • Go to your entreprises/entreprises/pipelines/ directory (the inner entreprises folder is your project's core module)
  • Create a blank file named __init__.py here. It doesn't need any code inside—just its presence tells Python this is an importable package.

2. Verify your project folder structure

Double-check that your files are arranged correctly. Your project should look like this:

entreprises/  # Outer project root folder
  ├── entreprises/  # Inner core module folder
  │   ├── __init__.py
  │   ├── items.py
  │   ├── middlewares.py
  │   ├── pipelines/
  │   │   ├── __init__.py
  │   │   └── tidyup.py  # Your pipeline file
  │   ├── settings.py
  │   └── spiders/
  │       └── usine-digital2.py  # Your spider
  └── scrapy.cfg

If your pipelines folder is in the outer entreprises directory instead of the inner one, that's a problem—Scrapy won't look there for modules.

3. Test the import manually

To confirm the path is correct, open a Python terminal from your project's root (the outer entreprises folder), and run this command:

from entreprises.pipelines.tidyup import TidyUp

If this works without errors, your Scrapy crawl should too. If it still throws the same ModuleNotFoundError, go back and recheck your folder structure and __init__.py files.

4. Clear Python's cache (if needed)

Sometimes old cached files can cause weird import issues. Delete any __pycache__ folders in your project directory, then try running your crawl command again:

scrapy crawl usine-digital2 -s CLOSESPIDER_ITEMCOUNT=30

The most common culprit here is missing the __init__.py in the pipelines folder—fixing that should resolve your error in most cases!

内容的提问来源于stack exchange,提问作者Gauthier Buttez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:32:22