You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy爬取网站遇ModuleNotFoundError问题排查

解决Scrapy报错:ModuleNotFoundError: No module named 'bookscraper.settings'

问题背景

执行scrapy或scrapy crawl bookspider -o bookdata.csv命令时,出现ModuleNotFoundError: No module named 'bookscraper.settings'错误。已确认项目scrapy.cfg文件正确、代码缩进无误,该项目属于FreeCodeCamp的Scrapy教程内容。

报错栈信息

Traceback (most recent call last):
  File "C:\Users\Tunansh Vatsa\AppData\Local\Programs\Python\Python310\lib\runpy.py", line 196, in _run_module_as_main
    return _run_code(code, main_globals, None,
  File "C:\Users\Tunansh Vatsa\AppData\Local\Programs\Python\Python310\lib\runpy.py", line 86, in _run_code
    exec(code, run_globals)
  File "C:\Users\Tunansh Vatsa\Desktop\New folder (3)\SCRAPY\venv\Scripts\scrapy.exe\__main__.py", line 7, in <module>
  File "C:\Users\Tunansh Vatsa\Desktop\New folder (3)\SCRAPY\venv\lib\site-packages\scrapy\cmdline.py", line 172, in execute
    settings = get_project_settings()
  File "C:\Users\Tunansh Vatsa\Desktop\New folder (3)\SCRAPY\venv\lib\site-packages\scrapy\utils\project.py", line 73, in get_project_settings
    settings.setmodule(settings_module_path, priority="project")
  File "C:\Users\Tunansh Vatsa\Desktop\New folder (3)\SCRAPY\venv\lib\site-packages\scrapy\settings\__init__.py", line 481, in setmodule    
    module = import_module(module)
  File "C:\Users\Tunansh Vatsa\AppData\Local\Programs\Python\Python310\lib\importlib\__init__.py", line 126, in import_module
    return _bootstrap._gcd_import(name[level:], package, level)
  File "<frozen importlib._bootstrap>", line 1050, in _gcd_import
  File "<frozen importlib._bootstrap>", line 1027, in _find_and_load
  File "<frozen importlib._bootstrap>", line 1004, in _find_and_load_unlocked
ModuleNotFoundError: No module named 'bookscraper.settings'

爬虫代码(bookspider.py)

import scrapy
class BookspiderSpider(scrapy.Spider):
    name = 'bookspider'
    allowed_domains = ['books.toscrape.com']
    start_urls = ['https://books.toscrape.com/']
    def parse(self, response):
        books = response.css('article.product_pod')
        for book in books:
            relative_url = book.css('h3 a ::attr(href)').get()

            if 'catalogue/' in relative_url:
                book_url = 'https://books.toscrape.com/' + relative_url
            else:
                book_url = 'https://books.toscrape.com/catalogue/' + relative_url
            yield response.follow(book_url, callback=self.parse_book_page)

        next_page = response.css('li.next a ::attr(href)').get()
        if next_page is not None:
            if 'catalogue/' in next_page:
                next_page_url = 'https://books.toscrape.com/' + next_page
            else:
                next_page_url = 'https://books.toscrape.com/catalogue/' + next_page
            yield response.follow(next_page_url, callback=self.parse)


    def parse_book_page(self, response):

        table_rows = response.css("table tr")
        
        yield {
            'url' : response.url,
            'title' : response.css('.product_main h1::text').get(),
            'product_type': table_rows[1].css("td ::text").get(),
            'price_excl_tax': table_rows[2].css("td ::text").get(),
            'price_incl_tax': table_rows[3].css("td ::text").get(),
            'tax': table_rows[4].css("td ::text").get(),
            'availability': table_rows[5].css("td ::text").get(),
            'num_reviews': table_rows[6].css("td ::text").get(),
            'stars' : response.css("p.star-rating").attrib['class'],
            'category' : response.xpath("//ul[@class='breadcrumb']/li[@class='active']/preceding-sibling::li[1]/a/text()").get(),
            'description' : response.xpath("//div[@id='product_description']/following-sibling::p/text()").get(),
            'price': response.css('p.price_color ::text').get(),
        }

排查方案

  • 确认命令执行路径:必须切换到包含scrapy.cfg的项目根目录下执行命令。比如你的项目根目录是C:\Users\Tunansh Vatsa\Desktop\New folder (3)\SCRAPY,先通过cd "C:\Users\Tunansh Vatsa\Desktop\New folder (3)\SCRAPY"切换路径,再执行爬虫命令。
  • 验证项目目录结构:确保项目符合Scrapy标准结构,重点检查:
    • 存在bookscraper文件夹,且内部有settings.py文件
    • bookscraper文件夹内有__init__.py(空文件即可,用于标识Python模块)
    • 爬虫文件bookspider.py放在bookscraper/spiders/目录下
      标准结构示例:
    SCRAPY/
    ├── scrapy.cfg
    └── bookscraper/
        ├── __init__.py
        ├── settings.py
        ├── items.py
        ├── pipelines.py
        └── spiders/
            ├── __init__.py
            └── bookspider.py
    
  • 二次校验scrapy.cfg配置:打开scrapy.cfg,确认[settings]部分的配置为:
    [settings]
    default = bookscraper.settings
    
    确保bookscraper和你的项目模块名完全一致(大小写保持统一)。
  • 激活虚拟环境:执行命令前先激活项目对应的虚拟环境,Windows下运行venv\Scripts\activate,激活后再执行Scrapy命令。
  • 排查命名冲突:检查项目目录下是否存在与bookscraper同名的文件或文件夹,避免Python导入模块时混淆。

内容的提问来源于stack exchange,提问作者noob_coder123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 02:27:06