使用Scrapy爬取网站遇ModuleNotFoundError问题排查
解决Scrapy报错:ModuleNotFoundError: No module named 'bookscraper.settings'
问题背景
执行scrapy或scrapy crawl bookspider -o bookdata.csv命令时,出现ModuleNotFoundError: No module named 'bookscraper.settings'错误。已确认项目scrapy.cfg文件正确、代码缩进无误,该项目属于FreeCodeCamp的Scrapy教程内容。
报错栈信息
Traceback (most recent call last): File "C:\Users\Tunansh Vatsa\AppData\Local\Programs\Python\Python310\lib\runpy.py", line 196, in _run_module_as_main return _run_code(code, main_globals, None, File "C:\Users\Tunansh Vatsa\AppData\Local\Programs\Python\Python310\lib\runpy.py", line 86, in _run_code exec(code, run_globals) File "C:\Users\Tunansh Vatsa\Desktop\New folder (3)\SCRAPY\venv\Scripts\scrapy.exe\__main__.py", line 7, in <module> File "C:\Users\Tunansh Vatsa\Desktop\New folder (3)\SCRAPY\venv\lib\site-packages\scrapy\cmdline.py", line 172, in execute settings = get_project_settings() File "C:\Users\Tunansh Vatsa\Desktop\New folder (3)\SCRAPY\venv\lib\site-packages\scrapy\utils\project.py", line 73, in get_project_settings settings.setmodule(settings_module_path, priority="project") File "C:\Users\Tunansh Vatsa\Desktop\New folder (3)\SCRAPY\venv\lib\site-packages\scrapy\settings\__init__.py", line 481, in setmodule module = import_module(module) File "C:\Users\Tunansh Vatsa\AppData\Local\Programs\Python\Python310\lib\importlib\__init__.py", line 126, in import_module return _bootstrap._gcd_import(name[level:], package, level) File "<frozen importlib._bootstrap>", line 1050, in _gcd_import File "<frozen importlib._bootstrap>", line 1027, in _find_and_load File "<frozen importlib._bootstrap>", line 1004, in _find_and_load_unlocked ModuleNotFoundError: No module named 'bookscraper.settings'
爬虫代码(bookspider.py)
import scrapy class BookspiderSpider(scrapy.Spider): name = 'bookspider' allowed_domains = ['books.toscrape.com'] start_urls = ['https://books.toscrape.com/'] def parse(self, response): books = response.css('article.product_pod') for book in books: relative_url = book.css('h3 a ::attr(href)').get() if 'catalogue/' in relative_url: book_url = 'https://books.toscrape.com/' + relative_url else: book_url = 'https://books.toscrape.com/catalogue/' + relative_url yield response.follow(book_url, callback=self.parse_book_page) next_page = response.css('li.next a ::attr(href)').get() if next_page is not None: if 'catalogue/' in next_page: next_page_url = 'https://books.toscrape.com/' + next_page else: next_page_url = 'https://books.toscrape.com/catalogue/' + next_page yield response.follow(next_page_url, callback=self.parse) def parse_book_page(self, response): table_rows = response.css("table tr") yield { 'url' : response.url, 'title' : response.css('.product_main h1::text').get(), 'product_type': table_rows[1].css("td ::text").get(), 'price_excl_tax': table_rows[2].css("td ::text").get(), 'price_incl_tax': table_rows[3].css("td ::text").get(), 'tax': table_rows[4].css("td ::text").get(), 'availability': table_rows[5].css("td ::text").get(), 'num_reviews': table_rows[6].css("td ::text").get(), 'stars' : response.css("p.star-rating").attrib['class'], 'category' : response.xpath("//ul[@class='breadcrumb']/li[@class='active']/preceding-sibling::li[1]/a/text()").get(), 'description' : response.xpath("//div[@id='product_description']/following-sibling::p/text()").get(), 'price': response.css('p.price_color ::text').get(), }
排查方案
- 确认命令执行路径:必须切换到包含
scrapy.cfg的项目根目录下执行命令。比如你的项目根目录是C:\Users\Tunansh Vatsa\Desktop\New folder (3)\SCRAPY,先通过cd "C:\Users\Tunansh Vatsa\Desktop\New folder (3)\SCRAPY"切换路径,再执行爬虫命令。 - 验证项目目录结构:确保项目符合Scrapy标准结构,重点检查:
- 存在
bookscraper文件夹,且内部有settings.py文件 bookscraper文件夹内有__init__.py(空文件即可,用于标识Python模块)- 爬虫文件
bookspider.py放在bookscraper/spiders/目录下
标准结构示例:
SCRAPY/ ├── scrapy.cfg └── bookscraper/ ├── __init__.py ├── settings.py ├── items.py ├── pipelines.py └── spiders/ ├── __init__.py └── bookspider.py - 存在
- 二次校验scrapy.cfg配置:打开
scrapy.cfg,确认[settings]部分的配置为:
确保[settings] default = bookscraper.settingsbookscraper和你的项目模块名完全一致(大小写保持统一)。 - 激活虚拟环境:执行命令前先激活项目对应的虚拟环境,Windows下运行
venv\Scripts\activate,激活后再执行Scrapy命令。 - 排查命名冲突:检查项目目录下是否存在与
bookscraper同名的文件或文件夹,避免Python导入模块时混淆。
内容的提问来源于stack exchange,提问作者noob_coder123
相关产品推荐
相关产品推荐

