You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy导入自定义BookItem类遇ModuleNotFoundError问题求助

Scrapy导入Item类触发ModuleNotFoundError问题

我使用Python + Scrapy框架开发网站爬虫,尝试通过Item类(scrapy.item)处理爬取数据,需要将items.py中的自定义BookItem类导入到bookspider.py中,但运行代码时出现ModuleNotFoundError: No module named 'bookscraper'错误。

文件目录结构

PYTHONSCRAPING
-bookscraper
  -bookscraper
   -spiders
    -bookspider.py
   -items.py
   -middlewares.py
   -pipelines.py
   -settings.py

items.py中的BookItem类代码

class BookItem(scrapy.Item):
    url = scrapy.Field()
    title = scrapy.Field()
    upc = scrapy.Field()
    product_type = scrapy.Field()
    price_excl_tax = scrapy.Field()
    price_incl_tax = scrapy.Field()
    tax = scrapy.Field()
    availability = scrapy.Field()
    num_reviews = scrapy.Field()
    stars = scrapy.Field()
    category = scrapy.Field()
    description = scrapy.Field()
    price = scrapy.Field()

bookspider.py中的导入语句

from bookscraper.items import BookItem

bookspider.py中使用BookItem的parse_book_page方法代码

def parse_book_page(self, response):
    table_rows = response.css('table tr')
    book_item = BookItem() # 关联items.py中的定义
    
    book_item['url'] = response.url,
    book_item['title'] = response.css('.product_main h1::text').get(),
    book_item['upc'] = table_rows[0].css('td ::text').get(),
    book_item['product_type'] = table_rows[1].css('td ::text').get(),
    book_item['price_excl_tax'] = table_rows[2].css('td ::text').get(),
    book_item['price_incl_tax'] = table_rows[3].css('td ::text').get(),
    book_item['tax'] = table_rows[4].css('td ::text').get(),
    book_item['availablitiy'] = table_rows[5].css('td ::text').get(),
    book_item['num_reviews'] = table_rows[6].css('td ::text').get(),
    book_item['stars'] = response.css('p.star-rating').attrib['class'],
    book_item['category'] = response.xpath("//ul[@class='breadcrumb']/li[@class='active']/preceding-sibling::li[1]/a/text()").get(),
    book_item['description'] = response.xpath("//div[@id='product_description']/following-sibling::p/text()").get(),
    book_item['price'] = response.css('p.price_color ::text').get(),
    yield book_item

运行命令及错误信息

(venv) (base) Tomass-MBP:bookscraper fullsnack$ /Users/fullsnack/Desktop/HelloWorld/PythonScraping/venv/bin/python /Users/fullsnack/Desktop/HelloWorld/PythonScraping/bookscraper/bookscraper/spiders/bookspider.py
Traceback (most recent call last):
  File "/Users/fullsnack/Desktop/HelloWorld/PythonScraping/bookscraper/bookscraper/spiders/bookspider.py", line 38, in <module>
    from bookscraper.items import BookItem #imports the items from item.py so that python can recognize when the items in the def parse_book_pages() 
ModuleNotFoundError: No module named 'bookscraper'

使用环境:Python 3.9.13,虚拟环境


问题原因及解决方法

  1. 核心原因:Scrapy爬虫不能直接单独运行spider脚本,你当前直接执行bookspider.py的方式不符合Scrapy的运行规范,Python无法识别上层的bookscraper包结构。

  2. 正确解决步骤:

    • 激活虚拟环境,切换到外层bookscraper目录(即包含scrapy.cfg文件的PYTHONSCRAPING/bookscraper目录)
    • 使用Scrapy官方命令启动爬虫:
      scrapy crawl bookspider
      
      注意:bookspider需要和你bookspider.py中爬虫类的name属性保持一致。
  3. 额外提示:
    检查parse_book_page方法中的字段拼写:book_item['availablitiy']应为availability,和items.py中的字段名保持一致,避免后续数据导出时出现字段缺失问题。

内容的提问来源于stack exchange,提问作者Tomas Di Leo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 17:23:18