Scrapy导入自定义BookItem类遇ModuleNotFoundError问题求助
Scrapy导入Item类触发ModuleNotFoundError问题
我使用Python + Scrapy框架开发网站爬虫,尝试通过Item类(scrapy.item)处理爬取数据,需要将items.py中的自定义BookItem类导入到bookspider.py中,但运行代码时出现ModuleNotFoundError: No module named 'bookscraper'错误。
文件目录结构
PYTHONSCRAPING -bookscraper -bookscraper -spiders -bookspider.py -items.py -middlewares.py -pipelines.py -settings.py
items.py中的BookItem类代码
class BookItem(scrapy.Item): url = scrapy.Field() title = scrapy.Field() upc = scrapy.Field() product_type = scrapy.Field() price_excl_tax = scrapy.Field() price_incl_tax = scrapy.Field() tax = scrapy.Field() availability = scrapy.Field() num_reviews = scrapy.Field() stars = scrapy.Field() category = scrapy.Field() description = scrapy.Field() price = scrapy.Field()
bookspider.py中的导入语句
from bookscraper.items import BookItem
bookspider.py中使用BookItem的parse_book_page方法代码
def parse_book_page(self, response): table_rows = response.css('table tr') book_item = BookItem() # 关联items.py中的定义 book_item['url'] = response.url, book_item['title'] = response.css('.product_main h1::text').get(), book_item['upc'] = table_rows[0].css('td ::text').get(), book_item['product_type'] = table_rows[1].css('td ::text').get(), book_item['price_excl_tax'] = table_rows[2].css('td ::text').get(), book_item['price_incl_tax'] = table_rows[3].css('td ::text').get(), book_item['tax'] = table_rows[4].css('td ::text').get(), book_item['availablitiy'] = table_rows[5].css('td ::text').get(), book_item['num_reviews'] = table_rows[6].css('td ::text').get(), book_item['stars'] = response.css('p.star-rating').attrib['class'], book_item['category'] = response.xpath("//ul[@class='breadcrumb']/li[@class='active']/preceding-sibling::li[1]/a/text()").get(), book_item['description'] = response.xpath("//div[@id='product_description']/following-sibling::p/text()").get(), book_item['price'] = response.css('p.price_color ::text').get(), yield book_item
运行命令及错误信息
(venv) (base) Tomass-MBP:bookscraper fullsnack$ /Users/fullsnack/Desktop/HelloWorld/PythonScraping/venv/bin/python /Users/fullsnack/Desktop/HelloWorld/PythonScraping/bookscraper/bookscraper/spiders/bookspider.py Traceback (most recent call last): File "/Users/fullsnack/Desktop/HelloWorld/PythonScraping/bookscraper/bookscraper/spiders/bookspider.py", line 38, in <module> from bookscraper.items import BookItem #imports the items from item.py so that python can recognize when the items in the def parse_book_pages() ModuleNotFoundError: No module named 'bookscraper'
使用环境:Python 3.9.13,虚拟环境
问题原因及解决方法
核心原因:Scrapy爬虫不能直接单独运行
spider脚本,你当前直接执行bookspider.py的方式不符合Scrapy的运行规范,Python无法识别上层的bookscraper包结构。正确解决步骤:
- 激活虚拟环境,切换到外层
bookscraper目录(即包含scrapy.cfg文件的PYTHONSCRAPING/bookscraper目录) - 使用Scrapy官方命令启动爬虫:
注意:scrapy crawl bookspiderbookspider需要和你bookspider.py中爬虫类的name属性保持一致。
- 激活虚拟环境,切换到外层
额外提示:
检查parse_book_page方法中的字段拼写:book_item['availablitiy']应为availability,和items.py中的字段名保持一致,避免后续数据导出时出现字段缺失问题。
内容的提问来源于stack exchange,提问作者Tomas Di Leo
相关产品推荐
相关产品推荐

