You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy运行提示找不到items模块且数据无法传入Pipeline问题求助

错误原因排查
  • 模块导入路径错误:Scrapy项目默认结构下,items.py和spiders文件夹同属项目根模块下,你的导入语句from watches.watches.items import WatchesItem多写了一层watches,触发watches.watches模块不存在的报错。同理Pipeline中from watches.watches.spiders import watchbot也存在路径冗余问题,且该导入无实际作用。
  • 运行工作目录错误:补充日志显示你在E:\semester\webcrawler_watches\watches\Crawler路径下运行爬虫,Scrapy找不到项目根目录下的scrapy.cfg配置文件,无法识别项目模块,触发watches模块不存在的报错。
  • items.py字段定义语法错误:你将output_processor参数写进了MapCompose的括号内,会导致后续Item处理异常。
  • Pipeline取值逻辑冲突:你已经给Item字段设置了TakeFirst()处理器,字段返回结果为字符串而非列表,item['name'][0]写法会触发字符串索引报错。
解决方案
  1. 修正导入路径
    把爬虫文件中的导入语句:
from watches.watches.items import WatchesItem

修改为:

from watches.items import WatchesItem

直接删除Pipeline中冗余的from watches.watches.spiders import watchbot导入语句。如果你的项目根模块(包含items.py、pipelines.py、settings.py、spiders文件夹的目录)名称不是watches,将导入语句中的watches替换为实际的根模块名称即可。

  1. 切换运行目录
    找到项目中scrapy.cfg文件所在的根目录,cd到该目录后再执行scrapy crawl watchbot命令。

  2. 修正Item字段定义
    把items.py中三个字段的参数拆分,将output_processor移到MapCompose外:

name = scrapy.Field(input_processor = MapCompose(remove_tags), output_processor = TakeFirst())
reference = scrapy.Field(input_processor = MapCompose(remove_tags), output_processor = TakeFirst())
year = scrapy.Field(input_processor = MapCompose(remove_tags), output_processor = TakeFirst())
  1. 修正Pipeline取值逻辑
    把store_db方法中的列表索引去掉:
self.curr.execute("""insert into test.watch values (%s, %s, %s)""", (
    item['name'],
    item['reference'],
    item['year'],
))

内容的提问来源于stack exchange,提问作者SyrixGG

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 21:45:05