Scrapy Spider未返回响应求助:爬虫代码运行异常排查
Scrapy爬虫代码运行失败求助
我是爬虫新手,编写了一段基础代码用于爬取reedsy网站。能在Scrapy Shell中定位并提取所需元素,但代码运行失败,报错信息过长无法确定具体问题,恳请帮助!
我的代码
import scrapy class PublisherSpider(scrapy.Spider): name = 'mycrawler' start_urls = ['https://blog.reedsy.com/publishers/african-american/'] def parse(self, response): for publishers in response.css('div.panel-body'): yield { 'Publisher': response.css('h3.text-heavy::text').get().replace('\n',''), 'url' : response.css('a.text-blue').attrib['href'], }
报错截图
问题分析与修正
你的代码存在两个核心问题:
- 错误使用全局response查找元素:循环遍历每个
div.panel-body时,你用的是response.css(),这会每次从整个页面查找元素,而非当前遍历的publishers节点下的子元素,导致提取内容重复或出错。 - 未处理元素不存在的情况:直接调用
.get()后链式调用.replace(),或者直接取attrib['href'],如果元素不存在会抛出AttributeError,这也是报错的主要原因。
修正后的代码:
import scrapy class PublisherSpider(scrapy.Spider): name = 'mycrawler' start_urls = ['https://blog.reedsy.com/publishers/african-american/'] def parse(self, response): # 遍历每个出版社节点 for publisher in response.css('div.panel-body'): # 从当前节点提取标题,处理空值和换行符 pub_name = publisher.css('h3.text-heavy::text').get() pub_name = pub_name.replace('\n', '').strip() if pub_name else '未知出版社' # 从当前节点提取链接,处理空值 pub_url = publisher.css('a.text-blue::attr(href)').get() yield { 'Publisher': pub_name, 'url': pub_url if pub_url else '无链接' }
修正说明
- 将循环内的
response.css()改为publisher.css(),确保从当前遍历的节点下提取子元素。 - 对提取内容做空值判断,避免因元素不存在抛出异常。
- 增加
.strip()去除多余空格,优化结果格式。
内容的提问来源于stack exchange,提问作者TLit
相关产品推荐
相关产品推荐

