You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy Spider未返回响应求助:爬虫代码运行异常排查

Scrapy爬虫代码运行失败求助

我是爬虫新手,编写了一段基础代码用于爬取reedsy网站。能在Scrapy Shell中定位并提取所需元素,但代码运行失败,报错信息过长无法确定具体问题,恳请帮助!

我的代码

import scrapy

class PublisherSpider(scrapy.Spider):
    name = 'mycrawler'
    start_urls = ['https://blog.reedsy.com/publishers/african-american/']
   
    def parse(self, response):
        for publishers in response.css('div.panel-body'):
            yield {
                'Publisher': response.css('h3.text-heavy::text').get().replace('\n',''),
                'url' : response.css('a.text-blue').attrib['href'],
            }

报错截图

问题分析与修正

你的代码存在两个核心问题:

  1. 错误使用全局response查找元素:循环遍历每个div.panel-body时,你用的是response.css(),这会每次从整个页面查找元素,而非当前遍历的publishers节点下的子元素,导致提取内容重复或出错。
  2. 未处理元素不存在的情况:直接调用.get()后链式调用.replace(),或者直接取attrib['href'],如果元素不存在会抛出AttributeError,这也是报错的主要原因。

修正后的代码:

import scrapy

class PublisherSpider(scrapy.Spider):
    name = 'mycrawler'
    start_urls = ['https://blog.reedsy.com/publishers/african-american/']
   
    def parse(self, response):
        # 遍历每个出版社节点
        for publisher in response.css('div.panel-body'):
            # 从当前节点提取标题,处理空值和换行符
            pub_name = publisher.css('h3.text-heavy::text').get()
            pub_name = pub_name.replace('\n', '').strip() if pub_name else '未知出版社'
            
            # 从当前节点提取链接,处理空值
            pub_url = publisher.css('a.text-blue::attr(href)').get()
            
            yield {
                'Publisher': pub_name,
                'url': pub_url if pub_url else '无链接'
            }

修正说明

  • 将循环内的response.css()改为publisher.css(),确保从当前遍历的节点下提取子元素。
  • 对提取内容做空值判断,避免因元素不存在抛出异常。
  • 增加.strip()去除多余空格,优化结果格式。

内容的提问来源于stack exchange,提问作者TLit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 05:47:21