You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy爬取作者名返回None或空白字符问题求助

解决Scrapy爬取作者名返回None或空白字符的问题

问题根源

你当前的CSS选择器span.elementor-post-info__item--type-author::text只获取了span标签本身的文本,但目标站点的作者名实际嵌套在该span下的<a>标签里,所以直接取span的自然是空值或空白。另外,即使拿到正确文本,也可能包含换行、制表符这类空白字符,需要清理。

修改后的代码

class NewsSpider(scrapy.Spider):
    name = "cruiseradio"

    def start_requests(self):
        url = input("Enter the article url: ")
        yield scrapy.Request(url, callback=self.parse_dir_contents)

    def parse_dir_contents(self, response):
        # 定位到span下的a标签获取作者文本,同时处理空白和None的情况
        author_text = response.css('span.elementor-post-info__item--type-author a::text').get(default='').strip()
        # 如果清理后为空,设为"NULL"
        author = author_text if author_text else "NULL"
        yield {
            'Author': author,
        }

关键调整说明

  • 修正选择器: 把选择器改为span.elementor-post-info__item--type-author a::text,直接定位到存储作者名的<a>标签文本。
  • 清理空白字符: 用strip()去掉文本前后的换行、制表符和空格,避免拿到一堆空白符号。
  • 处理空值: 使用get(default='')确保即使没找到元素也不会返回None,再判断清理后的文本是否为空,为空则设为"NULL"——原代码的try-except其实没用,因为get()不会抛出IndexError。

内容的提问来源于stack exchange,提问作者Info Rewind

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 17:55:20