使用Scrapy爬取作者名返回None或空白字符问题求助
解决Scrapy爬取作者名返回None或空白字符的问题
问题根源
你当前的CSS选择器span.elementor-post-info__item--type-author::text只获取了span标签本身的文本,但目标站点的作者名实际嵌套在该span下的<a>标签里,所以直接取span的自然是空值或空白。另外,即使拿到正确文本,也可能包含换行、制表符这类空白字符,需要清理。
修改后的代码
class NewsSpider(scrapy.Spider): name = "cruiseradio" def start_requests(self): url = input("Enter the article url: ") yield scrapy.Request(url, callback=self.parse_dir_contents) def parse_dir_contents(self, response): # 定位到span下的a标签获取作者文本,同时处理空白和None的情况 author_text = response.css('span.elementor-post-info__item--type-author a::text').get(default='').strip() # 如果清理后为空,设为"NULL" author = author_text if author_text else "NULL" yield { 'Author': author, }
关键调整说明
- 修正选择器: 把选择器改为
span.elementor-post-info__item--type-author a::text,直接定位到存储作者名的<a>标签文本。 - 清理空白字符: 用
strip()去掉文本前后的换行、制表符和空格,避免拿到一堆空白符号。 - 处理空值: 使用
get(default='')确保即使没找到元素也不会返回None,再判断清理后的文本是否为空,为空则设为"NULL"——原代码的try-except其实没用,因为get()不会抛出IndexError。
内容的提问来源于stack exchange,提问作者Info Rewind
相关产品推荐
相关产品推荐

