Scrapy爬虫爬取结果含冗余换行符,如何清理后正常存储为JSON
Scrapy爬虫JSON输出冗余空白字符修复方案
问题原因
当前输出JSON中大量冗余\n、\t、空白字符的来源是:直接提取了页面全部文本节点,包含HTML标签之间无实际意义的空白占位文本,未经过滤直接拼接导致,同时还可能混入script、style标签内的代码文本。
具体修复步骤
修改parse_item方法中的正文提取逻辑,添加清洗过滤规则,修改后代码如下:
import itemloaders from scrapy.spiders import CrawlSpider, Rule from scrapy.linkextractors import LinkExtractor #from crawl.items import SpideyItem class crawler(CrawlSpider): name = 'spidey' start_urls = ['https://quotes.toscrape.com/page/'] rules = ( Rule(LinkExtractor(), callback='parse_item', follow=True), ) custom_settings = { 'DEPTH_LIMIT': 1, 'DEPTH_PRIORITY': 1, } def parse_item(self, response): item = dict() item['url'] = response.url.strip() item['title'] = response.meta['link_text'].strip() # 排除script、style标签的文本,提取有效文本 text_list = [t.strip() for t in response.xpath('//*[not(self::script) and not(self::style)]/text()').extract()] # 过滤空文本片段 valid_text = [t for t in text_list if t] # 可按需选择拼接方式:用空格拼接则去除所有换行,用'\n'拼接保留段落分隔 item['body'] = ' '.join(valid_text) # 可选:额外去除剩余的特殊转义字符 item['body'] = item['body'].replace('\t', ' ').replace('\r', '') # 也可选择保存完整源码 #item['source'] = response.body yield item
效果说明
修改后输出的body字段会自动去除所有无意义的空白、换行、制表符,仅保留实际可读的页面文本内容,不会再出现大量冗余转义字符影响JSON使用。
内容的提问来源于stack exchange,提问作者nxd
相关产品推荐
相关产品推荐

