You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫爬取结果含冗余换行符,如何清理后正常存储为JSON

Scrapy爬虫JSON输出冗余空白字符修复方案

问题原因

当前输出JSON中大量冗余\n、\t、空白字符的来源是:直接提取了页面全部文本节点,包含HTML标签之间无实际意义的空白占位文本,未经过滤直接拼接导致,同时还可能混入script、style标签内的代码文本。

具体修复步骤

修改parse_item方法中的正文提取逻辑,添加清洗过滤规则,修改后代码如下:

import itemloaders
from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor
#from crawl.items import SpideyItem

class crawler(CrawlSpider):
    name = 'spidey'
    start_urls = ['https://quotes.toscrape.com/page/']

    rules = (
        Rule(LinkExtractor(), callback='parse_item', follow=True),
    )
    custom_settings = {
        'DEPTH_LIMIT': 1,
        'DEPTH_PRIORITY': 1,
    }

    def parse_item(self, response):
        item = dict()
        item['url'] = response.url.strip()
        item['title'] = response.meta['link_text'].strip()
        # 排除script、style标签的文本,提取有效文本
        text_list = [t.strip() for t in response.xpath('//*[not(self::script) and not(self::style)]/text()').extract()]
        # 过滤空文本片段
        valid_text = [t for t in text_list if t]
        # 可按需选择拼接方式:用空格拼接则去除所有换行,用'\n'拼接保留段落分隔
        item['body'] = ' '.join(valid_text)
        # 可选:额外去除剩余的特殊转义字符
        item['body'] = item['body'].replace('\t', ' ').replace('\r', '')
        # 也可选择保存完整源码
        #item['source'] = response.body
        yield item

效果说明

修改后输出的body字段会自动去除所有无意义的空白、换行、制表符,仅保留实际可读的页面文本内容,不会再出现大量冗余转义字符影响JSON使用。

内容的提问来源于stack exchange,提问作者nxd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 23:06:01