You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Scrapy爬取多URL后生成带scans层级的标准JSON?

问题描述

我使用Scrapy框架,通过以下代码从urllist.txt文件加载多个URL到start_urls参数:

with open('urllist.txt') as f:
    start_urls = f.read().splitlines()

随后用如下代码解析页面:

def parse(self, response):
    yield {
            'date': datetime.now().strftime("%d/%m/%Y %H:%M"),
            'item': response.xpath("//font[contains(@id, 'item')]/a/text()").get(),
            'price': response.xpath("//font[contains(@id, 'price')]/a/text()").get(),
          }

代码运行正常,但生成的JSON文件是每行一个JSON对象的格式:

{"date": "05/01/2023 02:39:49", "item": "Item 1", "price": "6.80"}
{"date": "05/01/2023 02:39:49", "item": "Item 2", "price": "8.50"}
{"date": "06/01/2023 02:39:49", "item": "Item 1", "price": "6.70"}
{"date": "06/01/2023 02:39:49", "item": "Item 2", "price": "8.50"}

我期望生成包含scans数组的标准JSON结构:

{"scans": [
    {"date": "05/01/2023 02:39:49", "item": "Item 1", "price": "6.80"},
    {"date": "05/01/2023 02:39:49", "item": "Item 2", "price": "8.50"},
    {"date": "06/01/2023 02:39:49", "item": "Item 1", "price": "6.70"},
    {"date": "06/01/2023 02:39:49", "item": "Item 2", "price": "8.50"}
]}

由于后续要将该JSON导入页面,AJAX不支持这种非标准格式,请问是否必须对结果进行后处理以添加scans层级?

解决方案

不需要非得依赖后处理,Scrapy内部可以直接生成目标格式,也可以用简单脚本做后处理,两种方式都可行:

方法1:Scrapy内部生成目标格式(推荐)

自定义Item Pipeline

编写一个Pipeline收集所有爬取到的item,最后统一输出包含scans数组的JSON:

在项目的pipelines.py中添加以下代码:

import json

class ScansPipeline:
    def open_spider(self, spider):
        self.items = []

    def process_item(self, item, spider):
        self.items.append(dict(item))
        return item

    def close_spider(self, spider):
        with open('output.json', 'w', encoding='utf-8') as f:
            json.dump({'scans': self.items}, f, indent=2)

然后在settings.py中启用这个Pipeline:

ITEM_PIPELINES = {
    '你的项目名称.pipelines.ScansPipeline': 300,
}

注意运行爬虫时不要使用scrapy crawl自带的-o参数导出文件,因为Pipeline会自动生成目标文件。

方法2:后处理脚本(简单直接)

如果不想修改Scrapy项目代码,可以写一个简单的Python脚本处理已生成的JSON文件:

import json

# 读取原JSON文件
with open('原始输出文件名.json', 'r', encoding='utf-8') as f:
    items = [json.loads(line) for line in f if line.strip()]

# 生成包含scans数组的结构
output_data = {'scans': items}

# 写入新文件
with open('格式化后文件名.json', 'w', encoding='utf-8') as f:
    json.dump(output_data, f, indent=2)

运行该脚本即可将每行一个对象的JSON转换成你需要的标准结构。


内容的提问来源于stack exchange,提问作者douglaslps

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 03:21:00