You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在JSON文件分类中添加子分类?Python爬取JSON格式调整

技术问题与解决方案

技术问题

如何在JSON文件的分类中添加子分类?

场景需求

我用Python结合BeautifulSoup爬取指定网页,结果保存为.json文件,当前代码如下:

from bs4 import BeautifulSoup
import json
import requests

url = 'https://storage.googleapis.com/infosimples-public/commercia/case/product.html#'

resposta_final = {}

response = requests.get(url)

parsed_html = BeautifulSoup(response.content, 'html.parser')
    
resposta_final['skus'] = [element.get_text(strip=True) for element in parsed_html.select(".skus-area")] 

json_resposta_final = json.dumps(resposta_final)

with open('produto.json','w' ) as arquivo_json:
    arquivo_json.write(json_resposta_final)

当前生成的produto.json内容为:

{
  "skus": [
    "Rubber Duck MK Ultra - Original$ 7.95$ 9.95Rubber Duck MK Ultra - Summer VersionOut of stockRubber Duck MK Ultra - Batman Version$ 14.95"
  ]
}

我需要调整为以下中文键名的格式:

{
  "skus": [
    {
      "名称": [
        "Rubber Duck MK Ultra - Original",
        "Rubber Duck MK Ultra - Summer Version",
        "Rubber Duck MK Ultra - Batman Version"
      ],
      "现价": [
        "$7.95",
        null,
        "$14.95"
      ],
      "原价": [
        "$9.95",
        null,
        null
      ],
      "是否可用": [
        true,
        false,
        true
      ]
    }
  ]
}

修改后的代码实现

from bs4 import BeautifulSoup
import json
import requests

url = 'https://storage.googleapis.com/infosimples-public/commercia/case/product.html#'

response = requests.get(url)
parsed_html = BeautifulSoup(response.content, 'html.parser')

# 定位每个独立的SKU容器
sku_items = parsed_html.select('.sku-item')

# 初始化各字段的存储列表
names = []
current_prices = []
old_prices = []
availabilities = []

for item in sku_items:
    # 提取SKU名称
    name = item.select_one('.sku-name').get_text(strip=True)
    names.append(name)
    
    # 处理现价和原价,没有对应价格就设为None(对应JSON的null)
    price_elements = item.select('.sku-price')
    current_price = price_elements[0].get_text(strip=True) if len(price_elements) >= 1 else None
    old_price = price_elements[1].get_text(strip=True) if len(price_elements) >= 2 else None
    
    # 判断库存状态
    stock_status = item.select_one('.sku-stock').get_text(strip=True)
    available = stock_status != "Out of stock"
    
    current_prices.append(current_price)
    old_prices.append(old_price)
    availabilities.append(available)

# 构建目标JSON结构
resposta_final = {
    "skus": [
        {
            "名称": names,
            "现价": current_prices,
            "原价": old_prices,
            "是否可用": availabilities
        }
    ]
}

# 保存为格式化的JSON文件
with open('produto.json', 'w', encoding='utf-8') as arquivo_json:
    json.dump(resposta_final, arquivo_json, indent=2, ensure_ascii=False)

核心调整说明

  • 别抓整块文本:原代码直接提取整个SKU区域的所有文本,导致内容混杂。现在先定位每个SKU的独立容器.sku-item,再逐个字段提取。
  • 分字段处理数据:
    • 名称直接从.sku-name节点提取,避免混入价格、库存信息
    • 价格区分现价和原价,没有对应价格就留空(对应JSON的null)
    • 库存状态通过判断文本是否为"Out of stock",转换为布尔值
  • 搭建子分类结构:把各字段的列表放入skus数组的子对象中,实现需求的分层结构。

内容的提问来源于stack exchange,提问作者zssp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 13:16:16