You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python调用json.dump追加写JSON文件未自动添加逗号分隔如何解决?

问题原因
  • 你使用了文件追加模式a循环写入单个字典对象:json.dump仅会将传入的单个Python对象序列化为合法JSON字符串,不会自动为多个独立对象添加分隔逗号,也不会自动包裹为JSON数组,直接拼接的多个字典不符合JSON语法规范。
  • 辅助变量specifications_list、specifications_dict定义在循环外部,每次遍历商品时不会清空,会导致后续商品的字典中混入前面商品的旧数据,属于逻辑bug。
修复方案

推荐优先使用「先全量收集数据再一次性写入」的方案,实现简单且不会出现格式错误,修复后代码如下:

import json
import requests
from bs4 import BeautifulSoup

with open("all_categories_dict.json") as file:
  all_products = json.load(file)

# 定义总列表存储所有商品的规格字典
all_specifications = []

for product_name, product_href in all_products.items(): 
  url = product_href
  response = requests.get(url)
  soup = BeautifulSoup(response.text, "lxml")

  price = soup.find("div", class_="new_price").text
  describe = soup.find("div", class_="additional_info").find(class_="tabs").find_all("li") 
  # 每个商品单独初始化变量,避免数据污染
  specifications_dict = {}
  # 固定字段直接写入字典,不需要额外中转列表
  specifications_dict["Наименование товара"] = product_name
  specifications_dict["Ссылка"] = product_href
  specifications_dict["Цена"] = price
  for specification in describe:
    k, *v = specification.text.replace('\r\n\t', "").strip().split(':')
    specifications_dict[k.strip()] = ':'.join(v).strip()
  # 将当前商品的字典加入总列表
  all_specifications.append(specifications_dict)

# 所有数据收集完成后一次性写入文件
with open("specifications.json", "w", encoding="utf-8") as f:
  json.dump(all_specifications, f, indent=4, ensure_ascii=False)

补充说明

如果爬取的商品量级极大,内存无法承载全量数据,可以使用边爬边写的方案:

  1. 首次打开文件时先写入左中括号[
  2. 每次循环写完单个字典后写入逗号
  3. 所有商品处理完成后,回退1个字节覆盖掉最后一个多余的逗号,再写入右中括号]
    该方案需要处理文件指针偏移,实现起来相对麻烦,非必要场景不推荐。

内容的提问来源于stack exchange,提问作者Dima Melnik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 00:57:03