You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python修改500MB大JSON文件并重新保存?

解决方案

基础实现:修改后直接保存

你只需要在循环中将修改后的changedTitle和changedText赋值回原数据列表的对应字段,之后使用json.dump()将整个数据结构写入文件即可。建议用with语句管理文件流,避免资源泄漏:

import json
from random import shuffle

# 读取原文件
with open('test.json', 'r', encoding='utf-8') as f:
    data = json.load(f)

# 遍历修改数据(直接遍历列表元素比用range(len(data))更简洁)
for item in data:
    # 修改title字段
    title_words = item['title'].split()
    shuffle(title_words)
    item['title'] = " ".join(title_words)
    
    # 修改text字段
    text_words = item['text']
    shuffle(text_words)
    item['text'] = " ".join(text_words)

# 保存修改后的JSON文件
with open('modified_test.json', 'w', encoding='utf-8') as f:
    # indent=2让输出格式更易读,ensure_ascii=False保留中文等非ASCII字符
    json.dump(data, f, ensure_ascii=False, indent=2)

针对大文件的优化方案

由于你的JSON文件约500MB,一次性加载到内存可能占用较多系统资源。可以采用流式处理的方式,逐行读取、修改并写入,大幅降低内存占用:

首先安装用于流式解析JSON的ijson库:

pip install ijson

然后使用以下代码:

import ijson
import json
from random import shuffle

# 流式读取原文件,同时写入修改后的内容
with open('test.json', 'r', encoding='utf-8') as infile, \
     open('modified_test.json', 'w', encoding='utf-8') as outfile:
    
    outfile.write('[')  # 先写入JSON数组的开头符号
    first_item = True
    
    # 流式遍历原文件中的每个JSON对象
    for item in ijson.items(infile, 'item'):
        # 修改title字段
        title_words = item['title'].split()
        shuffle(title_words)
        item['title'] = " ".join(title_words)
        
        # 修改text字段
        text_words = item['text']
        shuffle(text_words)
        item['text'] = " ".join(text_words)
        
        # 处理对象间的逗号分隔
        if not first_item:
            outfile.write(',')
        else:
            first_item = False
        json.dump(item, outfile, ensure_ascii=False)
    
    outfile.write(']')  # 写入JSON数组的结尾符号

这种方式无需一次性加载整个文件到内存,更适合处理大体积JSON文件。

内容的提问来源于stack exchange,提问作者allexstorm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 21:29:56