如何用Python修改500MB大JSON文件并重新保存?
解决方案
基础实现:修改后直接保存
你只需要在循环中将修改后的changedTitle和changedText赋值回原数据列表的对应字段,之后使用json.dump()将整个数据结构写入文件即可。建议用with语句管理文件流,避免资源泄漏:
import json from random import shuffle # 读取原文件 with open('test.json', 'r', encoding='utf-8') as f: data = json.load(f) # 遍历修改数据(直接遍历列表元素比用range(len(data))更简洁) for item in data: # 修改title字段 title_words = item['title'].split() shuffle(title_words) item['title'] = " ".join(title_words) # 修改text字段 text_words = item['text'] shuffle(text_words) item['text'] = " ".join(text_words) # 保存修改后的JSON文件 with open('modified_test.json', 'w', encoding='utf-8') as f: # indent=2让输出格式更易读,ensure_ascii=False保留中文等非ASCII字符 json.dump(data, f, ensure_ascii=False, indent=2)
针对大文件的优化方案
由于你的JSON文件约500MB,一次性加载到内存可能占用较多系统资源。可以采用流式处理的方式,逐行读取、修改并写入,大幅降低内存占用:
首先安装用于流式解析JSON的ijson库:
pip install ijson
然后使用以下代码:
import ijson import json from random import shuffle # 流式读取原文件,同时写入修改后的内容 with open('test.json', 'r', encoding='utf-8') as infile, \ open('modified_test.json', 'w', encoding='utf-8') as outfile: outfile.write('[') # 先写入JSON数组的开头符号 first_item = True # 流式遍历原文件中的每个JSON对象 for item in ijson.items(infile, 'item'): # 修改title字段 title_words = item['title'].split() shuffle(title_words) item['title'] = " ".join(title_words) # 修改text字段 text_words = item['text'] shuffle(text_words) item['text'] = " ".join(text_words) # 处理对象间的逗号分隔 if not first_item: outfile.write(',') else: first_item = False json.dump(item, outfile, ensure_ascii=False) outfile.write(']') # 写入JSON数组的结尾符号
这种方式无需一次性加载整个文件到内存,更适合处理大体积JSON文件。
内容的提问来源于stack exchange,提问作者allexstorm
相关产品推荐
相关产品推荐

