如何用Python、pandas、ijson处理86GB大JSON文件并过滤指定产品密钥?
处理超大JSON文件的筛选方案
完全可以处理,核心是用流式JSON解析替代一次性加载整个文件到内存,具体步骤如下:
- 预处理产品密钥列表:把所有需要匹配的密钥存入哈希集合(比如Python的
set),这样可以实现O(1)时间复杂度的快速查找,避免每次匹配都遍历整个列表。 - 使用流式JSON解析工具:这类工具会逐元素读取JSON内容,不会一次性加载全部数据到内存,常见的有Python的
ijson、jsonlines,Java的Jackson Streaming API等。以Python为例,操作流程如下:- 安装依赖:
pip install ijson - 编写处理逻辑:
import json import ijson # 加载产品密钥集合 with open('product_keys.txt', 'r') as f: target_keys = {line.strip() for line in f if line.strip()} # 流式解析大JSON并筛选(输出为JSON数组) output_file = open('filtered_groups.json', 'w') output_file.write('[') # 初始化JSON数组 first_entry = True with open('large_file.json', 'rb') as f: # 迭代读取每个group对象 for group in ijson.items(f, 'groups.item'): # 检查当前group的productCodes是否包含目标密钥 has_match = any( code['value'] in target_keys for code in group.get('productCodes', []) if code.get('type') == 'productkey' ) if has_match: if not first_entry: output_file.write(',') # 写入匹配的group json.dump(group, output_file) first_entry = False output_file.write(']') output_file.close()
- 安装依赖:
- 输出格式优化:如果不需要严格的JSON数组格式,也可以用JSON Lines格式输出(每个匹配的group单独一行),这样写入逻辑更简单,也避免了数组逗号处理的麻烦,示例:
import json import ijson with open('product_keys.txt', 'r') as f: target_keys = {line.strip() for line in f if line.strip()} with open('filtered_groups.jsonl', 'w') as output_file: with open('large_file.json', 'rb') as f: for group in ijson.items(f, 'groups.item'): has_match = any( code['value'] in target_keys for code in group.get('productCodes', []) if code.get('type') == 'productkey' ) if has_match: json.dump(group, output_file) output_file.write('\n')
这种方法的内存占用仅取决于单个group对象的大小,完全可以处理86GB级别的超大JSON文件。
内容的提问来源于stack exchange,提问作者Tyler Moore
相关产品推荐
相关产品推荐

