You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python、pandas、ijson处理86GB大JSON文件并过滤指定产品密钥?

处理超大JSON文件的筛选方案

完全可以处理,核心是用流式JSON解析替代一次性加载整个文件到内存,具体步骤如下:

  • 预处理产品密钥列表:把所有需要匹配的密钥存入哈希集合(比如Python的set),这样可以实现O(1)时间复杂度的快速查找,避免每次匹配都遍历整个列表。
  • 使用流式JSON解析工具:这类工具会逐元素读取JSON内容,不会一次性加载全部数据到内存,常见的有Python的ijson、jsonlines,Java的Jackson Streaming API等。以Python为例,操作流程如下:
    1. 安装依赖:pip install ijson
    2. 编写处理逻辑:
      import json
      import ijson
      
      # 加载产品密钥集合
      with open('product_keys.txt', 'r') as f:
          target_keys = {line.strip() for line in f if line.strip()}
      
      # 流式解析大JSON并筛选(输出为JSON数组)
      output_file = open('filtered_groups.json', 'w')
      output_file.write('[')  # 初始化JSON数组
      first_entry = True
      
      with open('large_file.json', 'rb') as f:
          # 迭代读取每个group对象
          for group in ijson.items(f, 'groups.item'):
              # 检查当前group的productCodes是否包含目标密钥
              has_match = any(
                  code['value'] in target_keys 
                  for code in group.get('productCodes', [])
                  if code.get('type') == 'productkey'
              )
              if has_match:
                  if not first_entry:
                      output_file.write(',')
                  # 写入匹配的group
                  json.dump(group, output_file)
                  first_entry = False
      
      output_file.write(']')
      output_file.close()
      
  • 输出格式优化:如果不需要严格的JSON数组格式,也可以用JSON Lines格式输出(每个匹配的group单独一行),这样写入逻辑更简单,也避免了数组逗号处理的麻烦,示例:
    import json
    import ijson
    
    with open('product_keys.txt', 'r') as f:
        target_keys = {line.strip() for line in f if line.strip()}
    
    with open('filtered_groups.jsonl', 'w') as output_file:
        with open('large_file.json', 'rb') as f:
            for group in ijson.items(f, 'groups.item'):
                has_match = any(
                    code['value'] in target_keys 
                    for code in group.get('productCodes', [])
                    if code.get('type') == 'productkey'
                )
                if has_match:
                    json.dump(group, output_file)
                    output_file.write('\n')
    

这种方法的内存占用仅取决于单个group对象的大小,完全可以处理86GB级别的超大JSON文件。

内容的提问来源于stack exchange,提问作者Tyler Moore

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 19:15:06