You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

处理大JSON文件拆分时Python出现MemoryError问题求助

Fixing MemoryError When Splitting Large JSON Files

Hey there! That MemoryError is totally expected—you’re trying to load a 500MB+ JSON file entirely into memory with json.load(), which is way too much for Python to handle in one go. JSON parsing expands data in memory (since Python dictionaries have overhead), so that 500MB file could easily balloon to 1GB+ once loaded. Let’s fix this with streaming parsing—processing the file piece by piece instead of all at once.

Why Your Current Code Fails

Your code reads the entire JSON array into d_list using json.load(json_data), which loads every single item into memory simultaneously. For small files this works, but for 500MB+ files, it’s guaranteed to hit a MemoryError.

ijson is a lightweight library that parses JSON incrementally, so it only loads one item at a time into memory. Here’s how to adjust your code:

First, install the library:

pip install ijson

Then update your code:

import ijson
import glob
import os

s_path = "D:\\User\\Desktop\\users\\SourceFolder\\"
extension = "*.json"
target_path = 'D:\\Desktop\\users\\DumpingFolder\\'

# Make sure the target folder exists (no errors if it already does)
os.makedirs(target_path, exist_ok=True)

for file in glob.iglob(s_path + extension):
    print(f"Processing file: {file}")
    with open(file, encoding="utf-8-sig") as json_data:
        # Stream each item from the JSON array
        for item in ijson.items(json_data, 'item'):
            file_name = f"Raw_Response_{item['id']}.json"
            print(f"Writing {file_name}")
            # Use os.path.join to avoid path separator issues
            with open(os.path.join(target_path, file_name), 'w+', encoding='utf-8') as file_obj:
                json.dump(item, file_obj, indent=2, sort_keys=False)

How This Works

ijson.items(json_data, 'item') iterates over each element in the top-level JSON array without loading the entire file into memory. Each item is processed and written to its own file immediately, keeping memory usage super low.

Solution 2: Manual Streaming (No Third-Party Libraries)

If you can’t install external libraries, you can use Python’s built-in json.JSONDecoder to parse the file incrementally:

import json
import glob
import os

s_path = "D:\\User\\Desktop\\users\\SourceFolder\\"
extension = "*.json"
target_path = 'D:\\Desktop\\users\\DumpingFolder\\'

os.makedirs(target_path, exist_ok=True)

for file in glob.iglob(s_path + extension):
    print(f"Processing file: {file}")
    with open(file, encoding="utf-8-sig") as f:
        decoder = json.JSONDecoder()
        buffer = ""
        for line in f:
            buffer += line.strip()
            # Skip the opening bracket of the array
            if buffer.startswith('['):
                buffer = buffer[1:]
            # Keep trying to parse complete objects from the buffer
            while buffer:
                try:
                    # Parse the next complete JSON object
                    item, idx = decoder.raw_decode(buffer)
                    file_name = f"Raw_Response_{item['id']}.json"
                    print(f"Writing {file_name}")
                    with open(os.path.join(target_path, file_name), 'w+', encoding='utf-8') as file_obj:
                        json.dump(item, file_obj, indent=2, sort_keys=False)
                    # Update the buffer to remove the parsed object and any leading commas
                    buffer = buffer[idx:].lstrip(',').strip()
                except json.JSONDecodeError:
                    # Not enough data for a complete object—read more lines
                    break

How This Works

This code builds a buffer of text from the file, then uses raw_decode() to parse the first complete JSON object it finds. Once parsed, it removes that object from the buffer and repeats, never loading more than a small chunk of the file into memory at once.

Extra Tips

  • Always use os.path.join() for file paths to avoid issues with backslashes/forward slashes across operating systems.
  • Specify encoding='utf-8' when writing files to ensure consistent encoding.
  • If your JSON file has a different structure (not a top-level array), adjust the ijson path (e.g., '*.items' for nested arrays) or tweak the manual parsing logic.

内容的提问来源于stack exchange,提问作者diksha ojha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:29:26