处理大JSON文件拆分时Python出现MemoryError问题求助
Hey there! That MemoryError is totally expected—you’re trying to load a 500MB+ JSON file entirely into memory with json.load(), which is way too much for Python to handle in one go. JSON parsing expands data in memory (since Python dictionaries have overhead), so that 500MB file could easily balloon to 1GB+ once loaded. Let’s fix this with streaming parsing—processing the file piece by piece instead of all at once.
Why Your Current Code Fails
Your code reads the entire JSON array into d_list using json.load(json_data), which loads every single item into memory simultaneously. For small files this works, but for 500MB+ files, it’s guaranteed to hit a MemoryError.
Solution 1: Use the ijson Library (Recommended)
ijson is a lightweight library that parses JSON incrementally, so it only loads one item at a time into memory. Here’s how to adjust your code:
First, install the library:
pip install ijson
Then update your code:
import ijson import glob import os s_path = "D:\\User\\Desktop\\users\\SourceFolder\\" extension = "*.json" target_path = 'D:\\Desktop\\users\\DumpingFolder\\' # Make sure the target folder exists (no errors if it already does) os.makedirs(target_path, exist_ok=True) for file in glob.iglob(s_path + extension): print(f"Processing file: {file}") with open(file, encoding="utf-8-sig") as json_data: # Stream each item from the JSON array for item in ijson.items(json_data, 'item'): file_name = f"Raw_Response_{item['id']}.json" print(f"Writing {file_name}") # Use os.path.join to avoid path separator issues with open(os.path.join(target_path, file_name), 'w+', encoding='utf-8') as file_obj: json.dump(item, file_obj, indent=2, sort_keys=False)
How This Works
ijson.items(json_data, 'item') iterates over each element in the top-level JSON array without loading the entire file into memory. Each item is processed and written to its own file immediately, keeping memory usage super low.
Solution 2: Manual Streaming (No Third-Party Libraries)
If you can’t install external libraries, you can use Python’s built-in json.JSONDecoder to parse the file incrementally:
import json import glob import os s_path = "D:\\User\\Desktop\\users\\SourceFolder\\" extension = "*.json" target_path = 'D:\\Desktop\\users\\DumpingFolder\\' os.makedirs(target_path, exist_ok=True) for file in glob.iglob(s_path + extension): print(f"Processing file: {file}") with open(file, encoding="utf-8-sig") as f: decoder = json.JSONDecoder() buffer = "" for line in f: buffer += line.strip() # Skip the opening bracket of the array if buffer.startswith('['): buffer = buffer[1:] # Keep trying to parse complete objects from the buffer while buffer: try: # Parse the next complete JSON object item, idx = decoder.raw_decode(buffer) file_name = f"Raw_Response_{item['id']}.json" print(f"Writing {file_name}") with open(os.path.join(target_path, file_name), 'w+', encoding='utf-8') as file_obj: json.dump(item, file_obj, indent=2, sort_keys=False) # Update the buffer to remove the parsed object and any leading commas buffer = buffer[idx:].lstrip(',').strip() except json.JSONDecodeError: # Not enough data for a complete object—read more lines break
How This Works
This code builds a buffer of text from the file, then uses raw_decode() to parse the first complete JSON object it finds. Once parsed, it removes that object from the buffer and repeats, never loading more than a small chunk of the file into memory at once.
Extra Tips
- Always use
os.path.join()for file paths to avoid issues with backslashes/forward slashes across operating systems. - Specify
encoding='utf-8'when writing files to ensure consistent encoding. - If your JSON file has a different structure (not a top-level array), adjust the
ijsonpath (e.g.,'*.items'for nested arrays) or tweak the manual parsing logic.
内容的提问来源于stack exchange,提问作者diksha ojha

