如何分块遍历JSON对象?已有Python可迭代对象分块函数求适配
Chunked Traversal for JSON Array Files
Got it, let's build on the solid foundation you already have with your chunked_iterable function to create a solution that handles JSON array files efficiently—especially helpful if you're working with large files that you don't want to load entirely into memory.
Complete Implementation
Here's the full code that combines your chunking logic with streaming JSON parsing:
from pathlib import Path import json from itertools import islice def chunked_iterable(iterable, chunk_size): it = iter(iterable) while True: chunk = tuple(islice(it, chunk_size)) if not chunk: break yield chunk def json_chunked_iterable(json_path, chunk_size): # Open the JSON file in read mode with open(json_path, 'r') as f: decoder = json.JSONDecoder() # Skip the opening array bracket '[' f.read(1) while True: # Skip whitespace (commas, newlines between elements) f.seek(f.tell()) char = f.read(1) if not char or char == ']': break if not char.isspace(): f.seek(f.tell() - 1) # Decode the next individual JSON object obj, _ = decoder.raw_decode(f) yield obj # Example usage if __name__ == "__main__": target_file = Path("your_large_data.json") # Stream JSON elements and chunk them into groups of 2 for chunk in chunked_iterable(json_chunked_iterable(target_file, chunk_size=2), chunk_size=2): print("Received chunk:") print(chunk)
How This Works
- Streaming JSON Parsing: The
json_chunked_iterablefunction usesjson.JSONDecoder.raw_decodeto parse one object at a time from the file, avoiding loading the entire array into memory. This is critical for large datasets. - Reuse Your Chunking Logic: We wrap the stream of JSON objects with your existing
chunked_iterablefunction to group elements into the chunk size you specify. - Edge Case Handling: It skips whitespace and properly stops when it hits the closing array bracket
].
Simplified Version (For Smaller Files)
If your JSON file is small enough that memory isn't an issue, you can take a more straightforward approach:
from pathlib import Path import json from itertools import islice def chunked_iterable(iterable, chunk_size): it = iter(iterable) while True: chunk = tuple(islice(it, chunk_size)) if not chunk: break yield chunk def json_chunked_iterable_simple(json_path, chunk_size): with open(json_path, 'r') as f: full_data = json.load(f) yield from chunked_iterable(full_data, chunk_size) # Usage example for chunk in json_chunked_iterable_simple("small_data.json", chunk_size=3): print(chunk)
Pick the streaming version for large files, and the simple version when you want minimal code and memory isn't a constraint.
内容的提问来源于stack exchange,提问作者GollyJer
相关产品推荐
相关产品推荐

