Python嵌套JSON转CSV:新手求助大型JSON文件转换方案
Hey there! Since you're new to Python and need to convert a large JSON file to CSV, let's break this down step by step. Your sample JSON has nested structures (like the Signatories array), so we need to handle that first—CSV works best with flat, tabular data. I'll cover two approaches: a native Python method (no extra libraries) and a simpler pandas-based method (perfect for beginners).
Method 1: Native Python (No Third-Party Libraries)
This is great if you don't want to install extra tools, or need to handle extremely large files with minimal memory usage.
Step 1: Write a JSON Flattening Function
We need to turn nested keys (like Signatories.Name) into flat column names (like Signatories_Name). Here's a function to do that:
def flatten_json(nested_json, parent_key='', sep='_'): items = [] for k, v in nested_json.items(): new_key = f"{parent_key}{sep}{k}" if parent_key else k if isinstance(v, dict): # Recursively flatten nested dictionaries items.extend(flatten_json(v, new_key, sep=sep).items()) elif isinstance(v, list): # Handle arrays: this example takes the first element (adjust if you need all) if v: items.extend(flatten_json(v[0], new_key, sep=sep).items()) else: items.append((new_key, None)) else: items.append((new_key, v)) return dict(items)
Step 2: Read and Process the JSON File
For a single large JSON object:
import json import csv # Load the JSON file (use streaming if it's *extremely* large—see note below) with open('input.json', 'r', encoding='utf-8') as f: raw_data = json.load(f) # Flatten the nested data flattened_data = flatten_json(raw_data) # Write to CSV with open('output.csv', 'w', newline='', encoding='utf-8') as csv_file: writer = csv.DictWriter(csv_file, fieldnames=flattened_data.keys()) writer.writeheader() writer.writerow(flattened_data)
For a JSON array of multiple objects:
If your file looks like [{}, {}, ...], loop through each object:
import json import csv with open('input.json', 'r', encoding='utf-8') as f: raw_data = json.load(f) # List of JSON objects # Flatten every item in the array flattened_list = [flatten_json(item) for item in raw_data] # Write all rows to CSV if flattened_list: headers = flattened_list[0].keys() with open('output.csv', 'w', newline='', encoding='utf-8') as csv_file: writer = csv.DictWriter(csv_file, fieldnames=headers) writer.writeheader() writer.writerows(flattened_list)
Streaming for Ultra-Large Files
If loading the entire file crashes your system, use the ijson library (install first with pip install ijson) to read nodes one at a time:
import ijson import csv with open('input.json', 'r', encoding='utf-8') as f, open('output.csv', 'w', newline='', encoding='utf-8') as csv_file: # Read items from the JSON array (adjust 'item' to match your structure) items = ijson.items(f, 'item') # Process the first item to set headers first_item = next(items) flattened_first = flatten_json(first_item) writer = csv.DictWriter(csv_file, fieldnames=flattened_first.keys()) writer.writeheader() writer.writerow(flattened_first) # Process remaining items for item in items: writer.writerow(flatten_json(item))
Method 2: Pandas (Simpler for Beginners)
Pandas has a built-in json_normalize function that handles flattening automatically. It's way less code! First install pandas:
pip install pandas
Basic Conversion Code
import pandas as pd import json # Load the JSON data with open('input.json', 'r', encoding='utf-8') as f: raw_data = json.load(f) # If it's a single object, convert it to a list for pandas if isinstance(raw_data, dict): raw_data = [raw_data] # Flatten and convert to DataFrame df = pd.json_normalize(raw_data) # Save to CSV (skip index to avoid extra columns) df.to_csv('output.csv', index=False, encoding='utf-8')
Handling Ultra-Large Files with Pandas
For files too big to fit in memory, combine pandas with streaming:
import ijson import pandas as pd from csv import DictWriter with open('input.json', 'r', encoding='utf-8') as f, open('output.csv', 'w', newline='', encoding='utf-8') as csv_file: items = ijson.items(f, 'item') first_item = next(items) flattened_first = pd.json_normalize(first_item).to_dict('records')[0] # Set up CSV writer writer = DictWriter(csv_file, fieldnames=flattened_first.keys()) writer.writeheader() writer.writerow(flattened_first) # Process in chunks to save memory chunk = [] chunk_size = 1000 # Adjust based on your system's memory for item in items: chunk.append(pd.json_normalize(item).to_dict('records')[0]) if len(chunk) >= chunk_size: writer.writerows(chunk) chunk = [] # Write any remaining items if chunk: writer.writerows(chunk)
Key Notes to Remember
- Nested Arrays: If your
Signatoriesarray has multiple entries, the default code only uses the first one. To split each array entry into a separate CSV row, modify the flattening function to duplicate parent fields for each array element. - Encoding: Always specify
encoding='utf-8'when reading/writing files to avoid character issues. - Memory: For truly massive files, stick to streaming methods (native + ijson or pandas + chunking) to avoid crashing your system.
内容的提问来源于stack exchange,提问作者SS360

