Python3下高效加载72个大JSON文件并分配变量的方案
Hey there! I totally get how tedious it is to manually write 72 load statements, and that 5GB merged file taking 12 hours is a total nightmare. Let's fix this with a clean, efficient approach that avoids both headaches.
Step 1: Batch Load All JSON Files
Instead of creating 72 separate variables (which is messy and impossible to maintain), use glob to automatically find all your target files and load them into a list. This cuts out all manual repetition in one go:
import json import glob import pandas as pd # Match all your netatmo JSON files in the target directory file_pattern = "/Users/Data/netatmo_20171231_*.json" all_file_paths = glob.glob(file_pattern) # Store each JSON's data in a list (way better than 72 individual variables!) json_data_list = [] for file_path in all_file_paths: with open(file_path, "r") as f: # Load the JSON data and add it to our list json_data = json.load(f) json_data_list.append(json_data)
Step 2: Flatten Nested JSON & Merge into a DataFrame
Pandas has a built-in tool json_normalize (or pd.json_normalize in newer versions) that handles nested JSON structures perfectly. We'll loop through our loaded data, flatten each entry, then combine everything into one cohesive DataFrame:
# Initialize a list to hold flattened DataFrames flattened_dfs = [] for data in json_data_list: # Flatten the nested JSON - adjust parameters if your structure has specific nested arrays # For example, if your data is wrapped in a "records" key, use: pd.json_normalize(data, record_path="records") flattened_df = pd.json_normalize(data) flattened_dfs.append(flattened_df) # Combine all flattened DataFrames into one final_df = pd.concat(flattened_dfs, ignore_index=True)
Why This Works Better
- No manual repetition:
globfinds all your files automatically, so you don't have to write 72openstatements by hand. - Memory-friendly: Loading files one at a time and processing them in small chunks avoids the 5GB file bottleneck. Each file is handled individually, so you won't hit that same 12-hour load time.
- Maintainable: Using lists instead of 72 scattered variables makes it easy to adjust your code later (like adding more files or tweaking processing steps).
Quick Tip for Custom Nested Structures
If your JSON has a specific nested array you need to extract (e.g., each file has a measurements key with the actual time-series data), tweak the json_normalize call to target that path explicitly:
# Example for a nested array under "measurements", with metadata fields flattened_df = pd.json_normalize( data, record_path="measurements", meta=["device_id", "collection_timestamp"] )
内容的提问来源于stack exchange,提问作者Aly Noyola

