如何在Pandas中将嵌套列表转为DataFrame及处理API空气质量JSON数据
Got it, let's walk through how to turn this nested air quality API response into a clean Pandas DataFrame. First, let's recap your data structure to make sure we're on the same page:
Your sample data is a list (
data[0]) containing a dictionary, where the key94103maps to a list of air quality records. Each record has a nestedCategorydictionary withNameandNumberfields.
Step 1: Extract the core data list
First, we need to pull out the actual list of air quality entries from that nested structure. Assuming your raw API response is stored in a variable called data, do this:
# Extract the list of air quality records air_quality_records = data[0]['94103']
Step 2: Convert to DataFrame (with nested fields expanded)
The easiest way to handle nested JSON structures like this is using Pandas' pd.json_normalize() function—it's built specifically for this use case.
Method 1: Use json_normalize (Recommended)
This function automatically flattens nested dictionaries into separate columns. You can even customize the separator for nested field names to make them more readable:
import pandas as pd # Flatten the nested data and create DataFrame df = pd.json_normalize(air_quality_records, sep='_')
This will give you columns like AQI, Category_Name, Category_Number, DateObserved, HourObserved, etc.—no manual parsing needed!
Method 2: Manual Parsing (If you want more control)
If you prefer to handle the nested fields manually, you can loop through each record, expand the Category data, and build a new list of processed dictionaries:
processed_records = [] for record in air_quality_records: # Make a copy of the original record to avoid modifying raw data processed = record.copy() # Extract nested Category fields into separate keys processed['Category_Name'] = record['Category']['Name'] processed['Category_Number'] = record['Category']['Number'] # Remove the original nested Category dictionary del processed['Category'] processed_records.append(processed) # Convert the processed list to DataFrame df = pd.DataFrame(processed_records)
Step 3: Clean up extra formatting (Optional)
Notice your DateObserved field has trailing spaces? You can clean that up easily with a string operation:
df['DateObserved'] = df['DateObserved'].str.strip()
Final Result
Your DataFrame will now have all the air quality data in a flat, easy-to-analyze format—perfect for filtering, plotting, or running further analysis!
内容的提问来源于stack exchange,提问作者HP-Nunes

