如何转换NomisWeb英国人口JSON-stat数据为可分析格式并完成指定分析?
Hey there, let's tackle this JSON-stat data conversion issue step by step. The key thing to remember about JSON-stat format is that the value array is a flattened representation of all combinations of the dimensions listed in dimension. So we need to map each value to its corresponding dimension categories first, then we can build a meaningful DataFrame for analysis.
Step 1: Load and Parse the JSON-stat Data
First, let's expand your existing code to extract dimension details and map them to the values properly:
import requests import json import pandas as pd from itertools import product # Download and load the data json_url = 'https://www.nomisweb.co.uk/api/v01/dataset/NM_31_1.jsonstat.json' response = requests.get(json_url) content = json.loads(response.content) # Extract core dimension info dimension_ids = content['id'] # Should be ['GEOGRAPHY', 'SEX', 'AGE', 'DATE', 'MEASURES'] dimensions = content['dimension'] # Create lists of labels for each dimension (in the correct order) dimension_label_lists = [] for dim_id in dimension_ids: dim = dimensions[dim_id] # Pull category labels in the order defined by the dimension's index category_labels = [dim['category']['label'][idx] for idx in dim['category']['index']] dimension_label_lists.append(category_labels) # Generate all possible combinations of dimension categories dim_combinations = list(product(*dimension_label_lists)) # Combine combinations with values to build a structured DataFrame df = pd.DataFrame(dim_combinations, columns=dimension_ids) df['POPULATION'] = content['value'] # Clean up rows with missing population values df = df.dropna(subset=['POPULATION'])
This code uses itertools.product to generate every possible combination of the dimension labels (geography, sex, age, date, measure), then pairs each combination with the corresponding value from the value array. The result is a DataFrame where each row represents a unique, context-rich population observation.
Step 2: Task 1 - Latest Year Population by Region and Sex
Let's pull the most recent year's data to create a clear table of regional population split by sex and total:
# First, confirm the latest available year latest_year = df['DATE'].max() print(f"Latest available year: {latest_year}") # Filter for latest year and total population measure (adjust label if needed) # First check what measures are available: print("Available measures:", df['MEASURES'].unique()) latest_pop_data = df[ (df['DATE'] == latest_year) & (df['MEASURES'] == 'Total population') ].copy() # Pivot the data to show sex as columns for readability latest_pop_table = latest_pop_data.pivot( index='GEOGRAPHY', columns='SEX', values='POPULATION' ).reset_index() # Rename columns for clarity latest_pop_table.columns = ['Region', 'Male', 'Female', 'Total'] # Display the table print("\nLatest Year Population by Region:") print(latest_pop_table)
Step 3: Task 2 - Exploratory Analysis: Population Trends by Region and Age Group
To visualize population changes over time, we can create line charts for selected regions and age groups. We'll use seaborn for clean, readable plots:
import seaborn as sns import matplotlib.pyplot as plt # Filter data to focus on total population trend_data = df[df['MEASURES'] == 'Total population'].copy() # Convert year strings to datetime for proper plotting trend_data['DATE'] = pd.to_datetime(trend_data['DATE'], format='%Y') # Set up plot parameters plt.figure(figsize=(14, 8)) # Pick a few regions to avoid clutter (adjust based on your interests) selected_regions = ['United Kingdom', 'London', 'North West', 'Scotland'] filtered_trends = trend_data[trend_data['GEOGRAPHY'].isin(selected_regions)] # Create line plot with region as hue and age group as style sns.lineplot( data=filtered_trends, x='DATE', y='POPULATION', hue='GEOGRAPHY', style='AGE', marker='o', linewidth=2 ) # Format plot for readability plt.title('Population Trends by Region and Age Group (1981-2017)', fontsize=16) plt.xlabel('Year', fontsize=12) plt.ylabel('Population', fontsize=12) plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left', fontsize=10) plt.grid(alpha=0.3) plt.tight_layout() plt.show()
This plot will let you easily spot how different age groups in specific regions have grown or shrunk over the 36-year period. You can tweak the selected_regions list or filter by specific age groups to dive deeper into subsets of the data.
Quick Notes
- Always double-check dimension labels (like
MEASURESorSEX) by printing their unique values first—exact labels might vary slightly from the assumptions here. - If you're working with a large subset of the data, filter early (e.g., only keep the regions/age groups you care about) to reduce memory usage.
- For missing values in the
POPULATIONcolumn, thedropnastep ensures you only work with valid observations.
内容的提问来源于stack exchange,提问作者dwalker

