You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何转换NomisWeb英国人口JSON-stat数据为可分析格式并完成指定分析?

解决NomisWeb JSON-stat数据转换与分析问题

Hey there, let's tackle this JSON-stat data conversion issue step by step. The key thing to remember about JSON-stat format is that the value array is a flattened representation of all combinations of the dimensions listed in dimension. So we need to map each value to its corresponding dimension categories first, then we can build a meaningful DataFrame for analysis.


Step 1: Load and Parse the JSON-stat Data

First, let's expand your existing code to extract dimension details and map them to the values properly:

import requests
import json
import pandas as pd
from itertools import product

# Download and load the data
json_url = 'https://www.nomisweb.co.uk/api/v01/dataset/NM_31_1.jsonstat.json'
response = requests.get(json_url)
content = json.loads(response.content)

# Extract core dimension info
dimension_ids = content['id']  # Should be ['GEOGRAPHY', 'SEX', 'AGE', 'DATE', 'MEASURES']
dimensions = content['dimension']

# Create lists of labels for each dimension (in the correct order)
dimension_label_lists = []
for dim_id in dimension_ids:
    dim = dimensions[dim_id]
    # Pull category labels in the order defined by the dimension's index
    category_labels = [dim['category']['label'][idx] for idx in dim['category']['index']]
    dimension_label_lists.append(category_labels)

# Generate all possible combinations of dimension categories
dim_combinations = list(product(*dimension_label_lists))

# Combine combinations with values to build a structured DataFrame
df = pd.DataFrame(dim_combinations, columns=dimension_ids)
df['POPULATION'] = content['value']

# Clean up rows with missing population values
df = df.dropna(subset=['POPULATION'])

This code uses itertools.product to generate every possible combination of the dimension labels (geography, sex, age, date, measure), then pairs each combination with the corresponding value from the value array. The result is a DataFrame where each row represents a unique, context-rich population observation.


Step 2: Task 1 - Latest Year Population by Region and Sex

Let's pull the most recent year's data to create a clear table of regional population split by sex and total:

# First, confirm the latest available year
latest_year = df['DATE'].max()
print(f"Latest available year: {latest_year}")

# Filter for latest year and total population measure (adjust label if needed)
# First check what measures are available:
print("Available measures:", df['MEASURES'].unique())

latest_pop_data = df[
    (df['DATE'] == latest_year) &
    (df['MEASURES'] == 'Total population')
].copy()

# Pivot the data to show sex as columns for readability
latest_pop_table = latest_pop_data.pivot(
    index='GEOGRAPHY',
    columns='SEX',
    values='POPULATION'
).reset_index()

# Rename columns for clarity
latest_pop_table.columns = ['Region', 'Male', 'Female', 'Total']

# Display the table
print("\nLatest Year Population by Region:")
print(latest_pop_table)

To visualize population changes over time, we can create line charts for selected regions and age groups. We'll use seaborn for clean, readable plots:

import seaborn as sns
import matplotlib.pyplot as plt

# Filter data to focus on total population
trend_data = df[df['MEASURES'] == 'Total population'].copy()

# Convert year strings to datetime for proper plotting
trend_data['DATE'] = pd.to_datetime(trend_data['DATE'], format='%Y')

# Set up plot parameters
plt.figure(figsize=(14, 8))

# Pick a few regions to avoid clutter (adjust based on your interests)
selected_regions = ['United Kingdom', 'London', 'North West', 'Scotland']
filtered_trends = trend_data[trend_data['GEOGRAPHY'].isin(selected_regions)]

# Create line plot with region as hue and age group as style
sns.lineplot(
    data=filtered_trends,
    x='DATE',
    y='POPULATION',
    hue='GEOGRAPHY',
    style='AGE',
    marker='o',
    linewidth=2
)

# Format plot for readability
plt.title('Population Trends by Region and Age Group (1981-2017)', fontsize=16)
plt.xlabel('Year', fontsize=12)
plt.ylabel('Population', fontsize=12)
plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left', fontsize=10)
plt.grid(alpha=0.3)
plt.tight_layout()
plt.show()

This plot will let you easily spot how different age groups in specific regions have grown or shrunk over the 36-year period. You can tweak the selected_regions list or filter by specific age groups to dive deeper into subsets of the data.


Quick Notes

  • Always double-check dimension labels (like MEASURES or SEX) by printing their unique values first—exact labels might vary slightly from the assumptions here.
  • If you're working with a large subset of the data, filter early (e.g., only keep the regions/age groups you care about) to reduce memory usage.
  • For missing values in the POPULATION column, the dropna step ensures you only work with valid observations.

内容的提问来源于stack exchange,提问作者dwalker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:08:02