如何使用次Y轴绘制带分组索引的DataFrame
Got it, let's break this down step by step—first we'll fix up the data loading code you started, then create a plot with a secondary Y-axis for your grouped WHO DataFrame. Here's how to do it:
Step 1: Complete Data Loading & Preprocessing
Your initial code cut off before finishing the JSON normalization, so let's fill that in and clean the data to get a usable grouped DataFrame:
import pandas as pd import urllib2 import json # Fetch and parse the WHO data url = "http://apps.who.int/gho/athena/data/GHO/MORT_100.json?profile=simple&filter=COUNTRY:*;CHILDCAUSE:CH6" response2 = urllib2.urlopen(url) response_json2 = json.loads(response2.read()) # Normalize the core data (stored in the 'fact' section of the JSON) dfWHO2 = pd.json_normalize(response_json2['fact']) # Map dimension codes to human-readable labels (e.g., country names instead of codes) dimensions = pd.json_normalize(response_json2['dimension']) dim_label_maps = {} for _, dim in dimensions.iterrows(): dim_code = dim['code'] dim_label_maps[dim_code] = {cat['code']: cat['label'] for cat in dim['category']['category']} # Add readable columns and clean numeric data dfWHO2['COUNTRY'] = dfWHO2['dim.COUNTRY'].map(dim_label_maps['COUNTRY']) dfWHO2['YEAR'] = dfWHO2['dim.YEAR'].map(dim_label_maps['YEAR']).astype(int) dfWHO2['UNDER5_MEASLES_DEATHS'] = dfWHO2['Value'].astype(float) # Set the grouped index (Country + Year) as requested dfWHO2.set_index(['COUNTRY', 'YEAR'], inplace=True) # Optional: Filter to a smaller set of countries for clearer plotting (adjust as needed) selected_countries = ['India', 'Nigeria', 'Democratic Republic of the Congo'] df_filtered = dfWHO2.loc[selected_countries, ['UNDER5_MEASLES_DEATHS']].unstack(level=0)
Step 2: Plot with Secondary Y-Axis
We'll use matplotlib to create a plot where the primary Y-axis shows absolute death counts, and the secondary Y-axis shows annual percentage changes (a common use case for dual axes). You can swap the secondary axis metric if you have another value to compare:
import matplotlib.pyplot as plt # Initialize figure and primary axis fig, ax1 = plt.subplots(figsize=(12, 6)) # Plot absolute death counts on primary Y-axis color = 'tab:blue' ax1.set_xlabel('Year') ax1.set_ylabel('Under-5 Measles Deaths', color=color) for country in selected_countries: ax1.plot(df_filtered.index, df_filtered['UNDER5_MEASLES_DEATHS'][country], marker='o', label=country) ax1.tick_params(axis='y', labelcolor=color) ax1.legend(loc='upper left') # Create secondary Y-axis for percentage change ax2 = ax1.twinx() color = 'tab:red' ax2.set_ylabel('Annual % Change in Deaths', color=color) # Calculate year-over-year percentage change for each country for country in selected_countries: pct_change = df_filtered['UNDER5_MEASLES_DEATHS'][country].pct_change() * 100 ax2.plot(df_filtered.index, pct_change, marker='s', linestyle='--', label=f'{country} % Change') ax2.tick_params(axis='y', labelcolor=color) ax2.legend(loc='upper right') # Finalize plot layout and title plt.title('Measles Deaths in Under-5 Children (Absolute Counts & Yearly Change)') fig.tight_layout() plt.show()
Key Details:
- Grouped Index Handling: We set
['COUNTRY', 'YEAR']as a multi-index, then useunstack()to reshape the data so years are on the x-axis and each country is a separate series—this makes plotting grouped data straightforward. - Secondary Y-Axis: The
twinx()method creates a second axis that shares the x-axis, perfect for comparing metrics on different scales (like raw counts vs. percentage changes). - Readability: We use distinct colors, markers, and separate legends to keep the two axes' data clear.
If you have a different secondary metric (e.g., death rates per 1000 live births), just replace the percentage change calculation with your target column from the DataFrame.
内容的提问来源于stack exchange,提问作者kiltannen

