如何在Python中基于月度数据计算季度平均值?
Hey there! Calculating quarterly averages from your monthly employment data in Python is totally straightforward with pandas—let me walk you through the process using your NSW 1978-1979 example as context.
Pandas is the go-to tool for time-series and tabular data tasks like this. First, we'll import it:
import pandas as pd
First, load your dataset (I’ll assume it’s a CSV, but adjust for Excel or other formats as needed). Then, make sure your Date column is recognized as a datetime type—this is crucial for extracting quarters later.
# Load your data (replace with your actual file path/name) df = pd.read_csv('employment_data.csv') # Convert 'Date' to datetime format (add a `format` parameter if your dates aren't standard, e.g., '%d/%m/%Y') df['Date'] = pd.to_datetime(df['Date'])
We need to group our monthly data by quarter. You can either create separate Year and Quarter columns, or a combined YearQuarter column (like 1978Q1) for clarity:
# Option 1: Create separate year and quarter columns df['Year'] = df['Date'].dt.year df['Quarter'] = df['Date'].dt.quarter # Option 2: Create a combined year-quarter column (e.g., '1978Q1') df['YearQuarter'] = df['Date'].dt.to_period('Q')
Now, group the data by State/Territory and your chosen quarter identifier, then calculate the mean for your rate columns.
Using the combined YearQuarter column:
quarterly_averages = df.groupby(['State/Territory', 'YearQuarter'])[ ['Employment Rate (15-64)', 'Unemployment Rate (15+)'] ].mean().reset_index()
Using separate Year + Quarter columns:
quarterly_averages = df.groupby(['State/Territory', 'Year', 'Quarter'])[ ['Employment Rate (15-64)', 'Unemployment Rate (15+)'] ].mean().reset_index()
The reset_index() call converts the grouped columns back to regular columns, making the result easier to read and work with.
You can print the full result or filter for a specific state like NSW:
# Print all quarterly averages print(quarterly_averages) # Print only NSW's 1978-1979 quarterly data print(quarterly_averages[quarterly_averages['State/Territory'] == 'NSW'])
Quick Note on Missing Data
If your dataset has missing values, pandas’ mean() function will automatically skip them (thanks to the default skipna=True). If you want to handle missing data differently (e.g., drop rows first), you can use df.dropna(subset=['Employment Rate (15-64)', 'Unemployment Rate (15+)']) before grouping.
内容的提问来源于stack exchange,提问作者Aayush Kapoor

