categorical变量City与Rainfall的corr分析及重复城市数据处理问询
Hey there, let's walk through how to tackle this analysis step by step—since you're working with repeated daily weather data across cities, we need to use methods tailored for categorical vs continuous variables, plus leverage the time-series aspect of your data.
1. First: Clean & Prep Your Data
Before diving into analysis, make sure your data is ready to work with:
- Fix data types: Convert the
Datecolumn to datetime format, and ensureLocationis treated as a categorical variable:import pandas as pd df['Date'] = pd.to_datetime(df['Date']) df['Location'] = df['Location'].astype('category') - Handle missing values: Check for missing
Rainfallentries. You can drop rows with missing data (if the count is small) or fill missing values with the city's average rainfall:# Drop rows with missing Rainfall df = df.dropna(subset=['Rainfall']) # OR fill missing values with city-specific mean df['Rainfall'] = df.groupby('Location')['Rainfall'].transform(lambda x: x.fillna(x.mean()))
2. Compare Rainfall Across Cities (Exploratory Analysis)
Start with simple stats and visuals to get a sense of differences between cities:
- Calculate city-level rainfall stats: Group by
Locationto get key metrics like average, median, and variability:city_rain_summary = df.groupby('Location')['Rainfall'].agg( ['mean', 'median', 'std', 'min', 'max', 'count'] ).sort_values('mean', ascending=False) print(city_rain_summary) - Visualize distributions: Use boxplots to compare rainfall spread and outliers across cities—this is way more intuitive than numbers alone:
import seaborn as sns import matplotlib.pyplot as plt plt.figure(figsize=(12, 6)) sns.boxplot(x='Location', y='Rainfall', data=df) plt.xticks(rotation=45) plt.title('Rainfall Distribution by City') plt.ylabel('Rainfall (mm)') plt.show() - Bonus: Rainy day frequency: Stats on how often it rains in each city can be just as insightful as total rainfall:
city_rainy_days = df.groupby('Location').apply( lambda x: round((x['Rainfall'] > 0).mean() * 100, 2) ).reset_index(name='Rainy_Day_Percentage') print(city_rainy_days.sort_values('Rainy_Day_Percentage', ascending=False))
3. Measure "Correlation" Between Categorical Location & Rainfall
You can’t use Pearson correlation here (that’s for two continuous variables). Instead, use these methods to quantify the relationship:
- ANOVA Test: Checks if there’s a statistically significant difference in average rainfall across cities. A small p-value (<0.05) means at least one city’s rainfall is significantly different from others:
import scipy.stats as stats # Extract rainfall data for each city as separate groups city_rain_groups = [group['Rainfall'].values for name, group in df.groupby('Location')] # Run one-way ANOVA f_stat, p_value = stats.f_oneway(*city_rain_groups) print(f"ANOVA F-statistic: {f_stat:.2f}, p-value: {p_value:.4f}") - Eta Squared (η²): If you want a single number (like a correlation score) to measure how much variance in rainfall is explained by location. Values range from 0 (no association) to 1 (perfect association):
total_variance = sum((df['Rainfall'] - df['Rainfall'].mean()) ** 2) between_city_variance = sum( city_rain_summary['count'] * (city_rain_summary['mean'] - df['Rainfall'].mean()) ** 2 ) eta_squared = between_city_variance / total_variance print(f"Eta Squared (Location vs Rainfall): {eta_squared:.4f}") - Non-parametric alternative (Kruskal-Wallis Test): Use this if your rainfall data isn’t normally distributed (test with Shapiro-Wilk if unsure)—it works without assuming normality:
h_stat, p_value = stats.kruskal(*city_rain_groups) print(f"Kruskal-Wallis H-statistic: {h_stat:.2f}, p-value: {p_value:.4f}")
4. Leverage the Time Dimension (Optional but Powerful)
Since you have daily data, you can dig into temporal patterns across cities:
- Monthly/seasonal rainfall trends: See how rainfall varies by month for each city:
df['Month'] = df['Date'].dt.month monthly_city_rain = df.groupby(['Location', 'Month'])['Rainfall'].mean().unstack() monthly_city_rain.plot(kind='line', figsize=(12, 6)) plt.title('Monthly Average Rainfall by City') plt.ylabel('Average Rainfall (mm)') plt.xlabel('Month') plt.legend(title='City', bbox_to_anchor=(1.05, 1), loc='upper left') plt.show()
内容的提问来源于stack exchange,提问作者Chukas Ebuka
相关产品推荐
相关产品推荐

