You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

categorical变量City与Rainfall的corr分析及重复城市数据处理问询

Hey there, let's walk through how to tackle this analysis step by step—since you're working with repeated daily weather data across cities, we need to use methods tailored for categorical vs continuous variables, plus leverage the time-series aspect of your data.

1. First: Clean & Prep Your Data

Before diving into analysis, make sure your data is ready to work with:

  • Fix data types: Convert the Date column to datetime format, and ensure Location is treated as a categorical variable:
    import pandas as pd
    
    df['Date'] = pd.to_datetime(df['Date'])
    df['Location'] = df['Location'].astype('category')
    
  • Handle missing values: Check for missing Rainfall entries. You can drop rows with missing data (if the count is small) or fill missing values with the city's average rainfall:
    # Drop rows with missing Rainfall
    df = df.dropna(subset=['Rainfall'])
    
    # OR fill missing values with city-specific mean
    df['Rainfall'] = df.groupby('Location')['Rainfall'].transform(lambda x: x.fillna(x.mean()))
    
2. Compare Rainfall Across Cities (Exploratory Analysis)

Start with simple stats and visuals to get a sense of differences between cities:

  • Calculate city-level rainfall stats: Group by Location to get key metrics like average, median, and variability:
    city_rain_summary = df.groupby('Location')['Rainfall'].agg(
        ['mean', 'median', 'std', 'min', 'max', 'count']
    ).sort_values('mean', ascending=False)
    print(city_rain_summary)
    
  • Visualize distributions: Use boxplots to compare rainfall spread and outliers across cities—this is way more intuitive than numbers alone:
    import seaborn as sns
    import matplotlib.pyplot as plt
    
    plt.figure(figsize=(12, 6))
    sns.boxplot(x='Location', y='Rainfall', data=df)
    plt.xticks(rotation=45)
    plt.title('Rainfall Distribution by City')
    plt.ylabel('Rainfall (mm)')
    plt.show()
    
  • Bonus: Rainy day frequency: Stats on how often it rains in each city can be just as insightful as total rainfall:
    city_rainy_days = df.groupby('Location').apply(
        lambda x: round((x['Rainfall'] > 0).mean() * 100, 2)
    ).reset_index(name='Rainy_Day_Percentage')
    print(city_rainy_days.sort_values('Rainy_Day_Percentage', ascending=False))
    
3. Measure "Correlation" Between Categorical Location & Rainfall

You can’t use Pearson correlation here (that’s for two continuous variables). Instead, use these methods to quantify the relationship:

  • ANOVA Test: Checks if there’s a statistically significant difference in average rainfall across cities. A small p-value (<0.05) means at least one city’s rainfall is significantly different from others:
    import scipy.stats as stats
    
    # Extract rainfall data for each city as separate groups
    city_rain_groups = [group['Rainfall'].values for name, group in df.groupby('Location')]
    
    # Run one-way ANOVA
    f_stat, p_value = stats.f_oneway(*city_rain_groups)
    print(f"ANOVA F-statistic: {f_stat:.2f}, p-value: {p_value:.4f}")
    
  • Eta Squared (η²): If you want a single number (like a correlation score) to measure how much variance in rainfall is explained by location. Values range from 0 (no association) to 1 (perfect association):
    total_variance = sum((df['Rainfall'] - df['Rainfall'].mean()) ** 2)
    between_city_variance = sum(
        city_rain_summary['count'] * (city_rain_summary['mean'] - df['Rainfall'].mean()) ** 2
    )
    eta_squared = between_city_variance / total_variance
    print(f"Eta Squared (Location vs Rainfall): {eta_squared:.4f}")
    
  • Non-parametric alternative (Kruskal-Wallis Test): Use this if your rainfall data isn’t normally distributed (test with Shapiro-Wilk if unsure)—it works without assuming normality:
    h_stat, p_value = stats.kruskal(*city_rain_groups)
    print(f"Kruskal-Wallis H-statistic: {h_stat:.2f}, p-value: {p_value:.4f}")
    
4. Leverage the Time Dimension (Optional but Powerful)

Since you have daily data, you can dig into temporal patterns across cities:

  • Monthly/seasonal rainfall trends: See how rainfall varies by month for each city:
    df['Month'] = df['Date'].dt.month
    monthly_city_rain = df.groupby(['Location', 'Month'])['Rainfall'].mean().unstack()
    
    monthly_city_rain.plot(kind='line', figsize=(12, 6))
    plt.title('Monthly Average Rainfall by City')
    plt.ylabel('Average Rainfall (mm)')
    plt.xlabel('Month')
    plt.legend(title='City', bbox_to_anchor=(1.05, 1), loc='upper left')
    plt.show()
    

内容的提问来源于stack exchange,提问作者Chukas Ebuka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:51:29