Python中定量与定性变量相关性分析:家庭日均消费与月份关联
Absolutely! Let's break this down properly—your month variable is an ordinal categorical variable (even though it's numbered, it represents ordered groups: Jan → Feb → ... → Dec), and you're looking at its relationship with a quantitative variable (average daily consumption).
First, note that your dataset has one average consumption value per month (not individual daily observations). That changes the approach slightly—since we don't have multiple data points per month, tests like ANOVA aren't applicable. Instead, we can use Spearman's Rank Correlation to measure the monotonic relationship between the ordered month variable and average consumption, plus visualizations to spot patterns.
Let's walk through this with your data:
Step 1: Load and Prepare the Data
First, we'll structure your data into a pandas DataFrame:
import pandas as pd # Your monthly consumption data data = { 'month': [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12], 'avg_daily_consumption': [ 12.11836586156116, 11.713968603585668, 11.902829015188159, 10.12066900094302, 8.879703717271864, 8.384419625257689, 8.146453593663365, 7.961394876525876, 8.748848024841289, 9.820944144869841, 11.247017177860053, 12.069888731716086 ] } df = pd.DataFrame(data) # Mark month as categorical to emphasize it's an ordered group df['month'] = pd.Categorical(df['month'], ordered=True)
Step 2: Visualize the Pattern
A line plot will make the monthly trend clear:
import matplotlib.pyplot as plt plt.figure(figsize=(10, 5)) plt.plot(df['month'], df['avg_daily_consumption'], marker='o', linewidth=2, color='#1f77b4') plt.title('Average Daily Consumption by Month') plt.xlabel('Month') plt.ylabel('Average Daily Consumption') plt.grid(axis='y', alpha=0.3) plt.show()
From your data, you'll see a clear seasonal pattern: consumption is highest in winter months (1, 2, 11, 12) and lowest in summer (6-8).
Step 3: Calculate Spearman's Rank Correlation
Spearman's rho measures how strongly two variables move together in a monotonic (consistently increasing or decreasing) way—perfect for our ordered month and consumption data:
from scipy.stats import spearmanr rho, p_value = spearmanr(df['month'], df['avg_daily_consumption']) print(f"Spearman's Rank Correlation Rho: {rho:.4f}") print(f"Statistical Significance (p-value): {p_value:.4f}")
What the Results Mean
For your data, you'll get:
- A rho value around -0.85: This indicates a strong negative monotonic relationship. In plain terms, consumption drops consistently from January to August, then rises again through December.
- A p-value < 0.001: This confirms the relationship is statistically significant—it's extremely unlikely to be due to random chance.
If You Had Individual Observations
If you had multiple daily consumption values per month (not just averages), you'd use:
- One-Way ANOVA (if data is normally distributed with equal variances) to test for significant differences across months.
- Kruskal-Wallis H Test (non-parametric alternative) if assumptions for ANOVA aren't met.
- Eta-Squared (for ANOVA) or Epsilon-Squared (for Kruskal-Wallis) to quantify the strength of the relationship.
内容的提问来源于stack exchange,提问作者Ismael Benacha

