如何为Pandas DataFrame中按logger分组的子集生成相关矩阵?
按Logger分组计算年份与气象指标的Pearson相关系数
你的需求是按logger分组,计算每个组内year与avg_max_temp、avg_min_temp、tot_precipitation的Pearson相关系数,最终得到包含100行(对应每个logger)和3个相关系数列的结果。
解决方案代码
import pandas as pd # 定义用于分组计算相关系数的函数 def calc_logger_correlations(group): # 仅保留需要计算相关系数的列,生成相关系数矩阵 corr_matrix = group[['year', 'avg_max_temp', 'avg_min_temp', 'tot_precipitation']].corr(method='pearson') # 提取year与三个气象指标的相关系数,返回结构化结果 return pd.Series({ 'corr_year_avg_max_temp': corr_matrix.loc['year', 'avg_max_temp'], 'corr_year_avg_min_temp': corr_matrix.loc['year', 'avg_min_temp'], 'corr_year_tot_precipitation': corr_matrix.loc['year', 'tot_precipitation'] }) # 按logger分组执行计算,重置索引后得到目标DataFrame result_df = df.groupby('logger').apply(calc_logger_correlations).reset_index() # 查看结果示例 result_df.head()
代码说明
- 分组逻辑:使用
groupby('logger')将原始DataFrame按logger ID拆分为100个独立的子数据集。 - 相关系数计算:对每个子数据集,仅选取
year和目标气象指标列计算Pearson相关系数矩阵,避免无关列(如yield)干扰结果。 - 结果结构化:从相关系数矩阵中提取
year与三个指标的相关系数,封装为Series,最终拼接成每行对应一个logger的结果DataFrame。
内容的提问来源于stack exchange,提问作者Tom
相关产品推荐
相关产品推荐

