You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为Pandas DataFrame中按logger分组的子集生成相关矩阵?

按Logger分组计算年份与气象指标的Pearson相关系数

你的需求是按logger分组,计算每个组内year与avg_max_temp、avg_min_temp、tot_precipitation的Pearson相关系数,最终得到包含100行(对应每个logger)和3个相关系数列的结果。

解决方案代码

import pandas as pd

# 定义用于分组计算相关系数的函数
def calc_logger_correlations(group):
    # 仅保留需要计算相关系数的列,生成相关系数矩阵
    corr_matrix = group[['year', 'avg_max_temp', 'avg_min_temp', 'tot_precipitation']].corr(method='pearson')
    # 提取year与三个气象指标的相关系数,返回结构化结果
    return pd.Series({
        'corr_year_avg_max_temp': corr_matrix.loc['year', 'avg_max_temp'],
        'corr_year_avg_min_temp': corr_matrix.loc['year', 'avg_min_temp'],
        'corr_year_tot_precipitation': corr_matrix.loc['year', 'tot_precipitation']
    })

# 按logger分组执行计算,重置索引后得到目标DataFrame
result_df = df.groupby('logger').apply(calc_logger_correlations).reset_index()

# 查看结果示例
result_df.head()

代码说明

  1. 分组逻辑:使用groupby('logger')将原始DataFrame按logger ID拆分为100个独立的子数据集。
  2. 相关系数计算:对每个子数据集,仅选取year和目标气象指标列计算Pearson相关系数矩阵,避免无关列(如yield)干扰结果。
  3. 结果结构化:从相关系数矩阵中提取year与三个指标的相关系数,封装为Series,最终拼接成每行对应一个logger的结果DataFrame。

内容的提问来源于stack exchange,提问作者Tom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 04:22:17