You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从DataFrame名称提取年份并关联数据计算相关系数

如何自动提取DataFrame变量名末尾的年份?

问题核心

直接将DataFrame转为字符串(str(df))无法获取变量名,因为该操作返回的是DataFrame的对象描述或内容预览,不是定义时的变量名——这也是你之前得到'm'而非年份的根本原因。

解决方案

方法1:字典映射(推荐,清晰可控)

把变量名和对应的DataFrame存入字典,循环时直接从字典键中提取年份:

import pandas as pd

# 基于已有DataFrame构建字典,无需重新赋值原变量
df_dict = {
    'grid_2010': grid_2010,
    'grid_2018': grid_2018,
    'grid_2022': grid_2022
}

for name, df in df_dict.items():
    year = name[-4:]  # 截取变量名最后4位作为年份
    print(f'{year}: col1 vs. col2')
    print(df['Col1'].corr(df['Col2']))

这种方法只是将现有变量组织到字典中,完全保留原变量的定义,逻辑清晰且不易出错。

方法2:利用作用域变量自动筛选(适合变量较多的场景)

通过locals()或globals()获取当前作用域的所有变量,筛选出目标DataFrame并提取变量名:

import pandas as pd

# 遍历当前局部作用域的变量
for name, df in locals().items():
    # 筛选出以'grid_'开头且是DataFrame类型的变量
    if name.startswith('grid_') and isinstance(df, pd.DataFrame):
        year = name[-4:]
        print(f'{year}: col1 vs. col2')
        print(df['Col1'].corr(df['Col2']))

注意:如果当前作用域中有其他以'grid_'开头的非DataFrame变量,需要调整过滤条件避免误选。

错误原因分析

你之前尝试的str(df)返回的是DataFrame的对象信息(如<pandas.core.frame.DataFrame object at 0x7fxxxx>)或内容预览,其字符串末尾的字符和变量名完全无关,所以截取后得到的是无意义的'm',这个思路从根本上走不通。

内容的提问来源于stack exchange,提问作者Vanessa_C

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 15:35:25