You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

导入process_excel函数后无法生成命名DataFrames的问题求助

问题:导入函数后无法生成命名DataFrames,仅返回无名称列表

我在function_general.py中编写了process_excel函数,该函数接收Excel文件路径、工作表名和字典列表作为输入,会生成与字典列表长度一致、按列表内容命名的DataFrames。在function_general.py内运行该函数可得到多个命名DataFrames,但将其导入其他文件调用时,仅返回无名称的DataFrame列表。使用Spyder环境,将函数代码复制到新文件中运行可得到预期结果,但希望通过导入方式使用该函数,寻求解决办法。

原函数代码

import pandas as pd

def process_excel(excel_file_path, sheet_name, result_list):
    df = pd.read_excel(excel_file_path, sheet_name, header=3)
    df['GC Injection'] = pd.to_numeric(df['GC Injection'], errors='coerce')
    df = df[~df['GC Injection'].isin(['Maximum', 'Average', 'Minimum', 'Standard Deviation', 'Relative Standard Deviation'])]
    variable = [b.get("experiment").replace(' ','_') + "_data" for b in result_list]
    
    for index, value in enumerate(result_list):
        dhp_test = df[(df['GC Injection'] >= value['Start Injection Number']) & 
                      (df['GC Injection'] <= value['End Injection Number'])]
        dhp_test = dhp_test.reset_index(drop=True)

        # Adding the 'minutes reaction' column
        dhp_test['minutes reaction'] = value['Start Minutes'] + 10 * dhp_test.index
        dhp_test['minutes reaction'] = dhp_test.pop(dhp_test.columns[-1])
        dhp_test.insert(3,"minutes",dhp_test['minutes reaction'])
        dhp_test.drop(['min'],axis = 1, inplace = True)
        dhp_test.drop(['minutes reaction'],axis = 1, inplace = True)       

        globals()[variable[index]] = dhp_test
    
    return [globals()[var] for var in variable]

导入调用代码

from function_general import process_excel
result_dataframes = process_excel(excel_file_path, "GC", injections_list)

问题原因

核心问题出在globals()[variable[index]] = dhp_test这行代码:globals()获取的是函数所在模块(function_general.py)的全局命名空间,而非调用该函数的文件的命名空间。因此,当你在其他文件导入调用时,这些命名好的DataFrames只会存在于function_general模块的全局空间中,调用者的文件无法直接访问,只能拿到函数返回的无名称列表。


解决方案

方案1:返回字典(推荐)

将函数修改为返回一个字典,键为DataFrame的命名,值为对应的DataFrame。这种方式符合Python函数设计原则,无副作用,且调用者能灵活访问数据。

修改后的函数代码:

import pandas as pd

def process_excel(excel_file_path, sheet_name, result_list):
    df = pd.read_excel(excel_file_path, sheet_name, header=3)
    df['GC Injection'] = pd.to_numeric(df['GC Injection'], errors='coerce')
    df = df[~df['GC Injection'].isin(['Maximum', 'Average', 'Minimum', 'Standard Deviation', 'Relative Standard Deviation'])]
    
    result_dict = {}
    for value in result_list:
        # 生成DataFrame的名称
        exp_name = value.get("experiment").replace(' ','_') + "_data"
        # 筛选数据
        dhp_test = df[(df['GC Injection'] >= value['Start Injection Number']) & 
                      (df['GC Injection'] <= value['End Injection Number'])]
        dhp_test = dhp_test.reset_index(drop=True)

        # 处理列逻辑
        dhp_test['minutes reaction'] = value['Start Minutes'] + 10 * dhp_test.index
        dhp_test.insert(3, "minutes", dhp_test['minutes reaction'])
        dhp_test.drop(['min', 'minutes reaction'], axis=1, inplace=True)       

        # 将DataFrame存入字典
        result_dict[exp_name] = dhp_test
    
    return result_dict

调用方式:

from function_general import process_excel
result_dataframes = process_excel(excel_file_path, "GC", injections_list)

# 通过键名访问指定DataFrame
print(result_dataframes['你的实验名称_data'])

方案2:传入调用者的全局命名空间(不推荐)

如果一定要在调用者的命名空间中直接生成命名变量,可以将调用者的globals()作为参数传入函数。这种方式耦合性高,容易引发命名冲突,不推荐使用。

修改后的函数代码:

import pandas as pd

def process_excel(excel_file_path, sheet_name, result_list, caller_globals=None):
    df = pd.read_excel(excel_file_path, sheet_name, header=3)
    df['GC Injection'] = pd.to_numeric(df['GC Injection'], errors='coerce')
    df = df[~df['GC Injection'].isin(['Maximum', 'Average', 'Minimum', 'Standard Deviation', 'Relative Standard Deviation'])]
    variable = [b.get("experiment").replace(' ','_') + "_data" for b in result_list]
    
    result_dfs = []
    for index, value in enumerate(result_list):
        dhp_test = df[(df['GC Injection'] >= value['Start Injection Number']) & 
                      (df['GC Injection'] <= value['End Injection Number'])]
        dhp_test = dhp_test.reset_index(drop=True)

        # 处理列逻辑
        dhp_test['minutes reaction'] = value['Start Minutes'] + 10 * dhp_test.index
        dhp_test.insert(3, "minutes", dhp_test['minutes reaction'])
        dhp_test.drop(['min', 'minutes reaction'], axis=1, inplace=True)       

        result_dfs.append(dhp_test)
        # 若传入调用者的globals,则写入变量
        if caller_globals is not None:
            caller_globals[variable[index]] = dhp_test
    
    return result_dfs

调用方式:

from function_general import process_excel
result_dataframes = process_excel(excel_file_path, "GC", injections_list, caller_globals=globals())

# 直接使用命名好的变量
print(你的实验名称_data)

内容的提问来源于stack exchange,提问作者RB_student

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 08:58:10