导入process_excel函数后无法生成命名DataFrames的问题求助
问题:导入函数后无法生成命名DataFrames,仅返回无名称列表
我在function_general.py中编写了process_excel函数,该函数接收Excel文件路径、工作表名和字典列表作为输入,会生成与字典列表长度一致、按列表内容命名的DataFrames。在function_general.py内运行该函数可得到多个命名DataFrames,但将其导入其他文件调用时,仅返回无名称的DataFrame列表。使用Spyder环境,将函数代码复制到新文件中运行可得到预期结果,但希望通过导入方式使用该函数,寻求解决办法。
原函数代码
import pandas as pd def process_excel(excel_file_path, sheet_name, result_list): df = pd.read_excel(excel_file_path, sheet_name, header=3) df['GC Injection'] = pd.to_numeric(df['GC Injection'], errors='coerce') df = df[~df['GC Injection'].isin(['Maximum', 'Average', 'Minimum', 'Standard Deviation', 'Relative Standard Deviation'])] variable = [b.get("experiment").replace(' ','_') + "_data" for b in result_list] for index, value in enumerate(result_list): dhp_test = df[(df['GC Injection'] >= value['Start Injection Number']) & (df['GC Injection'] <= value['End Injection Number'])] dhp_test = dhp_test.reset_index(drop=True) # Adding the 'minutes reaction' column dhp_test['minutes reaction'] = value['Start Minutes'] + 10 * dhp_test.index dhp_test['minutes reaction'] = dhp_test.pop(dhp_test.columns[-1]) dhp_test.insert(3,"minutes",dhp_test['minutes reaction']) dhp_test.drop(['min'],axis = 1, inplace = True) dhp_test.drop(['minutes reaction'],axis = 1, inplace = True) globals()[variable[index]] = dhp_test return [globals()[var] for var in variable]
导入调用代码
from function_general import process_excel result_dataframes = process_excel(excel_file_path, "GC", injections_list)
问题原因
核心问题出在globals()[variable[index]] = dhp_test这行代码:globals()获取的是函数所在模块(function_general.py)的全局命名空间,而非调用该函数的文件的命名空间。因此,当你在其他文件导入调用时,这些命名好的DataFrames只会存在于function_general模块的全局空间中,调用者的文件无法直接访问,只能拿到函数返回的无名称列表。
解决方案
方案1:返回字典(推荐)
将函数修改为返回一个字典,键为DataFrame的命名,值为对应的DataFrame。这种方式符合Python函数设计原则,无副作用,且调用者能灵活访问数据。
修改后的函数代码:
import pandas as pd def process_excel(excel_file_path, sheet_name, result_list): df = pd.read_excel(excel_file_path, sheet_name, header=3) df['GC Injection'] = pd.to_numeric(df['GC Injection'], errors='coerce') df = df[~df['GC Injection'].isin(['Maximum', 'Average', 'Minimum', 'Standard Deviation', 'Relative Standard Deviation'])] result_dict = {} for value in result_list: # 生成DataFrame的名称 exp_name = value.get("experiment").replace(' ','_') + "_data" # 筛选数据 dhp_test = df[(df['GC Injection'] >= value['Start Injection Number']) & (df['GC Injection'] <= value['End Injection Number'])] dhp_test = dhp_test.reset_index(drop=True) # 处理列逻辑 dhp_test['minutes reaction'] = value['Start Minutes'] + 10 * dhp_test.index dhp_test.insert(3, "minutes", dhp_test['minutes reaction']) dhp_test.drop(['min', 'minutes reaction'], axis=1, inplace=True) # 将DataFrame存入字典 result_dict[exp_name] = dhp_test return result_dict
调用方式:
from function_general import process_excel result_dataframes = process_excel(excel_file_path, "GC", injections_list) # 通过键名访问指定DataFrame print(result_dataframes['你的实验名称_data'])
方案2:传入调用者的全局命名空间(不推荐)
如果一定要在调用者的命名空间中直接生成命名变量,可以将调用者的globals()作为参数传入函数。这种方式耦合性高,容易引发命名冲突,不推荐使用。
修改后的函数代码:
import pandas as pd def process_excel(excel_file_path, sheet_name, result_list, caller_globals=None): df = pd.read_excel(excel_file_path, sheet_name, header=3) df['GC Injection'] = pd.to_numeric(df['GC Injection'], errors='coerce') df = df[~df['GC Injection'].isin(['Maximum', 'Average', 'Minimum', 'Standard Deviation', 'Relative Standard Deviation'])] variable = [b.get("experiment").replace(' ','_') + "_data" for b in result_list] result_dfs = [] for index, value in enumerate(result_list): dhp_test = df[(df['GC Injection'] >= value['Start Injection Number']) & (df['GC Injection'] <= value['End Injection Number'])] dhp_test = dhp_test.reset_index(drop=True) # 处理列逻辑 dhp_test['minutes reaction'] = value['Start Minutes'] + 10 * dhp_test.index dhp_test.insert(3, "minutes", dhp_test['minutes reaction']) dhp_test.drop(['min', 'minutes reaction'], axis=1, inplace=True) result_dfs.append(dhp_test) # 若传入调用者的globals,则写入变量 if caller_globals is not None: caller_globals[variable[index]] = dhp_test return result_dfs
调用方式:
from function_general import process_excel result_dataframes = process_excel(excel_file_path, "GC", injections_list, caller_globals=globals()) # 直接使用命名好的变量 print(你的实验名称_data)
内容的提问来源于stack exchange,提问作者RB_student
相关产品推荐
相关产品推荐

