You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

定义固定训练/测试集并在函数中按条件调用时遇变量未绑定错误求助

固定训练测试集下的多模型测试问题

问题背景

需要基于DataFrame测试不同模型,全程使用同一组训练集与测试集,但拆分后的变量在函数中无法正常使用,代码执行时出现UnboundLocalError: local variable 'X_train' referenced before assignment错误。

原始代码

import pandas as pd
from sklearn.model_selection import StratifiedKFold
from sklearn.preprocessing import StandardScaler

df = pd.read_excel(r'file.xlsx') # Load in dataset

X = df.iloc[:,2:] # features
y = df.iloc[:,0] # Metastases (1) or not (0)
skf = StratifiedKFold(n_splits=5)  # Define StratifiedKFold with 5 splits

for train_index, test_index in skf.split(X, y):
    X_train, X_test = X.loc[train_index], X.loc[test_index] 
    y_train, y_test = y.loc[train_index], y.loc[test_index] 

def model(variable):
    if model == 'clinical':
        X_train=pd.DataFrame(X_train.iloc[:,:6])
        X_test = pd.DataFrame(X_test.iloc[:, :6])
    return (X_train, X_test)
    if model == 'radiomic' or model == 'combined':
        X_train = pd.DataFrame(scaler.fit_transform(X_train))
        X_test = pd.DataFrame(scaler.fit_transform(X_test))
    return (X_train, X_test)

clinical_train, clinical_test = model('clinical')

错误原因分析

  • 变量作用域混淆:函数内直接使用外部的X_train、X_test,但又对同名变量赋值,Python会将其视为局部变量,导致赋值前引用报错。
  • 参数名错误:函数参数是variable,判断条件却用了model == 'clinical',应该用参数名而非函数名。
  • 逻辑执行顺序错误:第一个return会直接终止函数,后续的if判断永远不会执行。
  • 未初始化scaler:代码里用了scaler但没创建实例,会触发NameError。

解决方案

推荐方案:将训练测试集作为参数传入函数

避免全局变量,把拆分好的数据集传入函数,同时修正逻辑错误:

import pandas as pd
from sklearn.model_selection import StratifiedKFold
from sklearn.preprocessing import StandardScaler

df = pd.read_excel(r'file.xlsx') # Load in dataset

X = df.iloc[:,2:] # features
y = df.iloc[:,0] # Metastases (1) or not (0)
skf = StratifiedKFold(n_splits=5)  # Define StratifiedKFold with 5 splits

# 只保留第一组交叉验证的训练测试集(确保全程用同一组数据)
for train_index, test_index in skf.split(X, y):
    X_train, X_test = X.loc[train_index], X.loc[test_index] 
    y_train, y_test = y.loc[train_index], y.loc[test_index] 
    break  

def prepare_data(variable, X_train, X_test):
    scaler = StandardScaler()  # 初始化标准化器
    if variable == 'clinical':
        # 提取前6列临床特征
        processed_train = X_train.iloc[:,:6].copy()
        processed_test = X_test.iloc[:,:6].copy()
    elif variable == 'radiomic':
        # 标准化放射组学特征(假设为第6列之后的部分)
        processed_train = pd.DataFrame(scaler.fit_transform(X_train.iloc[:,6:]), index=X_train.index)
        processed_test = pd.DataFrame(scaler.transform(X_test.iloc[:,6:]), index=X_test.index)
    elif variable == 'combined':
        # 合并临床特征与标准化后的放射组学特征
        clinical_train = X_train.iloc[:,:6].copy()
        radiomic_train = pd.DataFrame(scaler.fit_transform(X_train.iloc[:,6:]), index=X_train.index)
        processed_train = pd.concat([clinical_train, radiomic_train], axis=1)
        
        clinical_test = X_test.iloc[:,:6].copy()
        radiomic_test = pd.DataFrame(scaler.transform(X_test.iloc[:,6:]), index=X_test.index)
        processed_test = pd.concat([clinical_test, radiomic_test], axis=1)
    else:
        raise ValueError("variable must be 'clinical', 'radiomic' or 'combined'")
    return processed_train, processed_test

# 用固定数据集生成不同模型的输入数据
clinical_train, clinical_test = prepare_data('clinical', X_train, X_test)
radiomic_train, radiomic_test = prepare_data('radiomic', X_train, X_test)
combined_train, combined_test = prepare_data('combined', X_train, X_test)

方案说明

  • 拆分数据集后用break固定一组,保证所有模型用同一训练测试集对比。
  • 函数接收数据集作为参数,彻底解决作用域问题。
  • 标准化时仅在训练集拟合scaler,测试集复用同一转换规则,避免数据泄露。
  • 保留原始索引,方便后续与标签数据对应。

之前尝试方案的问题

  1. 全局变量方案:虽能解决作用域问题,但会降低代码可维护性,容易引发意外的变量修改。
  2. 函数内拆分数据集:每次调用都会重新拆分,导致训练测试集不一致,无法保证模型对比的公平性。
  3. 函数置于for循环内:会重复定义函数,且若无break只会保留最后一组交叉验证数据,原逻辑错误依然存在。

内容的提问来源于stack exchange,提问作者Lieke Pullen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 23:48:23