AWS Step Function预处理与特征工程流程S3文件未保存问题咨询
问题根因分析
- 本地输出目录未创建:你代码里直接往
/opt/ml/processing/train、/opt/ml/processing/test路径写文件,但SageMaker Processing容器默认不会自动创建这两个输出目录,写入时会触发路径不存在的IO错误,部分运行环境会吞掉这类错误返回执行成功的状态,但实际文件没有写入成功。 - OneHotEncoder输出格式不兼容CSV写入:默认参数的
OneHotEncoder输出是稀疏矩阵,直接传入pd.DataFrame()生成的是存储稀疏对象的表格,写入CSV时要么内容为空,要么直接触发写入失败,不会生成有效文件。 - ProcessingStep输出通道未正确配置:即使本地文件写入成功,如果你在Step Functions定义ProcessingStep时,没有将本地的
/opt/ml/processing/train、/opt/ml/processing/test路径配置为输出通道、绑定对应的S3存储路径,SageMaker不会自动将本地生成的文件同步上传到指定的S3目录。 - 额外逻辑问题(非文件缺失直接原因,但会影响后续训练):代码中测试集特征处理调用了
OneHotEncoder().fit_transform(X_train),应该复用训练集拟合好的编码器调用transform(X_test),否则会造成特征泄漏、训练测试特征维度不一致的问题。
修复方案
1. 预处理代码修复
修改preprocessing.py代码,补充目录创建逻辑、调整OneHotEncoder参数、修正测试集处理逻辑:
%%writefile preprocessing.py import argparse import os import warnings import numpy as np import pandas as pd from sklearn.model_selection import train_test_split from sklearn.preprocessing import StandardScaler, OneHotEncoder, LabelBinarizer, KBinsDiscretizer from sklearn.preprocessing import PolynomialFeatures from sklearn.compose import make_column_transformer from sklearn.exceptions import DataConversionWarning warnings.filterwarnings(action="ignore", category=DataConversionWarning) if __name__ == "__main__": parser = argparse.ArgumentParser() parser.add_argument("--train-test-split-ratio", type=float, default=0.3) args, _ = parser.parse_known_args() print("Received arguments {}".format(args)) input_data_path = os.path.join("/opt/ml/processing/input", "raw-data.csv") print("Reading input data from {}".format(input_data_path)) df = pd.read_csv(input_data_path) # Handle null values df['Gender'][df['Gender'].isnull()]='Male' df['Married'][df['Married'].isnull()]='Yes' df['LoanAmount'][df['LoanAmount'].isnull()]= df['LoanAmount'].mean() df['Loan_Amount_Term'][df['Loan_Amount_Term'].isnull()]='360' df['Self_Employed'][df['Self_Employed'].isnull()]='No' df['Credit_History'][df['Credit_History'].isnull()]='1' df['Dependents'][df['Dependents'].isnull()]='0' df.loc[df.Dependents=='3+','Dependents']= 4 # Convert data types to numeric df.loc[df.Loan_Status=='N','Loan_Status']= 0 df.loc[df.Loan_Status=='Y','Loan_Status']=1 df.loc[df.Gender=='Male','Gender']= 0 df.loc[df.Gender=='Female','Gender']=1 df.loc[df.Married=='No','Married']= 0 df.loc[df.Married=='Yes','Married']=1 df.loc[df.Education=='Graduate','Education']= 0 df.loc[df.Education=='Not Graduate','Education']=1 df.loc[df.Self_Employed=='No','Self_Employed']= 0 df.loc[df.Self_Employed=='Yes','Self_Employed']=1 df['Married'] = df['Married'].astype(str).astype(int) df['Dependents'] = df['Dependents'].astype(str).astype(int) df['Education'] = df['Education'].astype(str).astype(int) df['Self_Employed'] = df['Self_Employed'].astype(str).astype(int) df['Loan_Amount_Term'] = df['Loan_Amount_Term'].astype(str).astype(float) df['Credit_History'] = df['Credit_History'].astype(str).astype(float) df['Loan_Status'] = df['Loan_Status'].astype(str).astype(int) df = df.drop('Loan_ID', axis=1) split_ratio = args.train_test_split_ratio print("Splitting data into train and test sets with ratio {}".format(split_ratio)) X_train, X_test, y_train, y_test = train_test_split( df.drop("Loan_Status", axis=1), df["Loan_Status"], test_size=split_ratio, random_state=0) print("Running preprocessing and feature engineering transformations") # 指定OneHotEncoder输出稠密矩阵,避免稀疏格式问题 enc = OneHotEncoder(sparse_output=False, handle_unknown='ignore') train_features = enc.fit_transform(X_train) # 测试集复用训练好的编码器,不要重新fit避免特征不一致 test_features = enc.transform(X_test) print("Train data shape after preprocessing: {}".format(train_features.shape)) print("Test data shape after preprocessing: {}".format(test_features.shape)) # 先创建输出目录,避免路径不存在错误 os.makedirs("/opt/ml/processing/train", exist_ok=True) os.makedirs("/opt/ml/processing/test", exist_ok=True) train_features_output_path = os.path.join("/opt/ml/processing/train", "train_features.csv") train_labels_output_path = os.path.join("/opt/ml/processing/train", "train_labels.csv") test_features_output_path = os.path.join("/opt/ml/processing/test", "test_features.csv") test_labels_output_path = os.path.join("/opt/ml/processing/test", "test_labels.csv") print("Saving training features to {}".format(train_features_output_path)) pd.DataFrame(train_features).to_csv(train_features_output_path, header=False, index=False) print("Saving test features to {}".format(test_features_output_path)) pd.DataFrame(test_features).to_csv(test_features_output_path, header=False, index=False) print("Saving training labels to {}".format(train_labels_output_path)) y_train.to_csv(train_labels_output_path, header=False, index=False) print("Saving test labels to {}".format(test_labels_output_path)) y_test.to_csv(test_labels_output_path, header=False, index=False)
2. ProcessingStep配置检查
确认Step Functions中ProcessingStep的输出配置已经绑定了本地路径和目标S3路径,参考配置逻辑如下(以Python SDK定义为例):
from sagemaker.processing import ProcessingOutput from sagemaker.workflow.steps import ProcessingStep processing_step = ProcessingStep( name="PreprocessingStep", processor=your_processor, inputs=[...], outputs=[ ProcessingOutput( output_name="train", source="/opt/ml/processing/train", destination="s3://你的存储桶/你的训练文件输出路径" ), ProcessingOutput( output_name="test", source="/opt/ml/processing/test", destination="s3://你的存储桶/你的测试文件输出路径" ) ], code="preprocessing.py" )
3. 权限验证
确认ProcessingStep使用的IAM角色具备对应S3存储桶的s3:PutObject权限,避免因权限不足导致文件上传失败。
内容的提问来源于stack exchange,提问作者Display Name is missing
相关产品推荐
相关产品推荐

