SageMaker FastFile输入模式配置疑问:input_data_config显示异常
问题描述
当使用Amazon SageMaker的input_mode为File或Pipe时,input_data_config中会对应显示TrainingInputMode为File或Pipe:
"input_data_config": {"train": {"TrainingInputMode": "File",
"S3DistributionType": "ShardedByS3Key",
"RecordWrapperType": "None"}
"input_data_config": {"train": {"TrainingInputMode": "Pipe",
"S3DistributionType": "ShardedByS3Key",
"RecordWrapperType": "None"}
但当指定input_mode为FastFile时,input_data_config仍显示TrainingInputMode为File:
"input_data_config": {"train": {"TrainingInputMode": "File",
"S3DistributionType": "ShardedByS3Key",
"RecordWrapperType": "None"}
请问这是否意味着我的实现未成功启用FastFile输入模式,或是FastFile模式本就不会在input_data_config中显示对应条目?
我的代码
import sagemaker from sagemaker import get_execution_role from sagemaker.tensorflow import TensorFlow hyperparameters = {"input_mode": "FastFile", #"FastFile", "Pipe", "File" "shards_on_input": 4 # 4, 120 (for input subfolders in S3) } train_input = sagemaker.inputs.TrainingInput("s3://mnist-tdrecords/train/{}".format(hyperparameters["shards_on_input"]), input_mode = hyperparameters["input_mode"], distribution = 'ShardedByS3Key' #'ShardedByS3Key', 'FullyReplicated' ) tf_estimator = TensorFlow(entry_point = "AWS_DataPipping_TFMirroredStrategy.py", source_dir = "./", framework_version = "2.3", py_version = "py37", instance_type = "ml.p3.2xlarge", # "ml.p3.2xlarge", "ml.p3.8xlarge", "ml.p3.16xlarge" instance_count = 1, role = sagemaker.get_execution_role(), hyperparameters = hyperparameters, output_path = f"s3://mnist-tdrecords/output", input_mode = hyperparameters["input_mode"], # "File", "Pipe", "FastFile") ) s3_data_channels = {"train": train_input} #s3_data_channels = {"train": "s3://mnist-tdrecords/train/{}".format(hyperparameters["shards_on_input"])} #"validation": f"s3://{bucket_name}/data/validation",} tf_estimator.fit(s3_data_channels)
解答
这是正常现象,FastFile模式本身不会在input_data_config中显示独立的TrainingInputMode条目——它本质是File模式的优化版本,底层仍会标记为File类型,但实际会启用FastFile的特性。
要确认FastFile是否成功启用,可以通过以下方式验证:
- 查看训练作业日志,搜索
FastFile关键词,会找到类似Using FastFile input mode的日志信息 - 观察训练启动速度:FastFile无需提前将S3数据下载到本地实例磁盘,而是直接流式访问S3,因此启动时间会比传统File模式短很多,尤其是数据量较大时
- 检查实例磁盘占用:传统File模式会将S3数据下载到本地,磁盘占用明显增加;FastFile模式下本地磁盘几乎不会有数据存储
另外,你的代码配置是正确的:
- 同时在
TrainingInput和TensorFlowestimator中指定了input_mode="FastFile" - 使用的ml.p3系列实例支持FastFile模式
只要代码中正确指定了input_mode="FastFile",且使用的实例类型和SageMaker版本支持该特性,FastFile就会正常生效,无需纠结input_data_config中的显示值。
内容的提问来源于stack exchange,提问作者sebtac

