使用Sagemaker Ground Truth增强清单文件时RecordIO格式解析错误排查
问题描述
我使用SageMaker Ground Truth创建了标注文件,参考官方文档和示例Notebook创建目标检测训练任务并执行超参数调优,但始终遇到以下错误:
ClientError: Unable to parse record. Please make sure input data is in correct recordio format. , exit code: 2
相关代码
#model config od_model = sagemaker.estimator.Estimator(training_image, role, instance_count=1, instance_type='ml.p3.2xlarge', train_volume_size=50, train_max_run=360000, input_mode='Pipe', output_path=s3_output_location, sagemaker_session=sagemaker_session) # Set hyperparameters od_model.set_hyperparameters(base_network='resnet-50', use_pretrained_model=1, num_classes=2, mini_batch_size=5, epochs=30, learning_rate=0.001, lr_scheduler_step='10,20', lr_scheduler_factor=0.1, optimizer='adam', num_training_samples=str(num_training_samples)) #data train_data = "s3://bucket/.../manifests/output/output.manifest" validation_data = "s3://bucket/.../manifests/output/.manifest" train_channel = sagemaker.inputs.TrainingInput(train_data, distribution='FullyReplicated', content_type='application/x-recordio', s3_data_type='AugmentedManifestFile', attribute_names=attribute_names, input_mode='Pipe', #record_wrapping='RecordIO') #did same for validation
可能的问题及修复方案
Content-Type参数配置错误
当前设置的content_type='application/x-recordio'不符合Augmented Manifest文件格式要求。Augmented Manifest是每行一个JSON的文本文件,正确的Content-Type应为application/jsonlines。训练过程中SageMaker会自动将Augmented Manifest转换为RecordIO格式,无需手动指定RecordIO类型。验证集路径错误
验证集路径指向了以.开头的隐藏文件.manifest,这大概率是错误路径。请修改为实际的验证集manifest文件名(例如validation.manifest),确保路径指向合法存在的文件。Record Wrapping配置冲突
代码中注释了record_wrapping='RecordIO',结合input_mode='Pipe'和AugmentedManifestFile类型,无需手动设置该参数,SageMaker会自动处理格式转换。若需启用该配置,请取消注释,但前提是已修正Content-Type参数。标注文件格式或权限问题
检查manifest文件是否符合规范:- 文件必须是每行一个独立JSON对象,不能用数组包裹整个文件
- 确认
attribute_names指定的字段(如标注任务中定义的标签键)在manifest的每行JSON中都存在 - 确保训练角色拥有对应S3桶的读取权限,且manifest中引用的图片S3路径可正常访问
内容的提问来源于stack exchange,提问作者lan jin

