在Google ML上重训SSD Mobilenet时Object Detection API运行报错
Hey there, let’s dig into why your local training works but the cloud job keeps failing—this is a common gotcha with distributed cloud training, and we can break down the fixes step by step:
1. Fix Cloud Storage Permissions (Top Culprit for UnavailableError: OS Error)
That UnavailableError: OS Error almost always points to permission issues with your GCS bucket. Even if your job can write to the job-dir, there might be missing access for other critical operations (like reading training data, writing checkpoints, or pulling model assets):
- Head over to your GCP project’s IAM section and make sure the ML Engine service account (it looks like
[PROJECT_NUMBER]-compute@developer.gserviceaccount.comby default) has Storage Object Admin or Storage Admin permissions on yourgs://BLAHBLAH-storagebucket. - Double-check every path in your pipeline config file—
train_input_reader,eval_input_reader,model_dir, etc.—all need to be full GCS paths (likegs://your-bucket/train.record), not local paths. Local paths won’t work across cloud replicas.
2. Validate Your Package Builds
Your setup.py and package tarballs might be missing key components the cloud environment needs:
- When building
object_detection-0.1.tar.gzandslim-0.1.tar.gz, run the build commands from the root of the Object Detection API repo. For example:# From the object_detection folder python setup.py sdist # From the slim folder python setup.py sdist - Add
tensorflow>=1.8to yourREQUIRED_PACKAGESinsetup.py. Even though you specify--runtime-version 1.8, the cloud environment might not pull the exact version unless it’s explicitly listed:from setuptools import find_packages from setuptools import setup REQUIRED_PACKAGES = ['tensorflow>=1.8', 'Pillow>=1.0', 'Matplotlib>=2.1', 'Cython>=0.28.1'] setup( name='object_detection', version='0.1', install_requires=REQUIRED_PACKAGES, include_package_data=True, packages=[p for p in find_packages() if p.startswith('object_detection')], description='Tensorflow Object Detection Library', )
3. Tweak Your Cloud Config for GPU & Distributed Training
Your cloud.yaml uses GPU instances, but distributed setups can fail due to network timeouts or resource mismatches:
- TensorFlow 1.8 requires CUDA 9.0 and cuDNN 7.0. While ML Engine’s TF 1.8 runtime should include these, try reducing the worker and parameter server counts temporarily to rule out distributed issues:
If this smaller setup works, you can scale back up gradually.trainingInput: runtimeVersion: "1.8" scaleTier: CUSTOM masterType: standard_gpu workerCount: 1 workerType: standard_gpu parameterServerCount: 1 parameterServerType: standard
4. Get the Full Error Traceback
The truncated logs you shared don’t show the real root cause. To get the full story:
- Go to the Jobs page in GCP ML Engine, click on your failed job (
grewe_object_detection_6), then go to the Logs tab. - Filter logs for the master replica to see the complete stack trace. Look for specific errors like missing GCS files, CUDA compatibility issues, or dependency conflicts—these will point you directly to the problem.
5. Test with a Minimal Working Example
If you’re still stuck, start with a known-good setup:
- Use the official Object Detection API cloud training steps to train a sample model on your bucket. Once that works, incrementally replace the sample data and config with your own to pinpoint what’s breaking your job.
内容的提问来源于stack exchange,提问作者Lynne Grewe

