SageMaker脚本模式训练:如何解决自定义模块导入失败问题?
Looks like the core issue here is that when you submit your training job with the TensorFlow estimator in script mode, SageMaker only uploads the entry_point file (train.py) to the training container by default. All your other custom modules (like dataset.py, models.py) and the config directory aren't being sent over, so the Python environment in the container can't locate them.
Here are three straightforward solutions to fix this:
1. Use the source_dir parameter (Recommended)
The simplest way to include all your local files and directories is to specify the source_dir parameter when initializing your TensorFlow estimator. This tells SageMaker to upload the entire contents of the specified directory to the training container.
Since your train.py lives in the root of your working directory, just set source_dir='.' (the dot represents your current directory):
estimator = TensorFlow( entry_point='train.py', source_dir='.', # Add this line to include all local files role=role, train_instance_count=1, train_instance_type='ml.p2.xlarge', framework_version='1.14', py_version='py3', script_mode=True, hyperparameters={ 'epochs': 10 } )
This will upload everything in your WORKDIR (including dataset.py, config/, etc.) to the container, so your imports will work as expected.
2. Use the dependencies parameter (For selective uploads)
If you don't want to upload your entire working directory (e.g., to exclude train.ipynb or other unnecessary files), you can use the dependencies parameter to list exactly which files/directories to include:
estimator = TensorFlow( entry_point='train.py', dependencies=[ 'dataset.py', 'models.py', 'densenet.py', 'resnet.py', 'imagenet_utils.py', 'keras_utils.py', 'utils.py', 'config/' # Includes the entire config directory ], role=role, train_instance_count=1, train_instance_type='ml.p2.xlarge', framework_version='1.14', py_version='py3', script_mode=True, hyperparameters={ 'epochs': 10 } )
SageMaker will upload all the items listed in dependencies alongside your train.py, making them accessible to your script.
3. Add the script directory to Python path (Quick fix)
If you prefer not to modify your estimator setup, you can add a few lines at the very top of train.py to add the current directory to Python's module search path:
import sys import os # Add the current script's directory to sys.path sys.path.append(os.path.dirname(os.path.abspath(__file__))) # Now your imports will work from dataset import WatchDataSet from models import BCNN # ... rest of your imports
This tells Python to look for modules in the same directory as train.py, which will resolve the ModuleNotFoundError. Note that this is more of a quick fix and less clean than using source_dir or dependencies.
After applying any of these solutions, re-run your training job and the import errors should be resolved.
内容的提问来源于stack exchange,提问作者Ashutosh Kumar

