本地测试正常,提交Cloud ML Engine远程训练任务遇Numpy版本兼容错误
Let's break down what's happening here and how to fix it quickly:
The Root Cause
You created your data.npz file using Numpy 1.14 (which uses API version 0xc) in Google Datalab, but Cloud ML Engine's runtime version 1.7 ships with an older Numpy version (1.13.x, API version 0xb). Numpy's binary formats like .npz are tied to the API version, so older versions can't read files saved by newer ones.
Solutions to Try
1. Upgrade your Cloud ML Engine Runtime Version
The simplest fix is to use a newer runtime version that includes Numpy 1.14 or higher. Runtime version 1.8 and above come with Numpy 1.14.x, which matches the version you used to create your data file. Just update your submission command:
gcloud ml-engine jobs submit training $JOB_NAME \ --job-dir gs://$BUCKET_NAME/$JOB_NAME \ --runtime-version 1.8 \ # Add any other flags you were using here
2. Regenerate the .npz File with a Compatible Numpy Version
If you can't upgrade the runtime (for example, if your code relies on older TensorFlow features in 1.7), you can re-save your data using Numpy 1.13.x. In Google Datalab, you can do this with:
# Install the matching Numpy version !pip install numpy==1.13.3 --quiet # Re-save your data (replace `your_data_dict` with your actual data structure) import numpy as np np.savez('data.npz', **your_data_dict) # Upload the new file back to GCS !gsutil cp data.npz gs://path_to_data.npz
This will create a .npz file that the runtime 1.7's Numpy can read without issues.
3. Force Install a Newer Numpy Version in the Training Job
As a last resort, you can override the runtime's default Numpy version by specifying it in your submission command. This can sometimes cause dependency conflicts, so test it first:
gcloud ml-engine jobs submit training $JOB_NAME \ --job-dir gs://$BUCKET_NAME/$JOB_NAME \ --runtime-version 1.7 \ --packages numpy==1.14.0
A Quick Note on Numpy API Versions
Numpy increments its API version when it makes breaking changes to its binary format. Version 1.14 introduced API 0xc, while 1.13 used 0xb—that's why the error message calls out this mismatch. Sticking to matching versions between data creation and training runtime avoids this headache entirely.
内容的提问来源于stack exchange,提问作者GRS

