如何重新编译TensorFlow以支持SSE4.2和AVX指令集?
Hey there! I’ve gone through this exact process a few times for older Ubuntu environments like yours, so let’s walk through step-by-step how to recompile TensorFlow from source to eliminate those SSE4.2/AVX warnings and speed up your CPU computations.
Prerequisites First
First, make sure your system is up-to-date and has all the necessary build tools installed:
- Update your package list and upgrade existing packages:
sudo apt-get update && sudo apt-get upgrade -y - Install core build dependencies:
sudo apt-get install build-essential git python-pip python-dev python-numpy python-wheel -y
Install Bazel (TensorFlow’s Build Tool)
TensorFlow uses Bazel for compilation, and you’ll need a version compatible with TensorFlow 1.x (since your Ubuntu 16.04 environment is paired with older libraries like Keras from PyImageSearch). We’ll use Bazel 0.26.1, which works well with TensorFlow 1.15 (a stable fit for your setup):
# Add Bazel repository echo "deb [arch=amd64] http://storage.googleapis.com/bazel-apt stable jdk1.8" | sudo tee /etc/apt/sources.list.d/bazel.list curl https://bazel.build/bazel-release.pub.gpg | sudo apt-key add - # Install Bazel 0.26.1 sudo apt-get update && sudo apt-get install bazel-0.26.1 -y
Clone TensorFlow Source Code
We’ll target TensorFlow 1.15, which is compatible with Ubuntu 16.04 and your pre-installed Keras setup:
git clone https://github.com/tensorflow/tensorflow.git cd tensorflow git checkout r1.15
Configure Compilation for SSE4.2 & AVX
This is the critical step to enable the instruction sets. We’ll use a flag that automatically detects all supported CPU optimizations (including SSE4.2 and AVX) instead of manually specifying them:
Set the compiler optimization flag first:
export CC_OPT_FLAGS="-march=native"The
-march=nativeflag tells GCC to optimize for your exact CPU, enabling all available instruction sets like SSE4.2, AVX, and more.Run the TensorFlow configuration script:
./configureAnswer the prompts as follows (adjust if you have a custom Python setup):
- Python interpreter path: Use the default (e.g.,
/usr/bin/python) or your virtual environment path if you’re using one. - CUDA support: Type
n(since we’re optimizing for CPU here). - All other prompts: Press Enter to accept the default values.
- Python interpreter path: Use the default (e.g.,
Compile TensorFlow
This will take a while (30 mins to 2+ hours depending on your VM’s CPU/memory), so be patient:
bazel build --config=opt //tensorflow/tools/pip_package:build_pip_package
The --config=opt flag ensures our earlier optimization settings are applied during compilation.
Build & Install the Custom TensorFlow Package
Once compilation finishes, generate a pip package and replace your existing TensorFlow installation:
- Generate the pip package:
bazel-bin/tensorflow/tools/pip_package/build_pip_package /tmp/tensorflow_pkg - Uninstall the original TensorFlow:
pip uninstall tensorflow -y - Install your custom-compiled version:
pip install /tmp/tensorflow_pkg/tensorflow-*.whl
Verify the Fix
Launch a Python shell and test if the warnings are gone:
import tensorflow as tf print(f"TensorFlow version: {tf.__version__}")
If you don’t see the SSE4.2/AVX warnings anymore, you’re all set!
Quick Troubleshooting Tip
If your VM runs out of memory during compilation, add a swap file to give it more headroom:
sudo fallocate -l 4G /swapfile sudo chmod 600 /swapfile sudo mkswap /swapfile sudo swapon /swapfile
内容的提问来源于stack exchange,提问作者Calvin Godfrey

