配置Spark on Yarn使用Hadoop原生库的问题咨询
Hey there! Let's walk through your problem and confirm that your solution is on the right track—plus cover some key details to make sure everything stays stable in a cluster environment.
Why the Error Happens
The java.lang.UnsatisfiedLinkError you're seeing in YARN mode boils down to one core issue: YARN containers don't automatically inherit your local environment's library paths. When running in non-YARN mode, Spark uses your machine's LD_LIBRARY_PATH (which you set in .bashrc), so it finds the Snappy native library easily. But in YARN mode, Driver and Executor processes run in isolated containers managed by YARN, so those local environment settings don't carry over.
Your Current Solution Is Valid (With Extra Checks)
Adding spark.driver.extraLibraryPath and spark.executor.extraLibraryPath to spark-defaults.conf is a standard, correct way to fix this. Here's how to make sure you're doing it right:
Verify the Snappy Library Path
First, confirm the exact path to your Snappy native library. On Ubuntu 20.04, this is typically:- System-installed:
/usr/lib/x86_64-linux-gnu/ - Hadoop-native bundled:
<your-hadoop-install-dir>/lib/native(if you compiled Hadoop with Snappy support)
Runhadoop checknative -ato verify Hadoop recognizes Snappy—look forsnappy: truein the output. If it'sfalse, install the system package first:
sudo apt-get install libsnappy-dev libsnappy-java- System-installed:
Configure Spark Properly
Inspark-defaults.conf, add these lines (replace the path with your actual Snappy library location):spark.driver.extraLibraryPath /usr/lib/x86_64-linux-gnu/ spark.executor.extraLibraryPath /usr/lib/x86_64-linux-gnu/Alternatively, you can pass these settings when launching
spark-shell:spark-shell --master yarn \ --conf spark.driver.extraLibraryPath=/usr/lib/x86_64-linux-gnu/ \ --conf spark.executor.extraLibraryPath=/usr/lib/x86_64-linux-gnu/Ensure Cluster-Wide Consistency
Critical note: This path must exist on every YARN node in your cluster. If some nodes have Snappy installed in a different location, those Executors will still throw errors. For a cluster, it's best to standardize the library path across all machines.
Alternative: Set Executor Environment Variables
Another reliable approach is to explicitly set the LD_LIBRARY_PATH for YARN containers using spark.executorEnv:
spark.executorEnv.LD_LIBRARY_PATH /usr/lib/x86_64-linux-gnu/:$LD_LIBRARY_PATH spark.driverEnv.LD_LIBRARY_PATH /usr/lib/x86_64-linux-gnu/:$LD_LIBRARY_PATH
This prepends the Snappy path to the container's existing library path, which can avoid conflicts with other native libraries.
Final Confirmation
Your initial fix is absolutely valid—this is the go-to method for passing native library paths to Spark in YARN mode. The key is ensuring all cluster nodes have the library available at the same path, and verifying Hadoop's native support for Snappy.
内容的提问来源于stack exchange,提问作者rbok78

