You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

配置Spark on Yarn使用Hadoop原生库的问题咨询

Snappy Compression Issue with Spark on YARN: Correct Configuration Approach

Hey there! Let's walk through your problem and confirm that your solution is on the right track—plus cover some key details to make sure everything stays stable in a cluster environment.

Why the Error Happens

The java.lang.UnsatisfiedLinkError you're seeing in YARN mode boils down to one core issue: YARN containers don't automatically inherit your local environment's library paths. When running in non-YARN mode, Spark uses your machine's LD_LIBRARY_PATH (which you set in .bashrc), so it finds the Snappy native library easily. But in YARN mode, Driver and Executor processes run in isolated containers managed by YARN, so those local environment settings don't carry over.

Your Current Solution Is Valid (With Extra Checks)

Adding spark.driver.extraLibraryPath and spark.executor.extraLibraryPath to spark-defaults.conf is a standard, correct way to fix this. Here's how to make sure you're doing it right:

  1. Verify the Snappy Library Path
    First, confirm the exact path to your Snappy native library. On Ubuntu 20.04, this is typically:

    • System-installed: /usr/lib/x86_64-linux-gnu/
    • Hadoop-native bundled: <your-hadoop-install-dir>/lib/native (if you compiled Hadoop with Snappy support)
      Run hadoop checknative -a to verify Hadoop recognizes Snappy—look for snappy: true in the output. If it's false, install the system package first:
    sudo apt-get install libsnappy-dev libsnappy-java
    
  2. Configure Spark Properly
    In spark-defaults.conf, add these lines (replace the path with your actual Snappy library location):

    spark.driver.extraLibraryPath /usr/lib/x86_64-linux-gnu/
    spark.executor.extraLibraryPath /usr/lib/x86_64-linux-gnu/
    

    Alternatively, you can pass these settings when launching spark-shell:

    spark-shell --master yarn \
      --conf spark.driver.extraLibraryPath=/usr/lib/x86_64-linux-gnu/ \
      --conf spark.executor.extraLibraryPath=/usr/lib/x86_64-linux-gnu/
    
  3. Ensure Cluster-Wide Consistency
    Critical note: This path must exist on every YARN node in your cluster. If some nodes have Snappy installed in a different location, those Executors will still throw errors. For a cluster, it's best to standardize the library path across all machines.

Alternative: Set Executor Environment Variables

Another reliable approach is to explicitly set the LD_LIBRARY_PATH for YARN containers using spark.executorEnv:

spark.executorEnv.LD_LIBRARY_PATH /usr/lib/x86_64-linux-gnu/:$LD_LIBRARY_PATH
spark.driverEnv.LD_LIBRARY_PATH /usr/lib/x86_64-linux-gnu/:$LD_LIBRARY_PATH

This prepends the Snappy path to the container's existing library path, which can avoid conflicts with other native libraries.

Final Confirmation

Your initial fix is absolutely valid—this is the go-to method for passing native library paths to Spark in YARN mode. The key is ensuring all cluster nodes have the library available at the same path, and verifying Hadoop's native support for Snappy.

内容的提问来源于stack exchange,提问作者rbok78

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 18:17:31