You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Debian11.6下独立Spark环境Python运行Delta Lake报错求助

问题描述

Debian 11.6系统中已安装独立版Spark 3.3.1与Anaconda,尝试通过Python使用Delta Lake,运行以下代码:

import pyspark
from delta import *

builder = pyspark.sql.SparkSession.builder.appName("MyApp") \
    .config("spark.sql.extensions", "io.delta.sql.DeltaSparkSessionExtension") \
    .config("spark.sql.catalog.spark_catalog", "org.apache.spark.sql.delta.catalog.DeltaCatalog")

spark = configure_spark_with_delta_pip(builder).getOrCreate()

运行后输出如下日志(含警告):

:: loading settings :: url = jar:file:/usr/bin/spark-3.3.1-bin-hadoop3/jars/ivy-2.5.0.jar!/org/apache/ivy/core/settings/ivysettings.xml

Ivy Default Cache set to: /home/boss/.ivy2/cache
The jars for the packages stored in: /home/boss/.ivy2/jars
io.delta#delta-core_2.12 added as a dependency
:: resolving dependencies :: org.apache.spark#spark-submit-parent-290d27e6-7e29-475f-81b5-1ab1331508fc;1.0
    confs: [default]
    found io.delta#delta-core_2.12;2.2.0 in central
    found io.delta#delta-storage;2.2.0 in central
    found org.antlr#antlr4-runtime;4.8 in central
:: resolution report :: resolve 272ms :: artifacts dl 10ms
    :: modules in use:
    io.delta#delta-core_2.12;2.2.0 from central in [default]
    io.delta#delta-storage;2.2.0 from central in [default]
    org.antlr#antlr4-runtime;4.8 from central in [default]
    ---------------------------------------------------------------------
    |                  |            modules            ||   artifacts   |
    |       conf       | number| search|dwnlded|evicted|| number|dwnlded|
    ---------------------------------------------------------------------
    |      default     |   3   |   0   |   0   |   0   ||   3   |   0   |
    ---------------------------------------------------------------------
:: retrieving :: org.apache.spark#spark-submit-parent-290d27e6-7e29-475f-81b5-1ab1331508fc
    confs: [default]
    0 artifacts copied, 3 already retrieved (0kB/11ms)

23/01/24 04:10:26 WARN NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable

Setting default log level to "WARN".
To adjust logging level use sc.setLogLevel(newLevel). For SparkR, use setLogLevel(newLevel).
解决方案

1. Ivy依赖日志处理(可选)

Ivy的输出是Spark自动检索Delta依赖的正常流程,说明Delta相关jar包已缓存到/home/boss/.ivy2/jars,无需额外操作。若嫌日志冗余,可在SparkSession配置中添加日志级别控制,过滤Ivy的INFO级日志。

2. NativeCodeLoader警告处理

这个警告是因为Debian系统缺少对应架构的Hadoop本地库,或Spark未正确配置本地库路径,有两种处理方式:

  • 编译适配本地库:从Hadoop源码编译适配Debian 11.6架构的本地库,将库文件路径添加到LD_LIBRARY_PATH环境变量,或在Spark的spark-env.sh中配置SPARK_LIBRARY_PATH指向该路径。
  • 直接忽略(推荐):如果不需要Hadoop本地库的性能优化(如压缩、IO加速),可直接忽略该警告,Spark会自动使用Java实现的替代类,不影响Delta Lake核心功能。

3. 日志级别调整(可选)

若想减少日志输出量,可在创建SparkSession前添加日志级别配置:

import pyspark
from delta import *

# 初始化SparkContext并设置日志级别
sc = pyspark.SparkContext.getOrCreate()
sc.setLogLevel("ERROR")

builder = pyspark.sql.SparkSession.builder.appName("MyApp") \
    .config("spark.sql.extensions", "io.delta.sql.DeltaSparkSessionExtension") \
    .config("spark.sql.catalog.spark_catalog", "org.apache.spark.sql.delta.catalog.DeltaCatalog")

spark = configure_spark_with_delta_pip(builder).getOrCreate()

内容的提问来源于stack exchange,提问作者Tavakoli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 06:56:54