You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Synapse中如何将容器文件列表转为PySpark DataFrame?

解决Azure Synapse中FileInfo列表转PySpark DataFrame全Null问题

问题原因

你原代码中的列表推导式逻辑错误:mssparkutils.fs.ls返回的是FileInfo对象的一维列表,并非嵌套列表。[item for sublist in list_for_dataframe for item in sublist]会错误地将每个FileInfo对象拆分为无效的单个元素,导致Spark无法匹配预设的Schema,最终生成全Null的DataFrame。

正确实现代码

方法一:通过字典列表转换(字段对应直观)

from pyspark.sql.types import StructType, StructField, StringType, LongType

# 先获取FileInfo对象列表
file_list = mssparkutils.fs.ls("abfss://config@datalake.dfs.core.windows.net/")

# 提取每个FileInfo的属性,生成字典列表
rows = [{"path": f.path, "name": f.name, "size": f.size} for f in file_list]

schema = StructType([
    StructField("path", StringType(), True),
    StructField("name", StringType(), True),
    StructField("size", LongType(), True)
])

df = spark.createDataFrame(rows, schema=schema)
df.show()

方法二:通过元组列表转换(需与Schema字段顺序一致)

from pyspark.sql.types import StructType, StructField, StringType, LongType

# 先获取FileInfo对象列表
file_list = mssparkutils.fs.ls("abfss://config@datalake.dfs.core.windows.net/")

# 提取属性生成元组列表,顺序必须与Schema的path/name/size对应
rows = [(f.path, f.name, f.size) for f in file_list]

schema = StructType([
    StructField("path", StringType(), True),
    StructField("name", StringType(), True),
    StructField("size", LongType(), True)
])

df = spark.createDataFrame(rows, schema=schema)
df.show()

效果说明

两种方法都能正确将FileInfo对象的属性映射到DataFrame的对应列,生成符合预期的包含path、name、size三列的数据集。

内容的提问来源于stack exchange,提问作者coding

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 01:43:19