Azure Synapse中如何将容器文件列表转为PySpark DataFrame?
解决Azure Synapse中FileInfo列表转PySpark DataFrame全Null问题
问题原因
你原代码中的列表推导式逻辑错误:mssparkutils.fs.ls返回的是FileInfo对象的一维列表,并非嵌套列表。[item for sublist in list_for_dataframe for item in sublist]会错误地将每个FileInfo对象拆分为无效的单个元素,导致Spark无法匹配预设的Schema,最终生成全Null的DataFrame。
正确实现代码
方法一:通过字典列表转换(字段对应直观)
from pyspark.sql.types import StructType, StructField, StringType, LongType # 先获取FileInfo对象列表 file_list = mssparkutils.fs.ls("abfss://config@datalake.dfs.core.windows.net/") # 提取每个FileInfo的属性,生成字典列表 rows = [{"path": f.path, "name": f.name, "size": f.size} for f in file_list] schema = StructType([ StructField("path", StringType(), True), StructField("name", StringType(), True), StructField("size", LongType(), True) ]) df = spark.createDataFrame(rows, schema=schema) df.show()
方法二:通过元组列表转换(需与Schema字段顺序一致)
from pyspark.sql.types import StructType, StructField, StringType, LongType # 先获取FileInfo对象列表 file_list = mssparkutils.fs.ls("abfss://config@datalake.dfs.core.windows.net/") # 提取属性生成元组列表,顺序必须与Schema的path/name/size对应 rows = [(f.path, f.name, f.size) for f in file_list] schema = StructType([ StructField("path", StringType(), True), StructField("name", StringType(), True), StructField("size", LongType(), True) ]) df = spark.createDataFrame(rows, schema=schema) df.show()
效果说明
两种方法都能正确将FileInfo对象的属性映射到DataFrame的对应列,生成符合预期的包含path、name、size三列的数据集。
内容的提问来源于stack exchange,提问作者coding
相关产品推荐
相关产品推荐

