You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark使用xpath提取XML元素返回空值数组的原因排查

问题原因与解决方案

为什么ContactName和PhoneNo返回空值数组?

你使用的XPath表达式仅指向了XML节点(比如/Root/Customers/Customer/Name),但没有明确指定提取节点的文本内容。PySpark的xpath函数在遇到这种仅指向节点的表达式时,返回的是元素节点对象,而非节点内的文本,因此在DataFrame中显示为null数组。

xpath函数能否提取XML节点的文本内容?

可以,但需要在XPath表达式末尾添加/text()来明确获取节点的文本值。

修正后的代码

将selectExpr中的XPath表达式修改为以下形式:

df_sample_CustomersOrders1 = df_Customers_Orders.selectExpr(
    "xpath(Data,'/Root/Customers/Customer/@CustomerID') as CustomerID",
    "xpath(Data,'/Root/Customers/Customer/Name/text()') as ContactName",
    "xpath(Data,'/Root/Customers/Customer/PhoneNo/text()') as PhoneNo",
)

可选:将数组展开为单行记录

如果希望每个客户的信息单独占一行(而非一行包含所有客户的数组),可以使用explode函数:

from pyspark.sql.functions import explode

# 展开数组,得到每行对应一个客户的记录
df_expanded = df_sample_CustomersOrders1.select(
    explode("CustomerID").alias("CustomerID"),
    explode("ContactName").alias("ContactName"),
    explode("PhoneNo").alias("PhoneNo")
)
df_expanded.show(truncate=False)

补充说明

你当前结果中CustomerID数组的内容([GREAL, HUNGC, LAZYK, LETSS])和提供的XML示例不符,推测是实际读取的CSV文件中的XML内容与示例不一致,但上述XPath的修正方法依然适用。

内容的提问来源于stack exchange,提问作者Govind Sajeev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 06:11:13