PySpark使用xpath提取XML元素返回空值数组的原因排查
问题原因与解决方案
为什么ContactName和PhoneNo返回空值数组?
你使用的XPath表达式仅指向了XML节点(比如/Root/Customers/Customer/Name),但没有明确指定提取节点的文本内容。PySpark的xpath函数在遇到这种仅指向节点的表达式时,返回的是元素节点对象,而非节点内的文本,因此在DataFrame中显示为null数组。
xpath函数能否提取XML节点的文本内容?
可以,但需要在XPath表达式末尾添加/text()来明确获取节点的文本值。
修正后的代码
将selectExpr中的XPath表达式修改为以下形式:
df_sample_CustomersOrders1 = df_Customers_Orders.selectExpr( "xpath(Data,'/Root/Customers/Customer/@CustomerID') as CustomerID", "xpath(Data,'/Root/Customers/Customer/Name/text()') as ContactName", "xpath(Data,'/Root/Customers/Customer/PhoneNo/text()') as PhoneNo", )
可选:将数组展开为单行记录
如果希望每个客户的信息单独占一行(而非一行包含所有客户的数组),可以使用explode函数:
from pyspark.sql.functions import explode # 展开数组,得到每行对应一个客户的记录 df_expanded = df_sample_CustomersOrders1.select( explode("CustomerID").alias("CustomerID"), explode("ContactName").alias("ContactName"), explode("PhoneNo").alias("PhoneNo") ) df_expanded.show(truncate=False)
补充说明
你当前结果中CustomerID数组的内容([GREAL, HUNGC, LAZYK, LETSS])和提供的XML示例不符,推测是实际读取的CSV文件中的XML内容与示例不一致,但上述XPath的修正方法依然适用。
内容的提问来源于stack exchange,提问作者Govind Sajeev
相关产品推荐
相关产品推荐

