Spark RDD Map报错TypeError:需类字节对象而非Row,XML转JSON求助
解决TypeError: a bytes-like object is required, not 'Row'问题
看起来你是在Spark DataFrame里处理XML转JSON时踩了个常见的坑——直接把整个Row对象传给了解析函数,而没有提取出XML_Data列里的实际XML内容。咱们一步步来搞定它:
1. 先搞清楚Row对象的结构
当你用testing.take(1)时,返回的是一个包含单个Row的列表,比如:
[Row(XML_Data='<root><name>test</name></root>')]
你需要先提取Row里的XML_Data字段,才能拿到真正要解析的XML字符串/字节,而不是把整个Row丢给xmlparse函数。
2. 修正你的解析函数调用(以Spark UDF为例)
假设你之前的代码是直接把Row传给UDF,现在要改成针对XML_Data列单独处理:
第一步:定义正确的XML转JSON函数
先处理可能的字节/字符串转换:如果你的XML_Data是bytes类型,记得先解码成字符串;如果已经是字符串就直接用。
import xml.etree.ElementTree as ET import json def xml_to_json(xml_content): # 处理字节类型的XML内容 if isinstance(xml_content, bytes): xml_content = xml_content.decode('utf-8') # 解析XML root = ET.fromstring(xml_content) # 这里可以根据你的XML结构自定义转JSON的逻辑,示例是简单递归转字典 def elem_to_dict(elem): result = {} for child in elem: child_dict = elem_to_dict(child) if child.tag in result: if not isinstance(result[child.tag], list): result[child.tag] = [result[child.tag]] result[child.tag].append(child_dict) else: result[child.tag] = child_dict if elem.text and elem.text.strip(): if result: result['text'] = elem.text.strip() else: result = elem.text.strip() return result return json.dumps(elem_to_dict(root))
第二步:注册UDF并应用到DataFrame
用Spark的UDF把这个函数映射到XML_Data列,而不是整个Row:
from pyspark.sql.functions import udf from pyspark.sql.types import StringType # 注册UDF xml_to_json_udf = udf(xml_to_json, StringType()) # 应用到DataFrame,生成新的JSON列 result_df = testing.withColumn('JSON_Data', xml_to_json_udf(testing['XML_Data']))
3. 针对60k行的性能优化
如果你的DataFrame有60k行,普通UDF可能效率不够,建议用Pandas UDF来加速:
from pyspark.sql.functions import pandas_udf import pandas as pd @pandas_udf(StringType()) def xml_to_json_pandas(xml_series: pd.Series) -> pd.Series: def process_xml(xml_content): if isinstance(xml_content, bytes): xml_content = xml_content.decode('utf-8') root = ET.fromstring(xml_content) # 复用上面的elem_to_dict函数 return json.dumps(elem_to_dict(root)) return xml_series.apply(process_xml) # 应用Pandas UDF result_df = testing.withColumn('JSON_Data', xml_to_json_pandas(testing['XML_Data']))
4. 调试小技巧
在写函数前,先单独提取一行测试,确认你的XML内容格式:
# 取第一行的Row对象 sample_row = testing.take(1)[0] # 提取XML内容 sample_xml = sample_row['XML_Data'] print(type(sample_xml)) # 看是str还是bytes print(sample_xml) # 确认XML内容是否正常
这样就能确保你的解析函数拿到的是正确的字符串/字节,而不是Row对象,解决那个TypeError问题啦。
内容的提问来源于stack exchange,提问作者mdeonte001
相关产品推荐
相关产品推荐

