将大型Pandas DataFrame转为Spark DataFrame时遇类型错误求解决
问题解决思路
问题本质
你定义的Spark Schema中,author字段是字符串数组类型(ArrayType(StringType())),但实际Pandas DataFrame的author列里混入了普通字符串值(错误信息里的长文本就是典型例子)。小数据集刚好没碰到这种异常数据所以正常运行,大数据集里出现了类型不匹配,就触发了报错。
解决步骤
- 确认异常数据类型
先执行以下代码,检查author列的所有数据类型:
print(data['author'].apply(type).unique())
执行后会输出类似[<class 'str'>, <class 'list'>]的结果,验证确实存在混合类型。
- 统一
author列的格式为数组
把所有值转成数组格式:单个作者的字符串变成单元素数组,原本是数组的保持不变,空值按需处理:
def fix_author_format(author_val): if isinstance(author_val, str): # 单个作者字符串转成数组 return [author_val] elif isinstance(author_val, list): # 已经是数组,直接返回 return author_val else: # 空值或其他类型,返回空数组(也可以改成return None) return [] # 批量处理整列 data['author'] = data['author'].apply(fix_author_format)
- 重新创建Spark DataFrame
处理完格式后,再执行你原来的创建代码即可:
sparkDF = spark.createDataFrame(data, df_schema) sparkDF.printSchema()
大数据集优化提示
如果用apply处理600多万行速度慢,可以改用矢量化操作提升效率:
import numpy as np data['author'] = np.where( data['author'].map(lambda x: isinstance(x, str)), data['author'].map(lambda x: [x]), data['author'] )
内容的提问来源于stack exchange,提问作者Lucas
相关产品推荐
相关产品推荐

