Spark SQL生成DataFrame后转CSV报错,求解决方法
问题修复:Spark DataFrame写入CSV报错(Interval类型不支持)
你的报错根源是查询生成的spend_time字段是时间间隔(Interval)类型,Spark的CSV输出组件不支持直接序列化这种类型到文本格式的CSV文件中。
修复方案1:将Interval类型转换为可序列化的类型
修改SQL查询,把(on_scene_datetime - request_datetime)转换为Spark支持写入CSV的类型,比如数值型秒数或字符串格式:
方式A:转换为总秒数(数值类型)
用unix_timestamp函数计算时间差,得到可写入的数值字段:
SELECT on_scene_datetime, request_datetime, unix_timestamp(on_scene_datetime) - unix_timestamp(request_datetime) as spend_time_seconds, trip_miles, sales_tax FROM ParquetTable WHERE on_scene_datetime IS NOT NULL AND request_datetime IS NOT NULL AND unix_timestamp(on_scene_datetime) - unix_timestamp(request_datetime) > 0 ORDER BY 4 DESC
方式B:转换为字符串格式
如果需要保留时间间隔的可读性,直接将Interval类型转为字符串:
SELECT on_scene_datetime, request_datetime, cast((on_scene_datetime - request_datetime) as string) as spend_time_str, trip_miles, sales_tax FROM ParquetTable WHERE on_scene_datetime IS NOT NULL AND request_datetime IS NOT NULL AND INT(on_scene_datetime - request_datetime) > 0 ORDER BY 4 DESC
修复方案2:移除不需要的Interval字段
如果业务上不需要spend_time字段,直接从查询中删除该字段即可:
SELECT on_scene_datetime, request_datetime, trip_miles, sales_tax FROM ParquetTable WHERE on_scene_datetime IS NOT NULL AND request_datetime IS NOT NULL AND INT(on_scene_datetime - request_datetime) > 0 ORDER BY 3 DESC
修改完成后,执行写入命令即可正常运行,还可以添加参数优化输出(比如写入表头、覆盖已有文件):
result.write.mode("overwrite").option("header", "true").csv("./result")
内容的提问来源于stack exchange,提问作者Parksubo
相关产品推荐
相关产品推荐

