Dataverse导出CSV转Parquet失败:Spark行分隔符解析异常
问题描述
从Dynamics 365 Dataverse导出至ADLS Gen2的CSV文件,通过Spark读取并转换为Parquet时遭遇解析错误。相关代码及报错如下:
读取CSV的Spark代码
df2 = (spark.read.format('csv') .option("delimiter", ",") .option("quote", '"') #.option("quoteAll", '"') .option("escape", '"') .option("header", "false") .option("path", '/mnt/d365/'+absolute+'/'+table_name+"/*.csv") .option("mode", "failfast") #.option("mode", "dropmalformed") #.option("mode", "permissive") .option("lineSep", "\r\n") .option("multiLine", "true") .schema(schema) .load() )
转换为Parquet的代码
for table in dataverse_dfs.keys(): try: (dataverse_dfs[table].write .mode('overwrite') .parquet(f'abfss://bronze@{storage_account_name_write}.dfs.core.windows.net/D365/{table}')) except Exception as e: print(e) raise Exception(e)
报错信息
Caused by: com.univocity.parsers.common.TextParsingException: java.lang.IllegalStateException - Error reading from input
Identified line separator characters in the parsed content. This may be the cause of the error. The line separator in your parser settings is set to '[crlf]'. Parsed content:
已尝试将lineSep改为"\n",问题仍未解决。
解决方案
- 自动检测换行符:Spark 3.0及以上版本支持
lineSep设为"auto",让解析器自动识别文件中的换行符类型,替换原有的固定换行符配置:.option("lineSep", "auto") - 排查实际换行符类型:下载样本CSV文件,用Notepad++等工具查看实际换行符(可能存在
\r单独使用或多种换行符混合的情况),再针对性设置lineSep。 - 调整解析模式与多行配置:如果CSV字段值中不包含换行符,关闭
multiLine并切换到permissive模式先定位问题:.option("multiLine", "false") .option("mode", "permissive") - 校验Schema匹配度:确认自定义的schema与CSV实际字段的数量、数据类型完全一致,字段不匹配也会触发解析异常。
- 清理特殊字符:添加忽略首尾空格的选项,避免Dataverse导出的CSV中存在的不可见字符干扰解析:
.option("ignoreLeadingWhiteSpace", "true") .option("ignoreTrailingWhiteSpace", "true")
内容的提问来源于stack exchange,提问作者Greencolor
相关产品推荐
相关产品推荐

