使用Dask追加写入Parquet时遇ValueError:数据类型不匹配
Dask追加写入Parquet时类型不匹配问题
我在Python中使用Dask将多个大型DataFrame写入单个Parquet数据集,所有列类型均为float或string[pyarrow]。通过循环遍历Dask DataFrame,调用to_parquet方法实现首次写入覆盖、后续追加,代码如下:
ddf.to_parquet(options.output_prefix, append=True if n> 0 else False, engine="pyarrow", write_index=False, write_metadata_file=True, compute=True)
第二次调用时触发如下错误:
ValueError: Appended dtypes differ. {('IJC_SAMPLE_1', 'object'), ('riExonStart_0base', dtype('float64')), ('riExonEnd', 'float64'), ('PValue', dtype('float64')), ('IncLevelDifference', 'float64'), ('chr', 'object'), ('SkipFormLen', dtype('float64')), ('ID', 'float64'), ('SJC_SAMPLE_1', 'object'), ('geneSymbol', 'object'), ('ID.1', 'float64'), ('downstreamEE', dtype('float64')), ('downstreamES', 'float64'), ('strand', 'object'), ('SJC_SAMPLE_2', 'object'), ('upstreamES', dtype('float64')), ('IncFormLen', 'float64'), ('FDR', 'float64'), ('IJC_SAMPLE_2', 'object'), ('ID.1', 'float64'), ('downstreamES', 'float64'), ('FDR', 'float64'), ('IJC_SAMPLE_1', string[pyarrow]), ('GeneID', 'object'), ('IncFormLen', dtype('float64')), ('chr', string[pyarrow]), ('GeneID', string[pyarrow]), ('IncLevel1', 'object'), ('SJC_SAMPLE_1', string[pyarrow]), ('geneSymbol', string[pyarrow]), ('riExonStart_0base', 'float64'), ('PValue', 'float64'), ('upstreamEE', dtype('float64')), ('IncLevelDifference', 'float64'), ('IncLevel1', string[pyarrow]), ('strand', string[pyarrow]), ('SJC_SAMPLE_2', string[pyarrow]), ('IncLevel2', 'object'), ('SkipFormLen', 'float64'), ('riExonEnd', dtype('float64')), ('IncLevel2', string[pyarrow]), ('upstreamEE', 'float64'), ('downstreamEE', 'float64'), ('ID', dtype('float64')), ('IJC_SAMPLE_2', string[pyarrow]), ('upstreamES', 'float64')}
我检查了两个DataFrame的数据类型,确认完全一致:
dtypes for frame 0: ID float64 GeneID string[pyarrow] geneSymbol string[pyarrow] chr string[pyarrow] strand string[pyarrow] riExonStart_0base float64 riExonEnd float64 upstreamES float64 upstreamEE float64 downstreamES float64 downstreamEE float64 ID.1 float64 IJC_SAMPLE_1 string[pyarrow] SJC_SAMPLE_1 string[pyarrow] IJC_SAMPLE_2 string[pyarrow] SJC_SAMPLE_2 string[pyarrow] IncFormLen float64 SkipFormLen float64 PValue float64 FDR float64 IncLevel1 string[pyarrow] IncLevel2 string[pyarrow] IncLevelDifference float64 dtype: object dtypes for frame 1: ID float64 GeneID string[pyarrow] geneSymbol string[pyarrow] chr string[pyarrow] strand string[pyarrow] riExonStart_0base float64 riExonEnd float64 upstreamES float64 upstreamEE float64 downstreamES float64 downstreamEE float64 ID.1 float64 IJC_SAMPLE_1 string[pyarrow] SJC_SAMPLE_1 string[pyarrow] IJC_SAMPLE_2 string[pyarrow] SJC_SAMPLE_2 string[pyarrow] IncFormLen float64 SkipFormLen float64 PValue float64 FDR float64 IncLevel1 string[pyarrow] IncLevel2 string[pyarrow] IncLevelDifference float64 dtype: object
问题根源应该是首次写入时,string[pyarrow]列被存储为object类型,后续追加时Dask/PyArrow判定类型不匹配。我尝试过设置to_parquet的schema参数,将string[pyarrow]列指定为object、np.object_等,但都没有效果。
内容的提问来源于stack exchange,提问作者Ian Sudbery
相关产品推荐
相关产品推荐

