You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Dask追加写入Parquet时遇ValueError:数据类型不匹配

Dask追加写入Parquet时类型不匹配问题

我在Python中使用Dask将多个大型DataFrame写入单个Parquet数据集,所有列类型均为float或string[pyarrow]。通过循环遍历Dask DataFrame,调用to_parquet方法实现首次写入覆盖、后续追加,代码如下:

ddf.to_parquet(options.output_prefix,
               append=True if n> 0 else False,
               engine="pyarrow",
               write_index=False,
               write_metadata_file=True,
               compute=True)

第二次调用时触发如下错误:

ValueError: Appended dtypes differ.
{('IJC_SAMPLE_1', 'object'), ('riExonStart_0base', dtype('float64')), ('riExonEnd', 'float64'), ('PValue', dtype('float64')), ('IncLevelDifference', 'float64'), ('chr', 'object'), ('SkipFormLen', dtype('float64')), ('ID', 'float64'), ('SJC_SAMPLE_1', 'object'), ('geneSymbol', 'object'), ('ID.1', 'float64'), ('downstreamEE', dtype('float64')), ('downstreamES', 'float64'), ('strand', 'object'), ('SJC_SAMPLE_2',
'object'), ('upstreamES', dtype('float64')), ('IncFormLen', 'float64'), ('FDR', 'float64'), ('IJC_SAMPLE_2', 'object'), ('ID.1', 'float64'), ('downstreamES', 'float64'), ('FDR', 'float64'), ('IJC_SAMPLE_1', string[pyarrow]), ('GeneID', 'object'), ('IncFormLen', dtype('float64')), ('chr', string[pyarrow]), ('GeneID', string[pyarrow]), ('IncLevel1', 'object'), ('SJC_SAMPLE_1', string[pyarrow]), ('geneSymbol', string[pyarrow]), ('riExonStart_0base', 'float64'), ('PValue', 'float64'), ('upstreamEE', dtype('float64')), ('IncLevelDifference', 'float64'), ('IncLevel1', string[pyarrow]), ('strand', string[pyarrow]), ('SJC_SAMPLE_2', string[pyarrow]), ('IncLevel2', 'object'), ('SkipFormLen', 'float64'), ('riExonEnd', dtype('float64')), ('IncLevel2', string[pyarrow]), ('upstreamEE', 'float64'), ('downstreamEE', 'float64'), ('ID', dtype('float64')), ('IJC_SAMPLE_2', string[pyarrow]), ('upstreamES', 'float64')}

我检查了两个DataFrame的数据类型,确认完全一致:

dtypes for frame 0:
ID                            float64 
GeneID                string[pyarrow] 
geneSymbol            string[pyarrow] 
chr                   string[pyarrow] 
strand                string[pyarrow] 
riExonStart_0base             float64 
riExonEnd                     float64 
upstreamES                    float64 
upstreamEE                    float64 
downstreamES                  float64 
downstreamEE                  float64 
ID.1                          float64 
IJC_SAMPLE_1          string[pyarrow] 
SJC_SAMPLE_1          string[pyarrow] 
IJC_SAMPLE_2          string[pyarrow] 
SJC_SAMPLE_2          string[pyarrow] 
IncFormLen                    float64 
SkipFormLen                   float64 
PValue                        float64 
FDR                           float64 
IncLevel1             string[pyarrow] 
IncLevel2             string[pyarrow] 
IncLevelDifference            float64 
dtype: object

dtypes for frame 1:
ID                            float64 
GeneID                string[pyarrow] 
geneSymbol            string[pyarrow] 
chr                   string[pyarrow] 
strand                string[pyarrow] 
riExonStart_0base             float64 
riExonEnd                     float64 
upstreamES                    float64 
upstreamEE                    float64 
downstreamES                  float64 
downstreamEE                  float64 
ID.1                          float64 
IJC_SAMPLE_1          string[pyarrow] 
SJC_SAMPLE_1          string[pyarrow] 
IJC_SAMPLE_2          string[pyarrow] 
SJC_SAMPLE_2          string[pyarrow] 
IncFormLen                    float64 
SkipFormLen                   float64 
PValue                        float64 
FDR                           float64 
IncLevel1             string[pyarrow] 
IncLevel2             string[pyarrow] 
IncLevelDifference            float64 
dtype: object

问题根源应该是首次写入时,string[pyarrow]列被存储为object类型,后续追加时Dask/PyArrow判定类型不匹配。我尝试过设置to_parquet的schema参数,将string[pyarrow]列指定为object、np.object_等,但都没有效果。


内容的提问来源于stack exchange,提问作者Ian Sudbery

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 03:34:53