You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Dask读取S3上Redshift导出的Parquet文件报schema重复字段错误

问题场景

通过UNLOAD命令将Redshift中的查询结果导出为Parquet格式存储到S3,导出语句如下:

UNLOAD
('SELECT
     delivered_at
    , flow_name
    , variant_name
    , user_id
') TO
            's3://data/raw/redshift/all_campaigns'
            IAM_ROLE 'arn:aws:iam::XYZ:role/redshift'
            FORMAT AS PARQUET
            maxfilesize 96 mb
            ALLOWOVERWRITE
            MANIFEST;

后续使用Dask框架读取对应S3路径下的Parquet分片文件,读取代码如下:

df = dd.read_parquet('s3://data/raw/redshift/all_campaigns0079_part_00.parquet')

实际执行时触发报错,完整报错栈信息如下:

/Users/fanbuch/Devel/NyRec/EDA.ipynb Cell 7' in <cell line: 1>()
----> 1 df = dd.read_parquet('s3://data/raw/redshift/crm/was_clicked/all_campaigns0079_part_00.parquet')

File ~/opt/miniconda3/envs/deeprec/lib/python3.9/site-packages/dask/dataframe/io/parquet/core.py:460, in read_parquet(path, columns, filters, categories, index, storage_options, engine, calculate_divisions, ignore_metadata_file, metadata_task_size, split_row_groups, chunksize, aggregate_files, parquet_file_extension, **kwargs)
    457 if index and isinstance(index, str):
    458     index = [index]
--> 460 read_metadata_result = engine.read_metadata(
    461     fs,
    462     paths,
    463     categories=categories,
    464     index=index,
    465     gather_statistics=calculate_divisions,
    466     filters=filters,
    467     split_row_groups=split_row_groups,
    468     chunksize=chunksize,
    469     aggregate_files=aggregate_files,
    470     ignore_metadata_file=ignore_metadata_file,
    471     metadata_task_size=metadata_task_size,
    472     parquet_file_extension=parquet_file_extension,
    473     **kwargs,
    474 )
    476 # In the future, we may want to give the engine the
    477 # option to return a dedicated element for `common_kwargs`.
...
File ~/opt/miniconda3/envs/deeprec/lib/python3.9/site-packages/pyarrow/error.pxi:144, in pyarrow.lib.pyarrow_internal_check_status()

File ~/opt/miniconda3/envs/deeprec/lib/python3.9/site-packages/pyarrow/error.pxi:100, in pyarrow.lib.check_status()

ArrowInvalid: Can't unify schema with duplicate field names.

已完成前置确认:

  • 导出数据使用的SELECT语句中不存在重复列名
  • 当前运行环境相关依赖版本如下:
dask version        : '2022.6.0'
pyarrow version     : '8.0.0'
fastparquet version : '0.8.1'

已就该问题在Dask官方GitHub仓库提交issue,现需要该问题的根因排查思路与可落地的解决方案。

内容的提问来源于stack exchange,提问作者Areza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 18:36:25