You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas read_parquet()带filters读取分区int64列触发ArrowNotImplementedError

解决Pandas read_parquet分区列类型推断错误导致的过滤失败问题

问题原因

Parquet分区列的类型会从目录结构的字符串名称推断,当分区目录命名为partition_column_name=123这类格式时,Pandas默认可能将其解析为字符串类型,而实际分区列存储的是int64类型,导致过滤时出现类型不匹配的ArrowNotImplementedError。

解决方法

方法1:用pyarrow Dataset显式指定schema

直接通过pyarrow的Dataset API定义分区列的正确类型,避免Pandas自动推断错误:

import pandas as pd
import pyarrow as pa
import pyarrow.dataset as ds

# 定义分区列的schema,指定为int64类型
schema = pa.schema([('partition_column_name', pa.int64())])

# 创建Dataset并指定schema与分区格式
dataset = ds.dataset(
    'path_to_file.parquet',
    format='parquet',
    schema=schema,
    partitioning='hive'  # 适配Hive风格的分区目录格式
)

# 应用过滤条件并转换为DataFrame
df = dataset.to_table(filter=ds.field('partition_column_name') == 123).to_pandas()

方法2:在Pandas read_parquet中指定schema与engine

如果坚持使用Pandas原生的read_parquet,可以通过engine='pyarrow'结合schema参数强制指定类型:

import pandas as pd
import pyarrow as pa

# 定义包含分区列的schema
schema = pa.schema([('partition_column_name', pa.int64())])

df = pd.read_parquet(
    'path_to_file.parquet',
    filters=[('partition_column_name', '==', 123)],
    engine='pyarrow',
    schema=schema
)

方法3:临时将过滤值转为字符串(应急方案)

若上述方法暂时无法生效,可临时将过滤条件的整数转为字符串,匹配Pandas错误推断的类型,但这仅为权宜之计,无法解决根本问题:

import pandas as pd

df = pd.read_parquet('path_to_file.parquet', filters=[('partition_column_name', '==', '123')])

内容的提问来源于stack exchange,提问作者Rafael Higa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 19:16:12