You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何读取Parquet文件中的dictionary_page值?代码正确性存疑

读取Parquet字典页的正确方法

你当前的代码问题在于,read_row_group是读取行组中的列数据,而非字典页本身,所以没法直接拿到字典值。要读取字典页,得用专门的处理方式:

  • 先通过col_meta.has_dictionary_page判断该列是否存在字典页
  • 如果存在,调用pq_file.read_dictionary_page(col_meta)读取字典页对象
  • 从字典页的data属性中提取字典值

修正后的代码示例:

for group in range(pq_file.metadata.num_row_groups):
    for col in range(pq_file.metadata.num_columns):
        col_meta = pq_file.metadata.row_group(group).column(col)
        # 判断当前列是否包含字典页
        if col_meta.has_dictionary_page:
            # 读取对应字典页
            dict_page = pq_file.read_dictionary_page(col_meta)
            # 提取字典值转为Python列表
            dict_values = dict_page.data.to_pylist()
            print(f"行组{group} 列{col}的字典值: {dict_values}")
        else:
            print(f"行组{group} 列{col}无字典页")

另外补充:dictionary_page_offset是字典页在文件中的字节偏移量,一般不需要你直接操作,read_dictionary_page会自动利用这个偏移量定位并读取对应内容。

内容的提问来源于stack exchange,提问作者naikordian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 16:45:25