You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在读取Parquet文件时为几何列应用PyArrow过滤谓词?

Parquet读取时如何用Arrow过滤几何列?

用pandas.read_parquet读取Parquet文件的DataFrame时,可借助pyarrow的filters API只读取行子集,geopandas也支持这种方式,示例:

gdf = geopandas.read_parquet(files, filters=[("station", "in", ["sjisjka", "kaitum"])])

但想把这种过滤逻辑应用到几何列时,目前只能先全量读取再显式过滤:

gdf = geopandas.read_parquet(files)
pol = shapely.geometry.Polygon(((65, 15), (65, 25), (70, 25), (70, 15), (65, 15)))
selection = gdf[gdf["location"].within(pol)]

这种方式速度很慢,哪怕用dask/dask-geopandas也没多大改善。

尝试直接用Arrow的filters做空间检索:

gdf = dask_geopandas.read_parquet(files, filters=[("location", "within", pol)])

但运行报错:

ValueError: "('location', 'within', <shapely.geometry.polygon.Polygon object at 0x7fd5ffa7ce80>)" is not a valid operator in predicates.

可行解决方案

目前pyarrow的Parquet过滤API不支持直接对几何列使用空间谓词(比如within),因为Parquet本身没有原生空间索引,pyarrow的filters只支持常规比较运算符(=, !=, <, >, in等),无法解析空间对象和空间操作。

要实现高效的空间过滤,有两种实用方案:

  1. 预生成边界列做粗过滤
    在写入Parquet前,提前计算几何列的外接矩形边界,把边界的最小/最大经纬度存成单独列(比如min_x, max_x, min_y, max_y)。读取时先用pyarrow filters过滤出边界与目标多边形有交集的行,再在内存中做精确的within过滤,能大幅减少加载的数据量:

    # 写入前预处理示例
    gdf["min_x"] = gdf["location"].bounds.minx
    gdf["max_x"] = gdf["location"].bounds.maxx
    gdf["min_y"] = gdf["location"].bounds.miny
    gdf["max_y"] = gdf["location"].bounds.maxy
    gdf.to_parquet("spatial_data.parquet")
    
    # 读取时先粗过滤,再精确筛选
    pol_bounds = pol.bounds  # 得到(minx, miny, maxx, maxy)
    gdf = dask_geopandas.read_parquet(
        "spatial_data.parquet",
        filters=[
            ("max_x", ">", pol_bounds[0]),
            ("min_x", "<", pol_bounds[2]),
            ("max_y", ">", pol_bounds[1]),
            ("min_y", "<", pol_bounds[3])
        ]
    )
    # 最后做精确的空间判断
    selection = gdf[gdf["location"].within(pol)]
    
  2. 用DuckDB直接执行空间查询
    DuckDB的空间扩展支持直接对Parquet中的几何列执行空间过滤,无需预生成额外列,能利用Parquet的列式存储特性加速查询:

    import duckdb
    import geopandas as gpd
    
    # 初始化DuckDB并加载空间扩展
    con = duckdb.connect()
    con.execute("INSTALL spatial;")
    con.execute("LOAD spatial;")
    
    # 直接查询Parquet文件,过滤符合空间条件的行
    query = """
        SELECT * FROM 'spatial_data.parquet'
        WHERE ST_Within(location, ST_GeomFromText('POLYGON ((65 15, 65 25, 70 25, 70 15, 65 15))'))
    """
    df = con.execute(query).df()
    # 转换为GeoDataFrame
    gdf = gpd.GeoDataFrame(df, geometry="location")
    

内容的提问来源于stack exchange,提问作者gerrit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 13:50:27