You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

遍历行政区多边形匹配内部点时仅处理单个区域的问题排查

问题:遍历行政区多边形匹配内部点时重复处理单个区域

数据集说明

  • admin_area.shp:行政区多边形数据,字段包含index、names、geometry
  • point_data.shp:点数据,字段包含index、lat、lon、geometry

需求

遍历admin_area中的每个多边形,识别每个区域内的相交点。

尝试的代码

# initialise an instance of an rtree Index object
idx = Index()

# loop through each row in the district dataset and extract the geometry of each one
for id, district in admin_areas.iterrows():
    idx.insert(id, district.geometry.bounds)
    
    # loop through each row in the point dataset and load into the spatial index
    for id, point in point_data.iterrows():
        idx.insert(id, point.geometry.bounds)
    
    # get the polygon of each district
    polygon = admin_areas.geometry.iloc[0]

    # how many rows are we starting with?
    print(f"boundaries: {len(admin_areas.index)}")

    # get the indexes of points that intersect bounds of the admin_areas
    possible_matches_index = list(idx.intersection(polygon.bounds))

    # use those indexes to extract the possible matches from the GeoDataFrame
    possible_matches = point_data.iloc[possible_matches_index]
    
    # how many points intersect the area? 
    print(f"Filtered points: {len(possible_matches.index)}")

    # then search the possible matches for precise matches using the slower but more precise method
    precise_matches = possible_matches.loc[possible_matches.within(polygon)]
    print(precise_matches.geometry)

    # how many points itersect the area now?
    print(f"precise points: {len(precise_matches.index)}")

运行结果(重复片段)

boundaries: 9
Filtered points: 111
647    POINT (371053.697 410002.844)
194    POINT (371053.697 410002.844)
21     POINT (371053.697 410002.844)
189    POINT (371053.697 410002.844)
640    POINT (371053.697 410002.844)
             
973    POINT (371848.442 409174.713)
982    POINT (371848.442 409174.713)
989    POINT (371848.442 409174.713)
990    POINT (371848.442 409174.713)
995    POINT (371848.442 409174.713)
Name: geometry, Length: 110, dtype: geometry
precise points: 110
boundaries: 9
Filtered points: 222
...(重复输出类似内容)

问题原因

  1. 固定读取第一个多边形:代码中polygon = admin_areas.geometry.iloc[0]始终取数据集的第一个行政区多边形,不管当前循环到哪个区域,导致每次都重复处理同一个行政区。
  2. 空间索引重复插入:在行政区循环内部重复遍历点数据并插入索引,每次循环都会把所有点再插入一遍,导致索引内的点数量翻倍,匹配结果混乱。
  3. 索引未合理规划:行政区的索引也被重复插入到同一个空间索引中,干扰后续的边界匹配逻辑。

修正后的代码

# 提前构建点数据的空间索引,仅执行一次
idx = Index()
for point_id, point in point_data.iterrows():
    idx.insert(point_id, point.geometry.bounds)

# 遍历每个行政区
for district_id, district in admin_areas.iterrows():
    polygon = district.geometry  # 取当前循环的行政区多边形
    print(f"当前处理行政区: {district['names']} (ID: {district_id})")
    print(f"总行政区数量: {len(admin_areas.index)}")

    # 获取与当前多边形边界相交的点索引
    possible_matches_index = list(idx.intersection(polygon.bounds))
    possible_matches = point_data.iloc[possible_matches_index]
    
    print(f"初步筛选点数量: {len(possible_matches.index)}")

    # 精确匹配:筛选真正在多边形内部的点
    precise_matches = possible_matches[possible_matches.within(polygon)]
    print(f"精确匹配点数量: {len(precise_matches.index)}")
    print(precise_matches.geometry)
    print("---")

修正说明

  • 预构建点索引:将点数据的空间索引构建移到行政区循环外,避免重复插入造成的索引膨胀。
  • 使用当前循环的多边形:用district.geometry替代固定取第一个多边形,确保每次处理当前遍历的行政区。
  • 优化匹配逻辑:单个空间索引仅存储点数据,用当前行政区的边界直接匹配,避免索引污染。
  • 增加进度标识:打印当前处理的行政区名称和ID,方便跟踪遍历进度。

内容的提问来源于stack exchange,提问作者Mailiaa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 19:25:14