遍历行政区多边形匹配内部点时仅处理单个区域的问题排查
问题:遍历行政区多边形匹配内部点时重复处理单个区域
数据集说明
- admin_area.shp:行政区多边形数据,字段包含
index、names、geometry - point_data.shp:点数据,字段包含
index、lat、lon、geometry
需求
遍历admin_area中的每个多边形,识别每个区域内的相交点。
尝试的代码
# initialise an instance of an rtree Index object idx = Index() # loop through each row in the district dataset and extract the geometry of each one for id, district in admin_areas.iterrows(): idx.insert(id, district.geometry.bounds) # loop through each row in the point dataset and load into the spatial index for id, point in point_data.iterrows(): idx.insert(id, point.geometry.bounds) # get the polygon of each district polygon = admin_areas.geometry.iloc[0] # how many rows are we starting with? print(f"boundaries: {len(admin_areas.index)}") # get the indexes of points that intersect bounds of the admin_areas possible_matches_index = list(idx.intersection(polygon.bounds)) # use those indexes to extract the possible matches from the GeoDataFrame possible_matches = point_data.iloc[possible_matches_index] # how many points intersect the area? print(f"Filtered points: {len(possible_matches.index)}") # then search the possible matches for precise matches using the slower but more precise method precise_matches = possible_matches.loc[possible_matches.within(polygon)] print(precise_matches.geometry) # how many points itersect the area now? print(f"precise points: {len(precise_matches.index)}")
运行结果(重复片段)
boundaries: 9 Filtered points: 111 647 POINT (371053.697 410002.844) 194 POINT (371053.697 410002.844) 21 POINT (371053.697 410002.844) 189 POINT (371053.697 410002.844) 640 POINT (371053.697 410002.844) 973 POINT (371848.442 409174.713) 982 POINT (371848.442 409174.713) 989 POINT (371848.442 409174.713) 990 POINT (371848.442 409174.713) 995 POINT (371848.442 409174.713) Name: geometry, Length: 110, dtype: geometry precise points: 110 boundaries: 9 Filtered points: 222 ...(重复输出类似内容)
问题原因
- 固定读取第一个多边形:代码中
polygon = admin_areas.geometry.iloc[0]始终取数据集的第一个行政区多边形,不管当前循环到哪个区域,导致每次都重复处理同一个行政区。 - 空间索引重复插入:在行政区循环内部重复遍历点数据并插入索引,每次循环都会把所有点再插入一遍,导致索引内的点数量翻倍,匹配结果混乱。
- 索引未合理规划:行政区的索引也被重复插入到同一个空间索引中,干扰后续的边界匹配逻辑。
修正后的代码
# 提前构建点数据的空间索引,仅执行一次 idx = Index() for point_id, point in point_data.iterrows(): idx.insert(point_id, point.geometry.bounds) # 遍历每个行政区 for district_id, district in admin_areas.iterrows(): polygon = district.geometry # 取当前循环的行政区多边形 print(f"当前处理行政区: {district['names']} (ID: {district_id})") print(f"总行政区数量: {len(admin_areas.index)}") # 获取与当前多边形边界相交的点索引 possible_matches_index = list(idx.intersection(polygon.bounds)) possible_matches = point_data.iloc[possible_matches_index] print(f"初步筛选点数量: {len(possible_matches.index)}") # 精确匹配:筛选真正在多边形内部的点 precise_matches = possible_matches[possible_matches.within(polygon)] print(f"精确匹配点数量: {len(precise_matches.index)}") print(precise_matches.geometry) print("---")
修正说明
- 预构建点索引:将点数据的空间索引构建移到行政区循环外,避免重复插入造成的索引膨胀。
- 使用当前循环的多边形:用
district.geometry替代固定取第一个多边形,确保每次处理当前遍历的行政区。 - 优化匹配逻辑:单个空间索引仅存储点数据,用当前行政区的边界直接匹配,避免索引污染。
- 增加进度标识:打印当前处理的行政区名称和ID,方便跟踪遍历进度。
内容的提问来源于stack exchange,提问作者Mailiaa
相关产品推荐
相关产品推荐

