如何遍历两个GeoDataFrame统计线要素相交数量并添加新列?
统计GeoDataFrame中线要素的相交数量
基础实现方法(简洁高效,适合小数据量)
直接使用GeoPandas的空间连接(sjoin)结合分组计数,无需手动遍历:
import geopandas as gpd from shapely.geometry import LineString # 构造示例数据 df = gpd.GeoDataFrame([['a',LineString([(1, 0.25), (2,1.25)])], ['b', LineString([(1.2, 1.0), (1.4, 1.5)])]], columns = ['name','geometry']) df2 = gpd.GeoDataFrame([['c', LineString([(2.0, 0.5), (1.0, .75)])], ['d', LineString([(2, 0.75), (1.25, 0.8)])]], columns = ['name', 'geometry']) # 空间连接,筛选相交的要素对 joined = gpd.sjoin(df, df2, how='left', predicate='intersects') # 按原df的索引分组,统计每条线的相交数量 df['intersect_count'] = joined.groupby(joined.index).size() # 无相交的线填充0 df['intersect_count'] = df['intersect_count'].fillna(0).astype(int) print(df)
运行后会得到你期望的结果:为df新增intersect_count列,记录每条线与df2中线要素的相交数量。
优化实现方法(适合大数据量)
如果数据量较大,直接空间连接效率偏低,可以结合空间索引和缓冲区过滤减少计算量:
import geopandas as gpd from shapely.geometry import LineString # 构造示例数据 df = gpd.GeoDataFrame([['a',LineString([(1, 0.25), (2,1.25)])], ['b', LineString([(1.2, 1.0), (1.4, 1.5)])]], columns = ['name','geometry']) df2 = gpd.GeoDataFrame([['c', LineString([(2.0, 0.5), (1.0, .75)])], ['d', LineString([(2, 0.75), (1.25, 0.8)])]], columns = ['name', 'geometry']) # 为df2创建空间索引,加速候选要素查询 sindex = df2.sindex def count_intersections(row): # 给当前线添加极小缓冲区,扩大查询范围(缓冲区大小需匹配数据精度) buffer = row.geometry.buffer(0.001) # 通过空间索引获取可能相交的候选要素索引 candidate_idx = list(sindex.intersection(buffer.bounds)) candidates = df2.iloc[candidate_idx] # 对候选要素做精确相交判断 matches = candidates[candidates.intersects(row.geometry)] return len(matches) # 批量计算每条线的相交数量 df['intersect_count'] = df.apply(count_intersections, axis=1) print(df)
你之前用缓冲区结果为空,大概率是缓冲区大小设置不合理——如果缓冲区太小,可能漏掉实际相交的要素;如果太大,过滤效果差。需要根据你的数据坐标精度调整buffer()的参数值。
内容的提问来源于stack exchange,提问作者kp42
相关产品推荐
相关产品推荐

