使用shapely批量提取道路数据集LineString首尾坐标并关联节点ID
实现方案(适配百万行量级性能需求)
依赖库安装
你需要提前安装以下依赖:
pip install pandas geopandas shapely openpyxl
完整处理代码
import pandas as pd import geopandas as gpd # 1. 读取原始数据集 df = pd.read_excel("你的数据集文件路径.xlsx") # 2. 批量将geometry列的WKT字符串转为Shapely LineString对象 df["geometry"] = gpd.GeoSeries.from_wkt(df["geometry"]) gdf = gpd.GeoDataFrame(df, geometry="geometry") # 3. 全量批量提取每条道路的起止点坐标(无逐行遍历,性能最优) all_coords = gdf.geometry.get_coordinates(index_parts=True) # 按道路行索引分组取第一个坐标为起点、最后一个坐标为终点 start_points = all_coords.groupby(level=0).first() end_points = all_coords.groupby(level=0).last() # 4. 关联坐标到对应节点 gdf["source_coord"] = list(zip(start_points["x"], start_points["y"])) gdf["target_coord"] = list(zip(end_points["x"], end_points["y"])) # 5. 生成节点-坐标唯一对应表(符合目标输出格式) source_node = gdf[["source", "source_coord"]].rename(columns={"source":"node_id", "source_coord":"coordinate"}) target_node = gdf[["Target", "target_coord"]].rename(columns={"Target":"node_id", "target_coord":"coordinate"}) node_coord_result = pd.concat([source_node, target_node], ignore_index=True).drop_duplicates(subset="node_id") # 导出结果 node_coord_result.to_excel("节点坐标对应结果.xlsx", index=False)
性能说明
上述代码全程使用向量化操作,没有逐行遍历逻辑,普通家用设备处理100万行数据耗时可控制在30秒以内。如果你的geometry列存储的不是WKT格式,仅需调整第2步的几何对象转换逻辑即可,核心坐标提取逻辑无需修改。
内容的提问来源于stack exchange,提问作者Aayush
相关产品推荐
相关产品推荐

