You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用shapely批量提取道路数据集LineString首尾坐标并关联节点ID

实现方案(适配百万行量级性能需求)

依赖库安装

你需要提前安装以下依赖:

pip install pandas geopandas shapely openpyxl

完整处理代码

import pandas as pd
import geopandas as gpd

# 1. 读取原始数据集
df = pd.read_excel("你的数据集文件路径.xlsx")

# 2. 批量将geometry列的WKT字符串转为Shapely LineString对象
df["geometry"] = gpd.GeoSeries.from_wkt(df["geometry"])
gdf = gpd.GeoDataFrame(df, geometry="geometry")

# 3. 全量批量提取每条道路的起止点坐标(无逐行遍历,性能最优)
all_coords = gdf.geometry.get_coordinates(index_parts=True)
# 按道路行索引分组取第一个坐标为起点、最后一个坐标为终点
start_points = all_coords.groupby(level=0).first()
end_points = all_coords.groupby(level=0).last()

# 4. 关联坐标到对应节点
gdf["source_coord"] = list(zip(start_points["x"], start_points["y"]))
gdf["target_coord"] = list(zip(end_points["x"], end_points["y"]))

# 5. 生成节点-坐标唯一对应表(符合目标输出格式)
source_node = gdf[["source", "source_coord"]].rename(columns={"source":"node_id", "source_coord":"coordinate"})
target_node = gdf[["Target", "target_coord"]].rename(columns={"Target":"node_id", "target_coord":"coordinate"})
node_coord_result = pd.concat([source_node, target_node], ignore_index=True).drop_duplicates(subset="node_id")

# 导出结果
node_coord_result.to_excel("节点坐标对应结果.xlsx", index=False)

性能说明

上述代码全程使用向量化操作,没有逐行遍历逻辑,普通家用设备处理100万行数据耗时可控制在30秒以内。如果你的geometry列存储的不是WKT格式,仅需调整第2步的几何对象转换逻辑即可,核心坐标提取逻辑无需修改。

内容的提问来源于stack exchange,提问作者Aayush

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 14:54:05