Python实现CSV经纬度点匹配最近OpenStreetMap建筑数据
解决思路与修正代码
原代码核心问题在于:错误使用OSMnx查找路网节点而非建筑,用生成器处理CSV未完成建筑匹配逻辑,且建筑数据的存储与检索方式效率低下。以下是优化后的实现方案,基于Geopandas和sklearn的BallTree完成高效的近邻建筑匹配:
完整代码
import geopandas as gpd import numpy as np from sklearn.neighbors import BallTree from pyrosm import OSM, get_data import pandas as pd import warnings from shapely.errors import ShapelyDeprecationWarning # 屏蔽无关警告 warnings.filterwarnings("ignore", category=ShapelyDeprecationWarning) def main(): # 1. 加载OSM建筑数据 chicago_osm = OSM(get_data("chicago")) buildings = chicago_osm.get_buildings() # 计算建筑中心点(用于近邻搜索) buildings["centroid"] = buildings.geometry.centroid # 转换为弧度格式的[纬度, 经度]数组(适配haversine距离计算) building_coords = np.radians(buildings["centroid"].apply(lambda geom: (geom.y, geom.x)).to_numpy()) # 2. 加载CSV点数据(替换为你的文件路径) gig_data = pd.read_csv("data_sample.csv", encoding="latin-1") # --- 解析经纬度(根据你的CSV格式调整)--- # 如果你的经纬度是类似"POINT(lon lat)"的字符串,用下面的代码解析: # coord_extract = gig_data["your_coord_column"].str.extract(r'POINT\((\d+\.\d+) (\d+\.\d+)\)') # gig_data["lon"] = coord_extract[0].astype(float) # gig_data["lat"] = coord_extract[1].astype(float) # 转换为弧度格式的[纬度, 经度]数组 point_coords = np.radians(gig_data[["lat", "lon"]].to_numpy()) # 3. 构建BallTree执行近邻搜索 tree = BallTree(building_coords, metric="haversine") # 查找每个点的最近1个建筑,返回弧度距离和建筑索引 distances_rad, indices = tree.query(point_coords, k=1) # 4. 转换距离为米并合并建筑属性 gig_data["nearest_building_distance_m"] = distances_rad.flatten() * 6371000 # 地球半径约6371km nearest_buildings = buildings.iloc[indices.flatten()] # 合并建筑核心属性到原数据 gig_data = pd.concat([ gig_data.reset_index(drop=True), nearest_buildings.reset_index(drop=True)[["id", "name", "building"]] ], axis=1) # 5. 保存结果或继续分析 gig_data.to_csv("gig_data_with_nearest_building.csv", index=False, encoding="latin-1") print(gig_data.head()) if __name__ == '__main__': main()
关键说明
- 数据预处理:将建筑几何图形转换为中心点,简化近邻计算;CSV中的经纬度需解析为数值型列,若为
POINT(lon lat)格式,使用代码中的正则解析逻辑即可。 - 高效近邻搜索:
BallTree搭配haversinemetric专门适配经纬度数据,计算球面距离更准确,效率远高于逐点遍历。 - 结果合并:直接将匹配到的建筑ID、名称、类型等属性合并到原CSV,同时添加距离列辅助分析。
原代码修正点
- 移除无关的路网加载逻辑,聚焦建筑数据处理;
- 替换生成器为Pandas加载CSV,简化数据处理与合并流程;
- 用BallTree替代低效字典检索,实现批量近邻匹配;
- 直接关联建筑属性而非路网节点,解决核心需求。
内容的提问来源于stack exchange,提问作者Tor-iv
相关产品推荐
相关产品推荐

