如何修改代码基于城市数据集为每个城市生成距离列?
批量计算数据集坐标到多个城市的距离列
问题背景
我之前使用haversine_distance函数计算数据集中的坐标与单个指定点(start_lat, start_lon = 40.6976637, -74.1197643)的距离,代码已成功生成Distance列。现在需要修改代码,改用包含多个城市坐标的数据集,为每个城市生成以城市名为列名的距离列。
原始城市坐标数组如下:
[['Nanaimo' -123.9364 49.1642] ['Prince Rupert' -130.3271 54.3122] ['Vancouver' -123.1386 49.2636] ['Victoria' -123.3673 48.4275] ['Edmonton' -113.4909 53.5445] ['Winnipeg' -97.1392 49.8994] ['Sarnia' -82.4065 42.9746] ['Sarnia' -82.4065 42.9746] ['North York' -79.4112 43.7598] ['Kingston' -76.4812 44.2305] ['St. Catharines' -79.2333 43.1833] ['Thunder Bay' -89.2461 48.3822] ['Gaspé' -64.4833 48.8333] ['Cap-aux-Meules' -61.8607 47.3801] ['Kangiqsujuaq' -71.9667 61.6] ['Montreal' -73.5534 45.5091] ['Quebec City' -71.2074 46.8142] ['Rimouski' -68.524 48.4489] ['Sept-Îles' -66.3833 50.2167] ['Bathurst' -65.6497 47.6186] ['Charlottetown' -63.1399 46.24] ['Corner Brook' -57.9711 48.9411] ['Dartmouth' -63.5714 44.6715] ['Lewisporte' -55.0667 49.2333] ['Port Hawkesbury' -61.3642 45.6153] ['Saint John' -66.0628 45.2796] ["St. John's" -52.7072 47.5675] ['Sydney' -60.1947 46.1381] ['Yarmouth' -66.1175 43.8361]]
解决方案
以下是修改后的代码,已优化效率并实现需求:
import numpy as np import pandas as pd # 保留原有的Haversine距离计算函数 def haversine_distance(lat1, lon1, lat2, lon2): r = 6371 # 地球半径(公里) phi1 = np.radians(lat1) phi2 = np.radians(lat2) delta_phi = np.radians(lat2 - lat1) delta_lambda = np.radians(lon2 - lon1) a = np.sin(delta_phi / 2)**2 + np.cos(phi1) * np.cos(phi2) * np.sin(delta_lambda / 2)**2 res = r * (2 * np.arctan2(np.sqrt(a), np.sqrt(1 - a))) return np.round(res, 2) # 整理城市数据:转成元组列表并去重重复的Sarnia条目 cities = [ ('Nanaimo', -123.9364, 49.1642), ('Prince Rupert', -130.3271, 54.3122), ('Vancouver', -123.1386, 49.2636), ('Victoria', -123.3673, 48.4275), ('Edmonton', -113.4909, 53.5445), ('Winnipeg', -97.1392, 49.8994), ('Sarnia', -82.4065, 42.9746), ('North York', -79.4112, 43.7598), ('Kingston', -76.4812, 44.2305), ('St. Catharines', -79.2333, 43.1833), ('Thunder Bay', -89.2461, 48.3822), ('Gaspé', -64.4833, 48.8333), ('Cap-aux-Meules', -61.8607, 47.3801), ('Kangiqsujuaq', -71.9667, 61.6), ('Montreal', -73.5534, 45.5091), ('Quebec City', -71.2074, 46.8142), ('Rimouski', -68.524, 48.4489), ('Sept-Îles', -66.3833, 50.2167), ('Bathurst', -65.6497, 47.6186), ('Charlottetown', -63.1399, 46.24), ('Corner Brook', -57.9711, 48.9411), ('Dartmouth', -63.5714, 44.6715), ('Lewisporte', -55.0667, 49.2333), ('Port Hawkesbury', -61.3642, 45.6153), ('Saint John', -66.0628, 45.2796), ("St. John's", -52.7072, 47.5675), ('Sydney', -60.1947, 46.1381), ('Yarmouth', -66.1175, 43.8361) ] # 遍历每个城市,生成对应距离列 for city_name, city_lon, city_lat in cities: # 使用pandas向量化计算,比逐行循环效率更高 pandas_df[city_name] = haversine_distance( pandas_df['lat'], pandas_df['lon'], city_lat, city_lon ) # 查看处理后的数据集 print(pandas_df)
关键说明
- 对原始城市数组做了整理:转成元组列表并移除重复的Sarnia条目,避免生成重复列
- 改用pandas向量化操作替代
itertuples逐行循环,大幅提升计算效率(尤其适合大数据集) - 直接以城市名称作为新列的列名,完全匹配需求
内容的提问来源于stack exchange,提问作者avxeesh
相关产品推荐
相关产品推荐

