双CSV坐标文件邻近性检测:内层循环仅执行一次问题及优化咨询
问题分析与解决方案
内层循环仅执行一次的原因及修复
原因
csv.reader返回的是迭代器,迭代器只能被完整遍历一次。第一次外层循环时,内层循环已经耗尽了file2的所有元素;后续外层循环再进入内层循环时,file2迭代器没有剩余元素,循环直接跳过。
修复代码
把file2的内容提前读取到列表中,同时用with语句管理文件句柄避免资源泄漏:
import csv def comparison_closeness_two_files(filename1, filename2): threshold_meters = 10/1000 # 10米转千米 # 提前读取并过滤file2的有效行,避免迭代器耗尽 with open(filename2, 'r') as f2: file2 = list(csv.reader(f2, delimiter=',')) file2 = [row for row in file2 if len(row) == 4] with open(filename1, 'r') as f1: file1 = csv.reader(f1, delimiter=',') for fc1 in file1: if len(fc1) != 4: continue lat1 = float(fc1[2]) lon1 = float(fc1[3]) for fc2 in file2: lat2 = float(fc2[2]) lon2 = float(fc2[3]) distance = get_distance_between(lat1, lon1, lat2, lon2) if distance <= threshold_meters: return True return False
更高效的实现方式
暴力两两比对的时间复杂度为O(n*m),数据量较大时效率极低。推荐用KD-Tree空间索引优化,将时间复杂度降至O(n log m):
优化代码(依赖scipy和numpy)
import csv import numpy as np from scipy.spatial import KDTree from math import radians, cos, sin, asin, sqrt # Haversine公式计算球面距离(单位:千米) def get_distance_between(lat1, lon1, lat2, lon2): lon1, lat1, lon2, lat2 = map(radians, [lon1, lat1, lon2, lat2]) dlon = lon2 - lon1 dlat = lat2 - lat1 a = sin(dlat/2)**2 + cos(lat1) * cos(lat2) * sin(dlon/2)**2 c = 2 * asin(sqrt(a)) return c * 6371 # 地球平均半径6371千米 # 加载文件中有效坐标(过滤无效行和非法数值) def load_valid_coords(filename): coords = [] with open(filename, 'r') as f: reader = csv.reader(f, delimiter=',') for row in reader: if len(row) == 4: try: lat = float(row[2]) lon = float(row[3]) coords.append([lon, lat]) # KD-Tree采用(x,y)格式,经度为x,纬度为y except ValueError: continue return np.array(coords) def comparison_closeness_two_files(filename1, filename2): threshold_km = 10 / 1000 # 10米转千米 coords1 = load_valid_coords(filename1) coords2 = load_valid_coords(filename2) if len(coords1) == 0 or len(coords2) == 0: return False # 用其中一组坐标构建KD-Tree tree = KDTree(coords2) # 对每个点查询阈值内的近邻,存在则返回True for point in coords1: if len(tree.query_ball_point(point, threshold_km)) > 0: return True return False
无第三方库的轻量优化
如果不能依赖第三方库,可以先按纬度过滤:10米对应的纬度差约为0.00009度,遍历点时先跳过纬度差超过该值的点,再计算精确距离,减少无效的距离计算量。
内容的提问来源于stack exchange,提问作者kregtopoli
相关产品推荐
相关产品推荐

