Python嵌套循环处理坐标匹配时KeyError问题求助
问题排查:DataFrame嵌套循环迭代时触发KeyError
问题描述
尝试用嵌套for循环比较两个DataFrame的坐标集:当两点距离小于0.2时,覆盖qinsy_file_2中的CMP坐标;若所有SEGY点都超出该距离则删除对应行。但脚本仅完成一次迭代,第二次计算距离时触发KeyError,相关代码及报错如下:
原代码
## Pull values from GUI qinsy_file=pd.read_csv(values["-QINSYInput-"] ,sep=',') segy_file=pd.read_csv(values["-SEGYInput-"] ,sep='\t') #print(segy_file) in_file=str(values["-QINSYInput-"]) ## Make the outfile name by replacing file suffix out_file=in_file.replace(".csv","_SEGY_NAV.csv").replace(".txt","_SEGY_NAV.txt") ## Correlation zone = 30cm buffer=0.2 ## Get required headers segy_vlookup=segy_file[['CDP_X','CDP_Y']] qinsy_file_2=qinsy_file[['Date','Time','Sparker CoG Easting','Sparker CoG Northing', 'Streamer CoG Easting','Streamer CoG Northing','CMP Easting', 'CMP Northing','Fix Number','CMP DTM Depth']] ## Loop through Qinsy file for index_qinsy,row_qinsy in qinsy_file_2.iterrows(): ## Loop through SEGY navigation for index_segy,row_segy in segy_vlookup.iterrows(): ## Calculate distance between points distance = (((segy_vlookup["CDP_X"][index_segy] - qinsy_file_2["CMP Easting"][index_qinsy])**2) + ((segy_vlookup["CDP_Y"][index_segy] - qinsy_file_2["CMP Northing"][index_qinsy])**2))**0.5 print(distance) ## If distance between points is less than or equal to the correlation value, replace the CMP X and Y values in the QINSY file if distance <= buffer: qinsy_file_2["CMP Easting"][index_qinsy]=segy_vlookup["CDP_X"][index_segy] qinsy_file_2["CMP Northing"][index_qinsy]=segy_vlookup["CDP_Y"][index_segy] print(qinsy_file_2) #qinsy_file_2["CMP Easting"]=segy_vlookup["CDP_X"] #qinsy_file_2["CMP Northing"]=segy_vlookup["CDP_Y"] else: ## Need to delete the row at this point qinsy_file_2.drop(index_qinsy,inplace=True) ## Export the "filtered" dataframe to csv, turning off index qinsy_file_2.to_csv(out_file,sep=',',index=False,header=True)
报错信息
71.10718458835196 # this is distance Traceback (most recent call last): File "C:\Users\tholgate\AppData\Local\Programs\Python\Python39\lib\site-packages\pandas\core\indexes\range.py", line 414, in get_loc return self._range.index(new_key) ValueError: 0 is not in range The above exception was the direct cause of the following exception: Traceback (most recent call last): File "p:\Xtra\Public\TH\Python Code\SEIS_NAV Comparison.py", line 52, in <module> distance = (((segy_vlookup["CDP_X"][index_segy] - qinsy_file_2["CMP Easting"][index_qinsy])**2) + ((segy_vlookup["CDP_Y"][index_segy] - qinsy_file_2["CMP Northing"][index_qinsy])**2))**0.5 File "C:\Users\tholgate\AppData\Local\Programs\Python\Python39\lib\site-packages\pandas\core\series.py", line 1040, in __getitem__ return self._get_value(key) File "C:\Users\tholgate\AppData\Local\Programs\Python\Python39\lib\site-packages\pandas\core\series.py", line 1156, in _get_value loc = self.index.get_loc(label) File "C:\Users\tholgate\AppData\Local\Programs\Python\Python39\lib\site-packages\pandas\core\indexes\range.py", line 416, in get_loc raise KeyError(key) from err KeyError: 0
错误原因分析
- 迭代时修改原DataFrame导致索引失效:在
iterrows()遍历qinsy_file_2的过程中,直接执行drop(index_qinsy, inplace=True)会实时删除行,改变原DataFrame的索引结构。当后续循环尝试访问已被删除的索引(比如0)时,就会触发KeyError。 - 匹配逻辑错误:原代码中只要某一个SEGY点与当前QINSY点的距离大于buffer,就立即删除该行,这不符合“所有SEGY点都超出距离才删除”的需求,而且会导致第一次不满足条件就删行,后续循环找不到原索引。
- 链式赋值风险:
qinsy_file_2["CMP Easting"][index_qinsy]属于链式赋值,可能触发SettingWithCopyWarning,且不是修改DataFrame的安全方式。
修正方案
核心思路
- 遍历过程中不直接修改原DataFrame,而是先标记需要保留/修改的行。
- 对每个QINSY点,检查是否存在至少一个SEGY点满足距离条件,有则替换坐标,无则标记删除。
- 使用
.loc进行安全赋值,避免链式赋值问题。
修正后的代码
## Pull values from GUI qinsy_file = pd.read_csv(values["-QINSYInput-"], sep=',') segy_file = pd.read_csv(values["-SEGYInput-"], sep='\t') in_file = str(values["-QINSYInput-"]) ## Make the outfile name by replacing file suffix out_file = in_file.replace(".csv", "_SEGY_NAV.csv").replace(".txt", "_SEGY_NAV.txt") ## Correlation zone = 30cm buffer = 0.2 ## Get required headers segy_vlookup = segy_file[['CDP_X', 'CDP_Y']].copy() # 创建副本避免修改原数据,同时确保是独立DataFrame qinsy_file_2 = qinsy_file[['Date', 'Time', 'Sparker CoG Easting', 'Sparker CoG Northing', 'Streamer CoG Easting', 'Streamer CoG Northing', 'CMP Easting', 'CMP Northing', 'Fix Number', 'CMP DTM Depth']].copy() # 标记需要删除的行 to_drop = [] ## Loop through Qinsy file for index_qinsy, row_qinsy in qinsy_file_2.iterrows(): found_match = False qinsy_x = row_qinsy['CMP Easting'] qinsy_y = row_qinsy['CMP Northing'] ## Loop through SEGY navigation for index_segy, row_segy in segy_vlookup.iterrows(): segy_x = row_segy['CDP_X'] segy_y = row_segy['CDP_Y'] ## Calculate distance between points distance = ((segy_x - qinsy_x)**2 + (segy_y - qinsy_y)**2)**0.5 print(distance) if distance <= buffer: # 使用.loc安全赋值 qinsy_file_2.loc[index_qinsy, 'CMP Easting'] = segy_x qinsy_file_2.loc[index_qinsy, 'CMP Northing'] = segy_y found_match = True # 找到匹配后可跳出内层循环,提升效率 break # 若没有找到任何匹配的SEGY点,标记为待删除 if not found_match: to_drop.append(index_qinsy) # 批量删除标记的行 qinsy_file_2.drop(to_drop, inplace=True) ## Export the "filtered" dataframe to csv, turning off index qinsy_file_2.to_csv(out_file, sep=',', index=False, header=True)
关键优化点
- 创建DataFrame副本进行操作,避免影响原数据。
- 增加
found_match标记,确保只有当所有SEGY点都不满足距离条件时才删除该行。 - 使用
.loc进行赋值,避免链式赋值的风险。 - 找到匹配后跳出内层循环,提升运行效率。
内容的提问来源于stack exchange,提问作者tholgate
相关产品推荐
相关产品推荐

