遍历字典列表与DataFrame列映射:行匹配错误修复求助
问题解决:修正经纬度与提取信息的映射错误
错误原因分析
你的自定义函数extract_details_sb_dsp_positive核心问题在于:函数内部循环遍历了整个data_metering DataFrame,导致每个response单元格提取出的storeExternalId和provider会与所有行的经纬度配对,而非仅对应行的经纬度,完全破坏了行与行之间的映射关系。
修正方案
我们需要让处理逻辑绑定到当前行:使用apply(axis=1)逐行处理,同时获取当前行的经纬度和对应的response数据,确保提取的信息只和当前行的地理坐标关联。
修正后的代码
def extract_details_sb_dsp_positive(row): res = [] response = row['RESPONSE'] # 校验必要键是否存在 if 'store-boundary-dsp' not in response: return res if 'estimates' not in response['store-boundary-dsp']: return res # 获取当前行的经纬度 lat = row['LATITUDE'] lng = row['LONGITUDE'] # 遍历当前response中的estimates for estimate in response['store-boundary-dsp']['estimates']: store_id = estimate.get('storeExternalId', '') provider = estimate.get('provider', '') # 仅绑定当前行的经纬度 res.append([store_id, provider, lat, lng]) return res # 逐行处理DataFrame response_values = data_metering.apply(extract_details_sb_dsp_positive, axis=1).tolist() # 扁平化结果列表 pair_values = [val for sublist in response_values for val in sublist] # 生成新的DataFrame import pandas as pd new_df = pd.DataFrame(pair_values, columns=['storeExternalId', 'provider', 'latitude', 'longitude'])
代码说明
- 函数接收整行数据作为参数,直接从行中获取当前的
response和经纬度,避免全局遍历DataFrame。 - 使用
dict.get()方法更简洁地提取目标字段,同时设置默认值避免键缺失报错。 - 提取的每组
storeExternalId/provider仅与当前行的经纬度组合,保证映射关系正确。 - 最后将扁平化后的结果转为新DataFrame,完全符合需求格式。
内容的提问来源于stack exchange,提问作者lifo
相关产品推荐
相关产品推荐

