在pandas DataFrame中如何逐行逐列写入?地理编码遍历代码问题排查
现有代码错误点
- 第一版代码问题:
df_fin.loc['address']是对行索引为address的整行赋值,而非对当前遍历行的address列赋值,每次循环都会覆盖全局值,最终所有行的新增字段都会被最后一次循环的结果覆盖,无匹配行索引时还会新增无效行。- 未记录当前遍历的行索引,无法定位到需要写入的具体行。
- 第二版代码问题:
iterrows()是DataFrame的专属方法,单独取market_address列得到的是Series对象,无该方法,运行会直接报错。iterrows()返回的row是原数据的副本,直接修改row不会回写到原DataFrame。- 直接将
row传入geolocator.geocode()会导致格式错误,需要取market_address列的字符串值传入。
正确实现方案
from geopy.geocoders import Nominatim import time # 数据量大时建议加延时避免接口限流 geolocator = Nominatim(user_agent="http") # 预初始化新增列,避免后续赋值出现格式错误 df_fin['address'] = '' df_fin['latitude'] = 0.0 df_fin['longitude'] = 0.0 df_fin['raw'] = '' # 遍历DataFrame全量行,拿到行索引和行数据 for index, row in df_fin.iterrows(): current_address = row['market_address'] try: location = geolocator.geocode(current_address) # 通过行索引定位原表的对应位置,直接修改原表数据 df_fin.loc[index, 'address'] = location.address df_fin.loc[index, 'latitude'] = location.latitude df_fin.loc[index, 'longitude'] = location.longitude df_fin.loc[index, 'raw'] = str(location.raw) print(location.raw) except: df_fin.loc[index, 'raw'] = f'no info for: {current_address}' print(f'no info for: {current_address}') time.sleep(1) # 按需调整延时时长,避免请求频率过高被封 df_fin.tail(10)
核心逻辑说明
通过iterrows()获取每行的原生索引,再用df.loc[行索引, 列名]直接操作原DataFrame的对应单元格,避免副本修改不生效的问题。
内容的提问来源于stack exchange,提问作者ASH
相关产品推荐
相关产品推荐

