You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在pandas DataFrame中如何逐行逐列写入?地理编码遍历代码问题排查

现有代码错误点

  • 第一版代码问题:
    • df_fin.loc['address'] 是对行索引为address的整行赋值,而非对当前遍历行的address列赋值,每次循环都会覆盖全局值,最终所有行的新增字段都会被最后一次循环的结果覆盖,无匹配行索引时还会新增无效行。
    • 未记录当前遍历的行索引,无法定位到需要写入的具体行。
  • 第二版代码问题:
    • iterrows()是DataFrame的专属方法,单独取market_address列得到的是Series对象,无该方法,运行会直接报错。
    • iterrows()返回的row是原数据的副本,直接修改row不会回写到原DataFrame。
    • 直接将row传入geolocator.geocode()会导致格式错误,需要取market_address列的字符串值传入。

正确实现方案

from geopy.geocoders import Nominatim
import time # 数据量大时建议加延时避免接口限流
geolocator = Nominatim(user_agent="http")

# 预初始化新增列,避免后续赋值出现格式错误
df_fin['address'] = ''
df_fin['latitude'] = 0.0
df_fin['longitude'] = 0.0
df_fin['raw'] = ''

# 遍历DataFrame全量行,拿到行索引和行数据
for index, row in df_fin.iterrows():
    current_address = row['market_address']
    try:
        location = geolocator.geocode(current_address)
        # 通过行索引定位原表的对应位置,直接修改原表数据
        df_fin.loc[index, 'address'] = location.address
        df_fin.loc[index, 'latitude'] = location.latitude
        df_fin.loc[index, 'longitude'] = location.longitude
        df_fin.loc[index, 'raw'] = str(location.raw)
        print(location.raw)
    except:
        df_fin.loc[index, 'raw'] = f'no info for: {current_address}'
        print(f'no info for: {current_address}')
    time.sleep(1) # 按需调整延时时长,避免请求频率过高被封

df_fin.tail(10)

核心逻辑说明

通过iterrows()获取每行的原生索引,再用df.loc[行索引, 列名]直接操作原DataFrame的对应单元格,避免副本修改不生效的问题。

内容的提问来源于stack exchange,提问作者ASH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 00:30:00