You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

批量更新经纬度遭请求限制,求高效解决方案(54000行数据)

解决大规模CEP转经纬度的请求限制问题

原代码的核心问题

  • 每次循环重复创建Nominatim实例,增加服务器识别风险
  • 未对已查询过的CEP/地址结果做缓存,重复请求浪费资源
  • 固定sleep时间不灵活,无法适配动态的限流规则
  • 未处理请求失败场景(如地址解析失败、API限流触发),容易导致程序崩溃

优化方案

1. 复用Geocoder实例

将Nominatim实例初始化移到循环外,避免重复建立连接,降低服务器端的识别压力。

2. 实现缓存机制

用字典缓存已查询过的CEP对应的经纬度,54000行数据中大概率存在大量重复CEP,缓存能直接减少请求量。

3. 动态退避处理限流

当触发Blocked by Flood时,采用指数退避策略(如第一次等30s,第二次60s,第三次120s),而不是固定sleep;同时捕获异常,避免程序中断。

4. 预去重(关键提速手段)

先对CSV中的CEP列去重,只查询唯一的CEP,再将结果映射回原表格,能大幅减少请求次数(若重复率50%,请求量直接减半)。

优化后的代码示例

import time as tm
import pandas as pa
import brazilcep as bc
from geopy.geocoders import Nominatim
from geopy.exc import GeocoderQuotaExceeded, GeocoderTimedOut

# 缓存已查询的CEP结果,key=CEP,value=(lat, lon)
cep_cache = {}
# 复用Nominatim实例,使用辨识度更高的user-agent
geolocator = Nominatim(user_agent="joao_brasil_geocoder")

coord = pa.read_csv("Arquivos\\Coordenadas.csv", sep=";")

# 先处理唯一CEP,减少重复请求
unique_ceps = coord["CEP"].unique()
for cep in unique_ceps:
    if cep in cep_cache:
        continue
    
    retry_count = 0
    max_retries = 5
    base_sleep = 15
    success = False
    
    while retry_count < max_retries and not success:
        try:
            endereco = bc.get_address_from_cep(cep)
            ad = f"{endereco['street'].split('-')[0]}, {endereco['city']}"
            location = geolocator.geocode(ad, timeout=10)
            
            if location:
                cep_cache[cep] = (location.latitude, location.longitude)
                success = True
                tm.sleep(base_sleep)
            else:
                print(f"CEP {cep} 无法解析地址")
                success = True
                tm.sleep(base_sleep)
        
        except GeocoderQuotaExceeded:
            sleep_time = base_sleep * (2 ** retry_count)
            print(f"触发限流,等待{sleep_time}秒后重试")
            tm.sleep(sleep_time)
            retry_count += 1
        
        except GeocoderTimedOut:
            sleep_time = base_sleep * (1.5 ** retry_count)
            print(f"请求超时,重试第{retry_count+1}次")
            tm.sleep(sleep_time)
            retry_count += 1
        
        except Exception as e:
            print(f"处理CEP {cep} 出错: {str(e)}")
            success = True
            tm.sleep(base_sleep)

# 将缓存结果映射回原表格
coord["LATITUDE"] = coord["CEP"].map(lambda x: cep_cache[x][0] if x in cep_cache else None)
coord["LONGITUDE"] = coord["CEP"].map(lambda x: cep_cache[x][1] if x in cep_cache else None)

# 保存结果
coord.to_csv("Arquivos\\Coordenadas_Atualizadas.csv", index=False)

额外建议

  • 更换User-Agent:避免使用test_app这类通用名称,用更具体的标识(如你的名字+用途),降低被判定为爬虫的概率
  • 分批处理:把唯一CEP分成若干批次(如每500个一批),每批处理完后休息5-10分钟,避免持续请求触发限流
  • 考虑付费API:如果免费API限制过严,可选用巴西本地的付费地理编码服务(如Google Maps Geocoding、Here Maps),这类服务限流宽松、响应速度更快

内容的提问来源于stack exchange,提问作者João Brugnolo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 08:30:25