Pandas调用英国邮编API逐行新增坐标报JSONDecodeError如何解决?
报错原因排查思路
- 参数异常:单条测试使用的是标准有效邮编,但数据集中
property_postcode列可能存在空值、格式错乱、无效邮编的情况,传入API后会返回404/500等错误响应,响应内容为HTML错误页面而非JSON格式,直接调用.json()方法就会抛出你遇到的错误。 - 代码隐患:你代码中
get_data函数使用r.get(url)发起请求,但你导入库时写的是import requests没有给别名r,若运行环境没有提前导入import requests as r会引发请求失败,返回非JSON内容。 - 限流问题:该公开API有请求频率限制,短时间批量发起请求会被限流,返回非JSON的拦截页面,触发解析错误。
- 缺少状态校验:你没有判断请求的HTTP状态码,只要请求失败就直接解析JSON,必然会触发报错。
解决方案
首先优化get_data函数,增加参数校验、状态判断、异常捕获逻辑,避免单个请求失败中断整个执行流程,同时可以按需打印错误日志定位异常邮编:
import requests import pandas as pd geo = spark.table('property_address').toPandas() def get_data(postcode): # 空值直接返回 if pd.isna(postcode): return (None, None, None, None) # 清洗邮编格式 postcode_clean = str(postcode).strip() if not postcode_clean: return (None, None, None, None) url = f"http://api.getthedata.com/postcode/{postcode_clean}" try: # 增加超时配置,修正请求写法 req = requests.get(url, timeout=10) # 校验请求状态,非200直接抛出异常走捕获逻辑 req.raise_for_status() res = req.json() # 校验返回结果结构合法 if res.get("status") == "success" and "data" in res: data = res["data"] return ( data.get("easting"), data.get("northing"), data.get("latitude"), data.get("longitude") ) return (None, None, None, None) except Exception as e: # 如需排查异常可打开下方注释,打印出问题的邮编和错误信息 # print(f"邮编{postcode_clean}处理失败,错误:{str(e)}") return (None, None, None, None)
然后优化批量调用逻辑,比原有的逐行修改row的写法性能更高:
geo['east'], geo['north'], geo['lat'], geo['long'] = zip(*geo['property_postcode'].apply(get_data)) display(geo)
可选优化建议
- 数据量超过1000条时,不要用普通apply串行请求,可以用
concurrent.futures.ThreadPoolExecutor多线程发起请求,速度提升5-10倍。 - 新增缓存逻辑,相同邮编只调用一次API,减少重复请求,避免触发限流。
- 原代码中的
call_api函数没有被使用,可以直接删除。
内容的提问来源于stack exchange,提问作者TJB
相关产品推荐
相关产品推荐

