You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark中调用geolocator.reverse()触发TypeError的问题求助

解决PySpark中Geopy调用的两个常见错误

第一个错误:TypeError: reverse() takes 2 positional arguments but 3 were given

问题根源

你踩了两个容易忽略的小坑:

  1. Geopy Reverse方法参数格式错误:Nominatim.reverse()仅接受一个location参数(加上实例自身的self),这个参数必须是纬度和经度组成的元组(比如(40.7128, -74.0060))。你直接传入了两个单独的lat和lon参数,导致方法接收到3个参数(self, lat, lon),触发了参数数量不匹配的错误。
  2. 直接用普通函数处理Spark Column:你写了direccion_func(F.col("Latitud"), F.col("Longitud")),但direccion_func是普通Python函数,根本无法识别Spark的Column对象——必须使用你注册好的UDF来调用才行。

修复后的代码

# 确保依赖安装
!pip install geopy
from geopy.geocoders import Nominatim
geolocator = Nominatim(user_agent='your_app_name')  # 建议替换为你的实际应用名称
from pyspark.sql.functions import col, udf, F
from pyspark.sql.types import StringType

def direccion_func(lat, lon):
    # 将经纬度打包成元组传给reverse方法
    try:
        return geolocator.reverse((lat, lon)).address
    except Exception as e:
        print(f"获取地址失败: {e}")
        return None  # 捕获异常,避免单个错误导致整个任务崩溃

# 注册UDF
direccion_udf = udf(direccion_func, StringType())
# 使用注册好的UDF处理Spark Column
paradasRuta1DF = paradasRuta1DF.withColumn('Direccion', direccion_udf(F.col("Latitud"), F.col("Longitud")))

第二个错误:TypeError: Failed to create Point instance from Column<b'LatLong'>

问题根源

  1. 混淆普通函数与Pandas UDF的调用方式:你直接把Spark Column传给了direccion_func,但这个函数是普通Python函数,无法处理Spark的Column类型——必须通过你注册的Pandas UDF来调用。
  2. 坐标列格式不兼容:如果LatLong是数组或字符串格式,reverse方法无法直接识别,需要先转换成它能接受的坐标元组格式。

修复后的代码(假设LatLong是数组类型,格式为[纬度, 经度])

from geopy.geocoders import Nominatim
geolocator = Nominatim(user_agent='your_app_name')
from pyspark.sql.functions import pandas_udf, F
from pyspark.sql.types import StringType
import pandas as pd

def direccion_func(coordenadas_series):
    # 定义单个坐标的处理逻辑
    def get_single_address(coords):
        if not coords or len(coords) != 2:
            return None
        try:
            return geolocator.reverse((coords[0], coords[1])).address
        except Exception as e:
            print(f"处理坐标{coords}失败: {e}")
            return None
    # 对整个Pandas Series应用处理函数
    return coordenadas_series.apply(get_single_address)

# 注册Pandas UDF
direccion_pandas_udf = pandas_udf(direccion_func, returnType=StringType())
# 通过注册好的UDF调用处理列
paradasRuta1DF = paradasRuta1DF.withColumn('Direccion', direccion_pandas_udf(F.col("LatLong")))

额外注意事项

  • Nominatim调用限制:OpenStreetMap的Nominatim API有调用频率限制,短时间内大量请求会被临时封禁,建议在函数中加入time.sleep(1)(需导入time包),或者缓存已查询过的坐标结果。
  • 空值与异常处理:实际业务中一定要处理空坐标、无效坐标的情况,否则单个错误就会导致整个Spark任务失败,上面的代码已经添加了基础的异常捕获和空值判断逻辑。

内容的提问来源于stack exchange,提问作者Jesus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 09:58:12