Databricks中PySpark DataFrame转pandas报错:'DataFrame'无dtype属性
PySpark DataFrame转pandas时触发AttributeError: 'DataFrame' object has no attribute 'dtype'
问题描述
在Databricks环境中,调用df.toPandas()将PySpark DataFrame转为pandas DataFrame时,持续报错:
/databricks/spark/python/pyspark/sql/pandas/conversion.py:145: UserWarning: toPandas attempted Arrow optimization because 'spark.sql.execution.arrow.pyspark.enabled' is set to true, but has reached the error below and can not continue. Note that 'spark.sql.execution.arrow.pyspark.fallback.enabled' does not have an effect on failures in the middle of computation. 'DataFrame' object has no attribute 'dtype' warnings.warn(msg) AttributeError: 'DataFrame' object has no attribute 'dtype'
已尝试禁用Arrow优化:
spark.conf.set("spark.sql.execution.arrow.enabled", "false")
但问题仍未解决,参考相关帖子也无帮助。
更新:执行df.printSchema()得到以下结构:
flight_id: string (nullable = true) |-- flight_direction: string (nullable = true) |-- service_type: string (nullable = true) |-- flight_designator: string (nullable = true) |-- flight_number: string (nullable = true) |-- callsign: string (nullable = true) |-- scheduled_datetime: timestamp (nullable = true) |-- connecting_flight_designator: string (nullable = true) |-- airport_iata_codes: array (nullable = true) | |-- element: string (containsNull = true) |-- airline_name: string (nullable = true) |-- airport_names: array (nullable = true) | |-- element: string (containsNull = true) |-- country_number: long (nullable = true) |-- eu_category: string (nullable = true) |-- safe_town_indicator: boolean (nullable = true) |-- sibt: timestamp (nullable = true) |-- aibt: timestamp (nullable = true) |-- sobt: timestamp (nullable = true) |-- aibt: timestamp (nullable = true) |-- tsat: timestamp (nullable = true) |-- aircraft_name: string (nullable = true) |-- aircraft_registration: string (nullable = true) |-- ramp: string (nullable = true) |-- ramp_previous: string (nullable = true) |-- seats: long (nullable = true) |-- actual_total_pax: integer (nullable = true) |-- handler_apron: string (nullable = true) |-- occupancy_rate: double (nullable = false)
解决方案
从报错信息和Schema可以看出,问题根源是DataFrame存在重复列名——你的Schema里出现了两次aibt字段。PySpark允许重复列,但pandas不支持,不管是否开启Arrow优化,转换时都会因为重复列触发异常。
步骤1:处理重复列
方法一:重命名重复列
先找出所有重复列,再给重复项添加后缀区分:
from collections import Counter # 检查重复列 col_counts = Counter(df.columns) duplicate_cols = [col for col, cnt in col_counts.items() if cnt > 1] print("重复列:", duplicate_cols) # 生成新列名,给重复列加序号后缀 new_col_names = [] col_counter = {} for col in df.columns: if col in col_counter: col_counter[col] += 1 new_col_names.append(f"{col}_{col_counter[col]}") else: col_counter[col] = 1 new_col_names.append(col) # 重命名DataFrame的列 df = df.toDF(*new_col_names)
方法二:删除重复列(如果不需要其中一个)
如果确定其中一个重复列无用,可以直接删除:
# 找到第二个aibt列的索引并删除 target_col = 'aibt' # 从索引1开始查找第二个出现的aibt col_index = df.columns.index(target_col, 1) df = df.drop(df.columns[col_index])
步骤2:重新尝试转换
处理完重复列后,执行df.toPandas()即可正常转换为pandas DataFrame。
内容的提问来源于stack exchange,提问作者Hans.nl
相关产品推荐
相关产品推荐

