You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark获取Course_Fees最大值对应记录时触发AttributeError问题求助

解决PySpark获取Course_Fees最大值对应记录的问题

错误原因

你遇到的AttributeError: 'DataFrame' object has no attribute '_get_object_id',是因为course_df.select(f.max("Course_Fees"))返回的是单行单列的DataFrame,而非可以直接与列比较的标量值。Spark无法将DataFrame当作表达式参与列的比较运算,因此触发错误。

正确实现方法

方法1:提取最大值标量(适合单条最大值记录场景)

先将最大值从DataFrame中提取为Python标量,再用于过滤:

from pyspark.sql import functions as f

# 获取Course_Fees的最大值标量
maximumCourseFee = course_df.select(f.max("Course_Fees")).first()[0]

# 过滤并展示结果
course_df.select("Name", "Duration_Months", "Payment_Done")
         .where(f.col("Course_Fees") == maximumCourseFee)
         .show()

方法2:窗口函数(适合多条最大值记录场景,更高效)

使用窗口函数可以避免将数据收集到Driver端,同时保留所有最大值对应的记录:

from pyspark.sql import Window
from pyspark.sql import functions as f

# 定义窗口:按Course_Fees降序排序
window_spec = Window.orderBy(f.col("Course_Fees").desc())

# 添加排名列并过滤排名第一的记录
course_df.select("Name", "Duration_Months", "Payment_Done", "Course_Fees")
         .withColumn("rank", f.rank().over(window_spec))  # 用rank()保留所有最大值记录,row_number()仅取一条
         .where(f.col("rank") == 1)
         .drop("rank")
         .show()

方法3:子查询(Spark SQL风格)

直接在过滤条件中嵌套最大值查询:

from pyspark.sql import functions as f

course_df.select("Name", "Duration_Months", "Payment_Done")
         .where(f.col("Course_Fees") == f.expr("(SELECT MAX(Course_Fees) FROM course_df)"))
         .show()

也可以用Spark SQL语句实现:

spark.sql("""
SELECT Name, Duration_Months, Payment_Done
FROM course_df
WHERE Course_Fees = (SELECT MAX(Course_Fees) FROM course_df)
""").show()

内容的提问来源于stack exchange,提问作者Peu Kgaphola

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 04:06:14