PySpark处理Watch Time字段时遇'function' object is not subscriptable错误求助
问题解决:移除Watch Time字段中的' min'并转换为整数
错误原因
你遇到的'function' object is not subscriptable错误,核心问题有两个:
- 错误调用
col函数:col是PySpark的函数,必须用括号()传入列名,而非下标[]语法; - 错误赋值
show()返回值:show()仅用于打印DataFrame,返回值是None,不能赋值给变量(比如df2)。
修正后的代码示例
方法1:用regexp_replace移除' min'并转整数
from pyspark.sql.types import * from pyspark.sql.functions import * # 修正列名引用,用括号调用col df1 = df.withColumn("Year Of Release", abs(col("Year of Release"))) # 移除" min"字符串,再转换为整数类型 df2 = df1.withColumn("Watch Time Cleaned", regexp_replace(col("Watch Time"), " min", "")) \ .withColumn("Watch Time (Minutes)", col("Watch Time Cleaned").cast(IntegerType())) # 处理完DataFrame后单独调用show() df2.show()
方法2:用split分割取数字部分并转整数
from pyspark.sql.types import * from pyspark.sql.functions import * df1 = df.withColumn("Year Of Release", abs(col("Year of Release"))) # 按空格分割字符串,取第一个元素后转整数 df2 = df1.withColumn("Watch Time (Minutes)", split(col("Watch Time"), " ").getItem(0).cast(IntegerType())) df2.show()
关键注意事项
- 所有列名引用必须用
col("列名")格式,禁止使用col["列名"]; - 避免将
show()的返回值赋值给变量,数据处理逻辑要在show()执行前完成; - 转换整数前确保字符串是纯数字,若存在非数字值,可使用
df2.na.drop(subset=["Watch Time (Minutes)"])过滤无效行。
内容的提问来源于stack exchange,提问作者Sanjay
相关产品推荐
相关产品推荐

