You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark报错cannot resolve '`timestamp`'问题排查求助

报错核心原因
  • 调用F.to_timestamp()时第一个参数传入了不存在的列名timestamp,你当前使用的DataFrame仅包含checkin_date列存储时间字符串,Spark无法找到指定列抛出解析异常。
  • date_trunc()参数顺序传反,该函数标准语法为date_trunc(截断单位, 时间列),你错误将列名放在了第一个参数位。
修复后的完整实现代码

优先使用Spark原生函数替代自定义UDF,执行性能更高,同时适配Yelp签到数据集的格式:

from pyspark.sql import functions as F

# 读取Yelp签到数据集
checkin = spark.read.json('yelp_academic_dataset_checkin.json.gz')
# 拆分多值签到时间、行转列展开
dates = checkin.select(F.split('date', ',').alias('date_list'))\
               .withColumn('checkin_date_str', F.explode('date_list'))
# 转换时间格式、按小时统计签到次数
hour_count = dates.withColumn('checkin_time', 
                        F.to_timestamp(F.trim('checkin_date_str'), 'yyyy-MM-dd HH:mm:ss \'UTC\''))\
                  # 如果只需要统计0-23的小时维度,替换下一行代码为 .groupBy(F.hour('checkin_time').alias('hour'))
                  .groupBy(F.date_trunc('hour', 'checkin_time').alias('checkin_hour'))\
                  .count()\
                  .orderBy(F.desc('count'))
# 输出结果,第一行即为签到次数最多的小时
hour_count.show(truncate=False)

内容的提问来源于stack exchange,提问作者Hefe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 23:15:03