PySpark报错cannot resolve '`timestamp`'问题排查求助
报错核心原因
- 调用
F.to_timestamp()时第一个参数传入了不存在的列名timestamp,你当前使用的DataFrame仅包含checkin_date列存储时间字符串,Spark无法找到指定列抛出解析异常。 date_trunc()参数顺序传反,该函数标准语法为date_trunc(截断单位, 时间列),你错误将列名放在了第一个参数位。
修复后的完整实现代码
优先使用Spark原生函数替代自定义UDF,执行性能更高,同时适配Yelp签到数据集的格式:
from pyspark.sql import functions as F # 读取Yelp签到数据集 checkin = spark.read.json('yelp_academic_dataset_checkin.json.gz') # 拆分多值签到时间、行转列展开 dates = checkin.select(F.split('date', ',').alias('date_list'))\ .withColumn('checkin_date_str', F.explode('date_list')) # 转换时间格式、按小时统计签到次数 hour_count = dates.withColumn('checkin_time', F.to_timestamp(F.trim('checkin_date_str'), 'yyyy-MM-dd HH:mm:ss \'UTC\''))\ # 如果只需要统计0-23的小时维度,替换下一行代码为 .groupBy(F.hour('checkin_time').alias('hour')) .groupBy(F.date_trunc('hour', 'checkin_time').alias('checkin_hour'))\ .count()\ .orderBy(F.desc('count')) # 输出结果,第一行即为签到次数最多的小时 hour_count.show(truncate=False)
内容的提问来源于stack exchange,提问作者Hefe
相关产品推荐
相关产品推荐

