为什么PySpark中使用isin函数实现列值映射的代码无法正常生效
PySpark isin映射不生效问题原因分析
核心原因
- 你在
withColumn操作中把映射后的结果存到了新增的handset_type字段,原handset_os字段没有被修改,后续groupBy时依然用了未修改的原字段统计,自然看不到映射效果。 when判断没有补充otherwise分支,即使你后续统计handset_type字段,不在匹配列表内的取值会默认返回null,无法保留原系统值。
修正代码
listOTHERS_hpos = ['BLACKBERRY 7', 'SYMBIAN', 'NOKIA OS', 'BLACKBERRY 10', 'WINDOWS'] lastvalue_month = lastvalue_month.withColumn('handset_os', when(col('handset_os').isin(listOTHERS_hpos) , lit('OTHERS')) # 非匹配项保留原handset_os取值 .otherwise(col('handset_os'))) lastvalue_month.groupBy('handset_os').count().orderBy('count').show()
如果需要保留原handset_os字段,只需要把withColumn的第一个参数改为handset_type,后续groupBy也改为handset_type即可。
内容的提问来源于stack exchange,提问作者Nabih Bawazir
相关产品推荐
相关产品推荐

