如何在PySpark DataFrame中映射列并避免生成空值?
解决PySpark列映射时保留未匹配值的方法
针对你遇到的问题,不用手动给所有现有值补充映射,有几种更高效的解决方式:
方法1:使用replace函数(最简洁)
PySpark的replace函数支持直接传入映射字典,仅替换匹配的键,未匹配的字段会自动保留原值,完全符合你的需求:
from pyspark.sql import functions as F mapping_dict = {'N': 'North', 'C': 'Central', 'S': 'South'} df_new = df.withColumn('City_New', F.replace(df['City'], mapping_dict))
方法2:结合create_map与coalesce
基于你原来的代码,只需要用coalesce函数判断映射结果:如果映射返回空值(即未匹配字典),就用原字段值替代:
from pyspark.sql import functions as F from itertools import chain mapping_dict = {'N': 'North', 'C': 'Central', 'S': 'South'} mapping_expr = F.create_map([F.lit(x) for x in chain(*mapping_dict.items())]) df_new = df.withColumn('City_New', F.coalesce(mapping_expr[df['City']], df['City']))
方法3:使用when+otherwise链式判断
通过逐个匹配字典中的键值对,最后用otherwise指定未匹配时返回原字段:
from pyspark.sql import functions as F mapping_dict = {'N': 'North', 'C': 'Central', 'S': 'South'} # 初始化when表达式 when_expr = F.lit(None) for key, value in mapping_dict.items(): when_expr = F.when(df['City'] == key, value).otherwise(when_expr) # 最后补充未匹配时返回原字段 df_new = df.withColumn('City_New', when_expr.otherwise(df['City']))
这三种方法都能避免手动补全大量重复映射,其中replace函数代码最简洁,推荐优先使用。
内容的提问来源于stack exchange,提问作者Mohammad
相关产品推荐
相关产品推荐

