PySpark中使用StringIndexer转换列时保留None值的方法问询
解决StringIndexer转换时保留None(Null)值的问题
我来帮你搞定这个问题!你遇到的情况是因为StringIndexer默认会把Null(也就是你说的None)要么直接报错,要么当设置handleInvalid="keep"时给它分配一个额外的数字索引。要实现转换字符串列为数值的同时保留Null值,我们需要调整TransformNominalToNumeric方法的逻辑,让Null值在转换后依然保持为Null。
具体修改方案
直接修改你的PreProcess类里的静态方法,通过先生成索引列、再将原Null值对应的索引转回Null的方式实现需求:
from pyspark.ml.feature import StringIndexer from pyspark.sql.functions import when, col class PreProcess: @staticmethod def TransformNominalToNumeric(df, column_name): # 初始化StringIndexer,设置handleInvalid="keep"避免Null值报错 indexer = StringIndexer( inputCol=column_name, outputCol=f"{column_name}_temp_index", handleInvalid="keep" ) # 训练索引模型并生成临时索引列 indexer_model = indexer.fit(df) temp_indexed_df = indexer_model.transform(df) # 获取Null值对应的索引(等于模型标签列表的长度) null_mapped_index = len(indexer_model.labels) # 将原列是Null的行,在目标列中保留Null;其他行用索引值替换 processed_df = temp_indexed_df.withColumn( column_name, when(col(column_name).isNull(), None).otherwise(col(f"{column_name}_temp_index")) ).drop(f"{column_name}_temp_index") # 删除临时索引列 return processed_df
逻辑说明
- 避免报错:设置
handleInvalid="keep"是为了让StringIndexer不会因为遇到Null值而终止运行,而是给Null分配一个特殊的索引值(等于模型训练得到的标签列表的长度)。 - 还原Null值:通过
when条件判断,把原列中是Null的行,在处理后的列里重新设置为Null;非Null的行则保留生成的索引数值。 - 清理临时列:最后删除临时生成的索引列,让原列直接替换为处理后的数值列,保持数据结构整洁。
你的循环代码无需修改
你原来遍历字符串列并调用转换方法的逻辑完全可以保留,因为修改后的TransformNominalToNumeric已经自动处理了Null值的保留:
for columnName, columnType in self.rawData.dtypes: if columnType == "string": self.rawData = PreProcess.TransformNominalToNumeric(self.rawData, columnName)
内容的提问来源于stack exchange,提问作者Tomas Goffa
相关产品推荐
相关产品推荐

