如何在PySpark中实现指定取值的选择性独热编码
自定义指定维度独热编码实现
你可以根据使用的计算框架选择对应代码,逻辑完全匹配需求:仅为列表内指定的apps_id生成编码列,值匹配标记1,不匹配(含apps_id不在列表内的场景)标记0。
PySpark 实现(对应你给出的DataFrame表格式场景)
# 1. 初始数据准备(替换为你自己的数据集即可) df = spark.createDataFrame([ (7445640, '146'), (5592981, '929'), (5103715, '929'), (386222, '114'), (7674331, '146') ], schema=["id", "apps_id"]) # 可自由编辑的指定apps_id列表 list_selected_apps_id = ['146', '929'] # 2. 核心处理:遍历列表逐列生成匹配标记 import pyspark.sql.functions as F for app_id in list_selected_apps_id: df = df.withColumn( app_id, F.when(F.col("apps_id") == app_id, 1).otherwise(0) ) # 3. 筛选输出列,得到最终结果 result_df = df.select("id", *list_selected_apps_id) result_df.show()
运行后输出结果和预期完全一致:
+-------+---+---+ | id|146|929| +-------+---+---+ |7445640| 1| 0| |5592981| 0| 1| |5103715| 0| 1| | 386222| 0| 0| |7674331| 1| 0| +-------+---+---+
Pandas 实现(本地表格处理场景)
import pandas as pd # 1. 初始数据准备 df = pd.DataFrame([ (7445640, '146'), (5592981, '929'), (5103715, '929'), (386222, '114'), (7674331, '146') ], columns=["id", "apps_id"]) list_selected_apps_id = ['146', '929'] # 2. 逐列生成编码 for app_id in list_selected_apps_id: df[app_id] = (df["apps_id"] == app_id).astype(int) # 3. 筛选结果 result_df = df[["id", *list_selected_apps_id]] print(result_df.to_string(index=False))
注意事项:处理前请确认原始数据中apps_id字段的类型和列表内元素类型一致(比如原始字段是整型的话,列表内元素要改为[146, 929]),避免类型不匹配导致判断逻辑失效。
内容的提问来源于stack exchange,提问作者Nabih Bawazir
相关产品推荐
相关产品推荐

