You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在PySpark中实现指定取值的选择性独热编码

自定义指定维度独热编码实现

你可以根据使用的计算框架选择对应代码,逻辑完全匹配需求:仅为列表内指定的apps_id生成编码列,值匹配标记1,不匹配(含apps_id不在列表内的场景)标记0。

PySpark 实现(对应你给出的DataFrame表格式场景)

# 1. 初始数据准备(替换为你自己的数据集即可)
df = spark.createDataFrame([
    (7445640, '146'),
    (5592981, '929'),
    (5103715, '929'),
    (386222, '114'),
    (7674331, '146')
], schema=["id", "apps_id"])

# 可自由编辑的指定apps_id列表
list_selected_apps_id = ['146', '929']

# 2. 核心处理:遍历列表逐列生成匹配标记
import pyspark.sql.functions as F
for app_id in list_selected_apps_id:
    df = df.withColumn(
        app_id,
        F.when(F.col("apps_id") == app_id, 1).otherwise(0)
    )

# 3. 筛选输出列,得到最终结果
result_df = df.select("id", *list_selected_apps_id)
result_df.show()

运行后输出结果和预期完全一致:

+-------+---+---+
|     id|146|929|
+-------+---+---+
|7445640|  1|  0|
|5592981|  0|  1|
|5103715|  0|  1|
| 386222|  0|  0|
|7674331|  1|  0|
+-------+---+---+

Pandas 实现(本地表格处理场景)

import pandas as pd

# 1. 初始数据准备
df = pd.DataFrame([
    (7445640, '146'),
    (5592981, '929'),
    (5103715, '929'),
    (386222, '114'),
    (7674331, '146')
], columns=["id", "apps_id"])

list_selected_apps_id = ['146', '929']

# 2. 逐列生成编码
for app_id in list_selected_apps_id:
    df[app_id] = (df["apps_id"] == app_id).astype(int)

# 3. 筛选结果
result_df = df[["id", *list_selected_apps_id]]
print(result_df.to_string(index=False))

注意事项:处理前请确认原始数据中apps_id字段的类型和列表内元素类型一致(比如原始字段是整型的话,列表内元素要改为[146, 929]),避免类型不匹配导致判断逻辑失效。

内容的提问来源于stack exchange,提问作者Nabih Bawazir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 16:45:43