You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用PySpark为DataFrame列中重复条目分配唯一ID?

为PySpark DataFrame的重复名称分配从0开始的唯一ID

可以直接使用PySpark MLlib中的StringIndexer工具实现,这是官方针对类别列生成索引的内置方法,比手动去重关联更简洁高效,能自动为重复名称分配相同的从0开始的ID。

实现步骤

  1. 导入所需类
from pyspark.ml.feature import StringIndexer
  1. 创建示例DataFrame
data = [("Alice",), ("Bob",), ("Alice",), ("Chloe",), ("Chloe",)]
df = spark.createDataFrame(data, ["name"])
  1. 使用StringIndexer生成ID列
# 初始化索引器,指定输入列、输出列,可自定义ID分配的排序规则
indexer = StringIndexer(
    inputCol="name",
    outputCol="id",
    stringOrderType="alphabetAsc"  # 按名称字母升序分配ID,匹配你要的示例结果
)

# 拟合数据并转换原DataFrame
indexed_df = indexer.fit(df).transform(df)
  1. 查看结果
indexed_df.show()

执行后会得到与你预期完全一致的结果:

+------+----+
| name | id |
+------+----+
|Alice | 0.0|
|Bob   | 1.0|
|Alice | 0.0|
|Chloe | 2.0|
|Chloe | 2.0|

自定义ID分配规则

如果需要调整ID的分配顺序,可修改stringOrderType参数:

  • frequencyDesc:按名称出现频率降序分配(默认规则)
  • frequencyAsc:按名称出现频率升序分配
  • alphabetDesc:按名称字母降序分配

类型转换(可选)

生成的id列默认是double类型,若需要转为整数类型,可添加以下代码:

from pyspark.sql.functions import col
indexed_df = indexed_df.withColumn("id", col("id").cast("int"))

内容的提问来源于stack exchange,提问作者Algorithman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 03:46:03