You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark文档TF-IDF处理后提取Top10词汇及对应分值的实现方法咨询

实现方案

推荐优先使用CountVectorizer方案,可避免HashingTF的哈希碰撞问题,映射关系更稳定。

方案1:CountVectorizer实现(推荐)

CountVectorizer训练完成后可直接获取全局词表,索引与词汇一一对应,无需额外映射计算。

from pyspark.ml.feature import CountVectorizer, IDF
from pyspark.sql.functions import udf
from pyspark.sql.types import ArrayType, StructType, StructField, StringType, FloatType

# 替换原有HashingTF为CountVectorizer
cv = CountVectorizer(inputCol="words", outputCol="rawFeatures")
cv_model = cv.fit(wordsData)
featurizedData = cv_model.transform(wordsData)

# 训练IDF模型
idf = IDF(inputCol="rawFeatures", outputCol="features")
idfModel = idf.fit(featurizedData)
rescaledData = idfModel.transform(featurizedData)

# 获取全局词表,索引对应对应词汇
vocab = cv_model.vocabulary

# 定义提取topN TF-IDF词汇的UDF
def get_top_tfidf(vector, top_n=10):
    indices = vector.indices
    values = vector.values
    # 配对索引与分值,倒序排序取前N
    idx_value_sorted = sorted(zip(indices, values), key=lambda x: x[1], reverse=True)[:top_n]
    # 索引转词汇
    return [(vocab[idx], float(val)) for idx, val in idx_value_sorted]

# 定义UDF返回结构
return_schema = ArrayType(StructType([
    StructField("word", StringType(), nullable=False),
    StructField("tfidf_score", FloatType(), nullable=False)
]))
top_tfidf_udf = udf(get_top_tfidf, return_schema)

# 生成top10词汇列
result_df = rescaledData.withColumn("top10_tfidf", top_tfidf_udf("features"))
# 查看结果
result_df.select("label", "words", "top10_tfidf").show(truncate=False)

方案2:HashingTF实现

HashingTF存在哈希碰撞风险,仅适合精度要求不高的场景,需要遍历当前文档的词汇完成映射:

# 沿用你已训练好的hashingTF与rescaledData
def get_top_tfidf_hash(vector, words, top_n=10):
    idx_score_map = dict(zip(vector.indices, vector.values))
    word_score = []
    used_words = set()
    for word in words:
        if word in used_words:
            continue
        used_words.add(word)
        word_idx = hashingTF.indexOf(word)
        score = idx_score_map.get(word_idx, 0.0)
        word_score.append((word, float(score)))
    return sorted(word_score, key=lambda x: x[1], reverse=True)[:top_n]

top_hash_udf = udf(get_top_tfidf_hash, return_schema)
result_df = rescaledData.withColumn("top10_tfidf", top_hash_udf("features", "words"))

扩展说明

如果需要将topN的词汇和分值拆成单独的列,可通过getItem方法实现,例如:

# 提取第一名的词汇和分值
result_df = result_df \
    .withColumn("top1_word", result_df.top10_tfidf[0].getItem("word")) \
    .withColumn("top1_score", result_df.top10_tfidf[0].getItem("tfidf_score"))

内容的提问来源于stack exchange,提问作者Aquen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 01:21:02