PySpark文档TF-IDF处理后提取Top10词汇及对应分值的实现方法咨询
实现方案
推荐优先使用CountVectorizer方案,可避免HashingTF的哈希碰撞问题,映射关系更稳定。
方案1:CountVectorizer实现(推荐)
CountVectorizer训练完成后可直接获取全局词表,索引与词汇一一对应,无需额外映射计算。
from pyspark.ml.feature import CountVectorizer, IDF from pyspark.sql.functions import udf from pyspark.sql.types import ArrayType, StructType, StructField, StringType, FloatType # 替换原有HashingTF为CountVectorizer cv = CountVectorizer(inputCol="words", outputCol="rawFeatures") cv_model = cv.fit(wordsData) featurizedData = cv_model.transform(wordsData) # 训练IDF模型 idf = IDF(inputCol="rawFeatures", outputCol="features") idfModel = idf.fit(featurizedData) rescaledData = idfModel.transform(featurizedData) # 获取全局词表,索引对应对应词汇 vocab = cv_model.vocabulary # 定义提取topN TF-IDF词汇的UDF def get_top_tfidf(vector, top_n=10): indices = vector.indices values = vector.values # 配对索引与分值,倒序排序取前N idx_value_sorted = sorted(zip(indices, values), key=lambda x: x[1], reverse=True)[:top_n] # 索引转词汇 return [(vocab[idx], float(val)) for idx, val in idx_value_sorted] # 定义UDF返回结构 return_schema = ArrayType(StructType([ StructField("word", StringType(), nullable=False), StructField("tfidf_score", FloatType(), nullable=False) ])) top_tfidf_udf = udf(get_top_tfidf, return_schema) # 生成top10词汇列 result_df = rescaledData.withColumn("top10_tfidf", top_tfidf_udf("features")) # 查看结果 result_df.select("label", "words", "top10_tfidf").show(truncate=False)
方案2:HashingTF实现
HashingTF存在哈希碰撞风险,仅适合精度要求不高的场景,需要遍历当前文档的词汇完成映射:
# 沿用你已训练好的hashingTF与rescaledData def get_top_tfidf_hash(vector, words, top_n=10): idx_score_map = dict(zip(vector.indices, vector.values)) word_score = [] used_words = set() for word in words: if word in used_words: continue used_words.add(word) word_idx = hashingTF.indexOf(word) score = idx_score_map.get(word_idx, 0.0) word_score.append((word, float(score))) return sorted(word_score, key=lambda x: x[1], reverse=True)[:top_n] top_hash_udf = udf(get_top_tfidf_hash, return_schema) result_df = rescaledData.withColumn("top10_tfidf", top_hash_udf("features", "words"))
扩展说明
如果需要将topN的词汇和分值拆成单独的列,可通过getItem方法实现,例如:
# 提取第一名的词汇和分值 result_df = result_df \ .withColumn("top1_word", result_df.top10_tfidf[0].getItem("word")) \ .withColumn("top1_score", result_df.top10_tfidf[0].getItem("tfidf_score"))
内容的提问来源于stack exchange,提问作者Aquen
相关产品推荐
相关产品推荐

